Stack: Django + DRF · Project: TicketFlow · OS: Linux Status: Published — taught after Lesson 04 is corrected Prerequisite: Lesson 04 — Data structures
Objectives
By the end of this lesson you will be able to:
- See, understand and operate your server's processes (Gunicorn, Celery, Postgres) with
ps,top/htopand signals. - Manage permissions and users with production judgement (principle of least privilege).
- Read logs with
journalctl,tail -f,grep,awkand find the error among millions of lines. - Configure the environment with variables (12-factor preview) without leaking secrets.
- Connect to production over SSH with keys and create secure tunnels for debugging.
1. Processes: see them, measure them, signal them
Your TicketFlow in production is several processes: Gunicorn (2-4), Celery (1-4), PostgreSQL, Redis, Nginx. Commands of the craft:
ps aux | grep gunicorn # who I am, how much memory (RSS), how much CPU
top # live; P sorts by CPU, M by memory
htop # the good top: trees and colors
ss -tlnp # who listens on which port (the modern "netstat")
uptime # load average: 1, 5 and 15 minutesThe load average reads like this: on a 2 vCPU machine, a load of 2.00 is "just full", 4.00 is "twice as much work as cores" (processes are waiting). If uptime says a sustained 8.00, your API is slow even though no error appears.
Signals: the polite way to talk to a process:
| Signal | What it does | Typical use |
|---|---|---|
| SIGTERM (15) | graceful shutdown | kill <pid>; Gunicorn finishes in-flight requests |
| SIGKILL (9) | violent shutdown | last resort; the process never even knows |
| SIGHUP (1) | reload config | Gunicorn uses it to re-read without dropping the service |
| SIGUSR1/2 | app-defined | rotate Gunicorn logs, debug workers |
Key:
kill -9is surrender: the process cannot save anything (logs, in-flight transactions). SIGTERM first, then wait; SIGKILL only if it is truly hung. Restarting withsystemctl restart gunicornsends SIGTERM and waits properly — that is why production services live in systemd, not innohup.
2. Permissions and users: least privilege
Production rule: your app does not run as root. Ever.
sudo useradd -r -s /usr/sbin/nologin ticketflow # system user, no shell
sudo chown -R ticketflow:ticketflow /srv/ticketflow
# in the systemd service: User=ticketflowPermissions are 3 groups × 3 bits: rwx for user/group/others. 755 = me everything, others read and execute. 600 = only I read/write (for secrets). 640 for .env: owner reads, group reads, others nothing.
chmod 600 /srv/ticketflow/.env # secrets: only the app's user
chmod 750 /srv/ticketflow # nobody else on the system can list your code
ls -la # look who is who before touching
umask 027 # prudent default permissionsClassic mistakes (and their consequences): .env with 644 (any server user reads your keys); running Gunicorn as root (one RCE in your app = the whole server compromised); chmod -R 777 as a "quick fix" (a fix that is itself the vulnerability).
3. Logs: finding the needle
A healthy server generates millions of lines. The cutting tools:
journalctl -u gunicorn -f # follow the service live
tail -n 100 error.log # the latest
grep -i "error" gunicorn-error.log # filter
grep -i error gunicorn-error.log | awk '{print $1}' | sort | uniq -c | sort -rn | headThe pipeline above is the diagnostic command: it counts errors grouped by the first column, sorted most-first. With it you know whether you have one error repeated 5000 times (a bug) or 5000 different errors (a storm).
What a useful log line contains (we formalize it in Lesson 45):
2026-09-28T10:14:22Z INFO [req_id=ab12cd] POST /api/reservations/ 201 45ms user=1042Level, UTC timestamp, correlation id, action, result, duration, actor. Without req_id, debugging a multi-service incident is divination.
And the retention rule: logs rotate (logrotate, journald with size limits) or they fill the disk — and a full disk is the #2 way to go down (the #1 is the DB growing without control; you will see it in Module 2).
4. Environment variables: the configuration that is not versioned
Keys, passwords and per-environment settings live in the environment, not in the code (12-factor: full Lesson 27):
# .env (outside Git: .gitignore already does it since Lesson 00)
DEBUG=0
SECRET_KEY=...
DATABASE_URL=postgres://ticketflow:...@localhost/ticketflow
CELERY_BROKER_URL=redis://localhost:6379/0
# load them in your dev shell
set -a; source .env; set +aIn Python: os.environ["SECRET_KEY"] (and with django-environ or pydantic-settings, validated at boot). Two hard rules:
- A versioned
.env.examplewith all the names and no real values. Whoever joins the project copies, fills in and works. - Secrets never go into logs or tracebacks:
DEBUG=1in production prints your configuration to the world (that is whyDEBUG=0is the first check of any audit).
5. SSH: the door to production
Key-based access, never passwords:
ssh-keygen -t ed25519 -C "you@email" # one key per machine (or per project)
ssh-copy-id user@ticketflow.app # install the public key on the server
ssh -i ~/.ssh/ed25519 user@ticketflow.app~/.ssh/config so you don't have to type:
Host prod
HostName ticketflow.app
User deploy
IdentityFile ~/.ssh/ed25519And the two tricks that change your backend life:
# Tunnel: the production DB on your localhost (read-only, with your app user)
ssh -L 5433:localhost:5432 prod
# now: psql -h localhost -p 5433 -U ticketflow ticketflow
# Execute without opening an interactive shell (for deployment scripts)
ssh prod "systemctl status gunicorn --no-pager"Key: the tunnel is for debugging with permission; production data access is governed by GDPR and company policy (Lesson 56). Less access than you need, and always auditable.
6. Scripting: automate what you do three times
Everything above combines into small, readable scripts. A real example: checking TicketFlow's full health:
#!/usr/bin/env bash
# healthcheck.sh — TicketFlow health in 4 checks
set -euo pipefail # fail fast, mandatory variables, strict pipes
for url in "http://127.0.0.1:8000/api/health/" "http://127.0.0.1:8000/api/events/"; do
code=$(curl -s -o /dev/null -w "%{http_code}" "$url")
if [[ "$code" != 200 ]]; then
echo "FAIL $url -> $code" >&2
exit 1
fi
done
systemctl is-active --quiet gunicorn celery
echo "OK: web and workers alive"The minimum bash that avoids 90% of the scares: set -euo pipefail (exit on the first error), quotes on every variable ("$var"), and [[ ]] instead of [ ]. And the rule of the craft: if you have done it by hand three times, it is a script; if a cron runs it, it is a job with logs and alerting (Lesson 31).
Self-assessment (answer me in the chat)
- Why
systemctl restart gunicornand notkill -9on the workers? What is lost with SIGKILL? - Your
/srv/ticketflow/.envhas 644 permissions and the server has 5 users. What concrete risks exist and how do you close them? - The pipeline
grep... | awk | sort | uniq -c | sort -rn | head: explain each link and which question it answers in an incident. - Why is
DEBUG=1in production the first finding of any audit? What exactly leaks? - Design TicketFlow's minimal
deploy.sh: pull, install deps, migrate, collect static, reload services. Which order did you choose and why does "migrate" come before reloading?
Continue with the exercises. The solutions only after trying it yourself.