The terminal is the definitive diagnostic tool: these exercises are muscle, not theory.
Exercise 1 — Anatomy of your processes
- The parent (arbiter) shows the base command
gunicorn config.wsgi --bind 127.0.0.1:8000; the workers show[1]/[2](or similar) after the command. The parent's RSS is small; the workers carry your app (more memory). - With SIGTERM the worker finishes politely (it completes the in-flight request) and the arbiter replaces it with a new PID: the service never stops answering. That is the mechanism of zero-downtime reload.
- With SIGKILL the worker dies without cleaning up: the in-flight request is cut (the client sees a closed connection) and the arbiter detects the dead one and raises another. Key difference: SIGTERM respects what was being done; SIGKILL doesn't even ask. That is why deployments use TERM and orchestrators (systemd, Docker) wait a graceful timeout before sending KILL.
Exercise 2 — Ports and load
- Gunicorn (or
runserver) on 8000; PostgreSQL on 5432; Redis on 6379. If something doesn't show up, that is your first real diagnosis. - The load spreads across the 2 workers (and
runserverif you used it: fine note,runserveris development and single-threaded; the numbers don't transfer to production). Thecurls compete with each other; with 200 launched in parallel you will see contention on CPU or on the socket's accept queue. - The three numbers are load at 1, 5 and 15 min. On an n-core machine, load > n means work waiting. Trend matters: if 1min > 15min, the system is getting worse right now; the trend matters more than the point value.
Exercise 3 — The log pipeline
$9is the ninth field of the combined access log: the HTTP status code. The pipeline answers "how many responses of each code do I have?" — in an incident, the first question: are 200s dominating (all good) or are 500s spiking (our bug)?- If the format has the duration in the last field (configurable in Nginx/Gunicorn):
awk '{print $NF, $7}' access.log | sort -rn | head -5(duration + route, descending, top 5). If your log has no duration, adding it is half the exercise: no data, no analysis — the same idea that reappears in Lesson 37 (profiling).
Exercise 4 — Real permissions
- Yes it could: 644 gives read to "others", and
app_testis "other". Any system user reads your secrets. - With 640 and the app's group: owner and that group read; the rest nothing. With 600: only the owner (stricter: use it when no other app shares the group). The hierarchy to remember: every bit you open is one more server user who can read.
chmod 600.envand owner = the user running the app (not root). And in the systemdservice:User=ticketflow, so only that user (and root) can read the file.
Exercise 5 — The well-made.env
- A complete
env.example: the names are contracts; the values, secrets. Ifsettings.pyasks for a variable missing from the example, a teammate's onboarding breaks for no reason. git check-ignore -v.envmust print the rule ignoring it (e.g..gitignore:6:.env). If it prints nothing, your.envis an accident waiting for a commit.- If secrets showed up in history: rotate (change the credential at the provider) — deleting the commit is not enough, the history stays cloned. Rotation is the only real repair. (The history-rewriting tools come in Lesson 06; rotation remains mandatory.)
Exercise 6 — Local SSH
- Yes, it got in without a password: your public key landed in
~/.ssh/authorized_keysand the private one proved identity. Password prompts stop being asked (and you can disablePasswordAuthenticationon the server: the brute-force door closes). ssh localtest hostnameworks the same: the config maps nicknames to host/user/key. With 3 servers you can't live without it.- On port 5433 listens your ssh client (a local process), which forwards whatever it receives through the encrypted tunnel to the server's 5432. With a real remote DB: you point your psql/DBeaver at
localhost:5433and traffic travels encrypted over SSH even though the DB is not exposed to the Internet — the debugging pattern you will use forever.
Exercise 7 — The operations script
- With Gunicorn down, curl returns 000/refused → the script prints FAIL and exits 1 (which is what a monitor reads as an alert). Up: OK.
- Redis down warns but does not fail because it doesn't take the web down (the API keeps serving; you lose cache/queue: degradation, not unavailability). The API down is a failure. Telling severities apart is the difference between a script and an instrument: the exit code is your API towards alerting systems.
- In production: a systemd timer (or cron) runs the script every minute and, on failure, an alerting mechanism kicks in (email, Slack, Prometheus — Lesson 46).
watchis the toy; the timer is the toy with logs, alerts and restart after a crash.
Professor's summary
- Processes are operated with signals: polite TERM, KILL as last resort; Gunicorn's arbiter is your first example of resilience.
- Least privilege: system user, 600 for secrets, and every open bit is one more door.
- The log pipeline answers "what is happening and how many times?" in one shell line.
- The environment lives outside the code; secrets never in Git nor logs; SSH with keys and tunnels for debugging.
- If you did it three times by hand, it is already a script. If it runs on its own, it is already operations.
When you submit, we close Lesson 05 and move to Lesson 06 — Professional Git: rebase, conflicts and the team flow that holds everything else up.