Module 1 · Foundations that hold everything up

Lesson 05 — Linux and the terminal

Processes, permissions, logs, environment variables, SSH and scripting to operate your backend.

Published
In this lesson
  1. Exercise 1 — Anatomy of your processes
  2. Exercise 2 — Ports and load
  3. Exercise 3 — The log pipeline
  4. Exercise 4 — Real permissions
  5. Exercise 5 — The well-made.env
  6. Exercise 6 — Local SSH
  7. Exercise 7 — The operations script
  8. Professor's summary

The terminal is the definitive diagnostic tool: these exercises are muscle, not theory.


Exercise 1 — Anatomy of your processes

  1. The parent (arbiter) shows the base command gunicorn config.wsgi --bind 127.0.0.1:8000; the workers show [1]/[2] (or similar) after the command. The parent's RSS is small; the workers carry your app (more memory).
  2. With SIGTERM the worker finishes politely (it completes the in-flight request) and the arbiter replaces it with a new PID: the service never stops answering. That is the mechanism of zero-downtime reload.
  3. With SIGKILL the worker dies without cleaning up: the in-flight request is cut (the client sees a closed connection) and the arbiter detects the dead one and raises another. Key difference: SIGTERM respects what was being done; SIGKILL doesn't even ask. That is why deployments use TERM and orchestrators (systemd, Docker) wait a graceful timeout before sending KILL.

Exercise 2 — Ports and load

  1. Gunicorn (or runserver) on 8000; PostgreSQL on 5432; Redis on 6379. If something doesn't show up, that is your first real diagnosis.
  2. The load spreads across the 2 workers (and runserver if you used it: fine note, runserver is development and single-threaded; the numbers don't transfer to production). The curls compete with each other; with 200 launched in parallel you will see contention on CPU or on the socket's accept queue.
  3. The three numbers are load at 1, 5 and 15 min. On an n-core machine, load > n means work waiting. Trend matters: if 1min > 15min, the system is getting worse right now; the trend matters more than the point value.

Exercise 3 — The log pipeline

  1. $9 is the ninth field of the combined access log: the HTTP status code. The pipeline answers "how many responses of each code do I have?" — in an incident, the first question: are 200s dominating (all good) or are 500s spiking (our bug)?
  2. If the format has the duration in the last field (configurable in Nginx/Gunicorn):
bash
awk '{print $NF, $7}' access.log | sort -rn | head -5

(duration + route, descending, top 5). If your log has no duration, adding it is half the exercise: no data, no analysis — the same idea that reappears in Lesson 37 (profiling).

Exercise 4 — Real permissions

  1. Yes it could: 644 gives read to "others", and app_test is "other". Any system user reads your secrets.
  2. With 640 and the app's group: owner and that group read; the rest nothing. With 600: only the owner (stricter: use it when no other app shares the group). The hierarchy to remember: every bit you open is one more server user who can read.
  3. chmod 600.env and owner = the user running the app (not root). And in the systemd service: User=ticketflow, so only that user (and root) can read the file.

Exercise 5 — The well-made.env

  1. A complete env.example: the names are contracts; the values, secrets. If settings.py asks for a variable missing from the example, a teammate's onboarding breaks for no reason.
  2. git check-ignore -v.env must print the rule ignoring it (e.g. .gitignore:6:.env). If it prints nothing, your .env is an accident waiting for a commit.
  3. If secrets showed up in history: rotate (change the credential at the provider) — deleting the commit is not enough, the history stays cloned. Rotation is the only real repair. (The history-rewriting tools come in Lesson 06; rotation remains mandatory.)

Exercise 6 — Local SSH

  1. Yes, it got in without a password: your public key landed in ~/.ssh/authorized_keys and the private one proved identity. Password prompts stop being asked (and you can disable PasswordAuthentication on the server: the brute-force door closes).
  2. ssh localtest hostname works the same: the config maps nicknames to host/user/key. With 3 servers you can't live without it.
  3. On port 5433 listens your ssh client (a local process), which forwards whatever it receives through the encrypted tunnel to the server's 5432. With a real remote DB: you point your psql/DBeaver at localhost:5433 and traffic travels encrypted over SSH even though the DB is not exposed to the Internet — the debugging pattern you will use forever.

Exercise 7 — The operations script

  1. With Gunicorn down, curl returns 000/refused → the script prints FAIL and exits 1 (which is what a monitor reads as an alert). Up: OK.
  2. Redis down warns but does not fail because it doesn't take the web down (the API keeps serving; you lose cache/queue: degradation, not unavailability). The API down is a failure. Telling severities apart is the difference between a script and an instrument: the exit code is your API towards alerting systems.
  3. In production: a systemd timer (or cron) runs the script every minute and, on failure, an alerting mechanism kicks in (email, Slack, Prometheus — Lesson 46). watch is the toy; the timer is the toy with logs, alerts and restart after a crash.

Professor's summary

  • Processes are operated with signals: polite TERM, KILL as last resort; Gunicorn's arbiter is your first example of resilience.
  • Least privilege: system user, 600 for secrets, and every open bit is one more door.
  • The log pipeline answers "what is happening and how many times?" in one shell line.
  • The environment lives outside the code; secrets never in Git nor logs; SSH with keys and tunnels for debugging.
  • If you did it three times by hand, it is already a script. If it runs on its own, it is already operations.

When you submit, we close Lesson 05 and move to Lesson 06 — Professional Git: rebase, conflicts and the team flow that holds everything else up.