Module 9 · Deployment and operations

Lesson 40 — Docker

Images, multi-stage Dockerfiles and compose for TicketFlow development.

Published
In this lesson
  1. Exercise 1 — The Dockerfile
  2. Exercise 2 — The secret and the CVE
  3. Exercise 3 — The compose
  4. Exercise 4 — Migrations
  5. Exercise 5 — The production image
  6. Professor's summary

Exercise 1 — The Dockerfile

  1. The typical saving: the naïve image (FROM python:3.12 + pip install + COPY.) ≈ 1.1 GB → multi-stage slim ≈ 210 MB (−80%). Every deployment pull goes from ~40 s to ~7 s: ×15 deploys/day × 5 min saved — size is deployment speed, not aesthetics.
  1. The evidence:
$ docker run --rm ticketflow whoami
10001
$ docker run --rm ticketflow touch /usr/x
touch: cannot touch '/usr/x': Permission denied

With 22 in hand: the attacker executing code in your container inherits uid 10001 — with no write permission on the container's filesystem (and with read_only: true in prod, the whole root filesystem), persistence (cron, replaced binaries) gets radically harder. It is not the wall (that's the host, 43): it is the entry price.

  1. Reproducibility: two consecutive builds → the SAME image ID (build cache produces identical bytes). Changing a .py comment invalidates ONLY the COPY.. layer and the ones after (collectstatic, useradd): requirements layers survive (the builder's pip cache intact). The build's hash changes (the content changed — correct), but the CHANGE is minimal and traceable: the image is the project's reproducible binary.

Exercise 2 — The secret and the CVE

  1. The disaster's evidence (COPY without dockerignore with a.env present):
$ docker history ticketflow:naif --no-trunc | grep -i secret
COPY . . → includes .env with PAYMENT_GATEWAY_KEY=sk_live_...

The history shows the COPY layer with the whole context; docker save/export give the literal bytes: the gateway's live secret lives in the Docker registry forever (until the registry's docker image prune, which nobody does). The fix: .dockerignore + never COPY.. without it + secrets ONLY via --env-file/orchestrator (27). The clean history grep is this lesson's regression test.

  1. The CVE triage (3 model findings): CVE-2024-XXXX libssl 1.3 (DoS) — relevant if TLS terminates IN the container (here it doesn't: 43's proxy terminates TLS → operational noise, patched with the base anyway); CVE-2024-YYYY pip (malicious installation) — doesn't apply at runtime (no pip: it stayed in builder); CVE-2024-ZZZZ glibc (local privilege escalation) — RELEVANT: combined with any app's remote-code vector, it escalates to the host. The criterion: exploitable IN my context (is the component exposed? does it demand prerequisites that don't exist?) — a high CVSS without context is a shopping list, not a triage.
  1. Pinning by digest: "slim" is a MOVING tag — any random Tuesday Debian ships a patch and the tag points to new bytes: today's build ≠ yesterday's build WITHOUT you changing anything (reproducibility breaks silently; the CI scan that passed yesterday fails today with no commit). The digest freezes the bytes: the base upgrade is ONE conscious commit with the new digest — which 41 automates (the bot that PRs the new digest with the scan green).

Exercise 3 — The compose

  1. The boot log:
db-1      | 2027-09-28 10:00:01 [1] LOG:  database system is ready to accept connections
redis-1   | Ready to accept connections
db-1      | [healthcheck] pg_isready OK (retries 1/10)
web-1     | Waiting for db: service_healthy...
web-1     | ✔ db healthy → starting gunicorn

"Started" ≠ "healthy": PG accepts connections BEFORE finishing recovery; the healthcheck (pg_isready) asks what the app needs to know. Without it, web boots, the first query explodes with OperationalError and the restart policy spins the boot roulette.

  1. Hot-reload worked (the ./ticketflow:/app/ticketflow bind mount and runserver's autoreloader); the edited requirements.txt + docker compose build: the RUN pip install step invalidated ONLY if the COPY file's hash changed — changing code does NOT touch the deps layer: the 90 s install survives the 2 s code deploy. The COPY order IS the development cache policy.
  1. The worker shell's use: migrate by hand (python manage.py shell) against compose's real DB to debug the reservation with locks (10), run 24's FakeClock in a realistic environment, and "reproduce the customer's bug with their anonymized data (23)" — the shell's environment IS the worker's environment: the job's bug reproduces where it lives.

Exercise 4 — Migrations

  1. The breaking evidence:
web-1 | Applying events.0012_add_status_index...
web-2 | Applying events.0012_add_status_index... (LOCKED: waiting for web-1)
web-3 | Applying events.0012_add_status_index... (LOCKED)
web-1 | OK
web-2 | OK

django_migrations' lock (10) avoids corruption (only one applies), but: the other 2 block boot for 30-120 s (11's CONCURRENTLY index is the worst: the CREATE INDEX lock inside migrate serializes all three), and if the migration is NOT idempotent (a badly written data migration), the second runner DUPLICATES data. The annoyance is not theoretical.

  1. The grown-up pattern: migrate as a JOB (CI/CD runs it against the DB once, 41), then --scale web=3 boots three gunicorns that ONLY serve: boot logs with no "Applying migrations" and the first 200 in <2 s. 11's mandate is met: migration and deployment are separate, observable phases.
  1. The clean boot: docker compose restart web → healthcheck + gunicorn bind: first 200 in ~1.5 s. What it does NOT wait for: neither DB nor Redis (the compose's healthchecks only apply at compose START; on a restart the container trusts the living infra — and if the infra died, 26's 500 says so, not a crash loop).

Exercise 5 — The production image

  1. The weight top: /usr/local/lib/python3.12/site-packages (Django+DRF+Celery+psycopg: 142 MB — the venv IS the app, no trimming it), collectstatic's statics (18 MB: not served from the app in prod, the CDN/white noise carries them — they can leave the image if the CDN receives them by pipeline upload, 41), libc + libpq5 (~40 MB, irreducible on slim). Conclusion: distroless (no shell, no apt) would save ~30 MB and LOSES the debug shell (exercise 3's compose run --rm web shell): slim with no-root is the sweet spot for a team of 1.
  1. The E2E from zero: build (cached: 15 s) + up + healthchecks (8 s) + seed (3 s) + pytest -m e2e (160 s) ≈ 3 min per PR: fits in CI (41) with the build cache uploaded as an artifact. The number decides the E2E policy: without the reproducible environment, "run E2E on every PR" is a wish.
  1. The ADR (excerpt): "Multi-stage image python:3.12-slim@digest, user 10001, /healthz/ healthcheck, config per environment. Reasons: reproducible (digest), secure (no-root + CI scan), fast (210 MB). Discarded: one image with everything (1.1 GB, secrets in history) and distroless (loses the debug shell a team of 1 uses daily). Review: if the base's scan accumulates relevant CVEs, migrate the base or go distroless + debug sidecar".

Professor's summary

  • Multi-stage + lockfile + no-root + pinned digest: a small, secure, reproducible image; the history without secrets is the regression test.
  • The dev compose is the deployment's prototype: healthchecks as conditions (not wishes), bind mounts for hot-reload, one image with per-service commands.
  • Migrations are a deployment JOB, not the container's boot: django_migrations' lock saves the data, not the peace.