Exercise 1 — The pipeline
- The parallelism:
lint+unit+integration+buildrun IN PARALLEL when the PR opens;e2ewaits forbuild(it needs the image);check-deployruns in parallel with the tests (it doesn't depend on code executing, only on settings). The critical path: e2e (~4 min with the image cached; without cache, the 90 s build + e2e). The PR's total pipeline: ~6-8 min cold, ~4-5 min warm — the number deciding whether the team pushes several times a day or every two hours.
- The caches (second run):
lint: 31 s → 9 s (pip cache)
unit: 28 s → 11 s (pip cache + .pytest_cache)
integration: 3m12 → 2m40 (services already pulled)
build: 88 s → 14 s (cache-from registry: the deps layers reused)
e2e: 4m10 → 3m55 (image from cache; the E2E is never cached: it ALWAYS runs)- The flaky's effect: 5 PRs × 1/5 failure chance = ~64% probability of at least one red in the day — the observed mean: 3 of 5 PRs needed 1 re-run, 1 needed 2. The invisible cost: trust — by the third re-run nobody looks anymore. The project's policy: manual re-run (never automatic in the gate), immediate issue, quarantine with expiry (35).
Exercise 2 — The digest
- The chain:
docker pushreturnssha256:9f2c...; the job exports it (::set-output) and the deploy consumes it.docker pull ticketflow@sha256:9f2c...→docker inspectshows the same digest: the artifact's identity is the digest, not the tag (the tag is a mutable alias; the digest, immutable).
- The scan gate: the old pinned base → trivy
--exit-code 1 --severity CRITICAL→ job RED with the CVE report. The MEDIUMs: to the report (job artifact), not the gate — 40's triage: a high CVSS without context blocks important deploys over noise; the exploitable CRITICAL never passes.
- The audit command: 1 command:
docker inspect --format '{{index .RepoDigests 0}}' prod-app | xargs -I{} docker image inspect {} --format '{{index .Config.Labels "org.opencontainers.image.revision"}}'With the org.opencontainers.image.revision label baked at build (the standard OCI label): digest → commit sha. One command, full chain — 47's audit ("what code ran when the customer called?") is one line.
Exercise 3 — The deployment
- The switch log:
10:14:02 deploy.sh green → green up (image a1b2c3d)
10:14:18 wait_healthy green → /healthz/ OK ×3
10:14:22 switch_traffic.sh → nginx upstream → green
10:14:23 smoke.sh: /events 200, /reservations 201, /healthz 200 → GREEN CONFIRMED
0 500 responses during the switch (nginx drains blue's keepalive connections)- The measured rollback: the smoke red →
rollback.sh→ switch back to blue → traffic restored in 11 s. A deployment without a tested rollback is a bet: the exercise costs 20 minutes and buys the peace of mind of ALL future deploys.
- The decision: blue-green (canary demands percentage routing + per-version metrics the local nginx doesn't give; TicketFlow's 2× infra cost is ~€30/month — the 11 s rollback pays for it). The paragraph: "Blue-green: rollback <15 s tested, 2× infra accepted. If we move to canary: the automatic rollback fires on
http_5xx_rate > 2%for 2 min orp95 > 2× baseline(46), with the canary alarm as its owner".
Exercise 4 — Migrations
- The release guard: the job runs
python manage.py makemigrations --check(nobody forgets to generate them) + 11's linter (django-migration-linter) with--fail-on-danger: a migration renaming a column without expand-contract reds the job with the remedy in the message.
- The intermediate state demonstrated: 11's test runs the suite against the EXPANDED schema with the OLD CODE (blue app's checkout): green — the new column is nullable/with no breaking default, the old code never sees it. That is the proof that expand-contract allows blue/green coexistence.
- The NEVERs documented: (a) new app without schema:
column "org_id" does not existon the first request → 500s in a chain (~14 in 3 s until rollback) — migrate BEFORE deploy is the law; (b)migrate && runserverin the container: the rollback redeploys blue... which RE-RUNS the boot migrate (again, with a lock, slowly) and, worse, the rollback cannot "not migrate": the migration is already applied and the blue code with the new schema coexists — the rollback stops being atomic. 47's runbook carries both warnings.
Exercise 5 — The circle
- The annotation: the pipeline's webhook calls the log (45):
{"event": "deployment_completed", "commit": "a1b2c3d", "digest": "sha256:9f2c...", "actor": "buffy", "ts": "..."}The p95 graph with the vertical line: 47's "what changed at 14:20?" is answered by hovering.
- The project's 4 DORA over the last 10 deploys: lead time PR→prod: 1.8 h (healthy); frequency: 2.1/day (healthy); deploy failure rate: 20% (2 of 10: the smoke caught them — a DETECTED deploy failure is nearly free; the shame would be the 502 in prod); MTTR: 14 min (the 11 s rollback + the diagnosis). The team of 1's shame: Friday-evening deploys keep happening — the document forbids them and the calendar doesn't comply.
docs/ci-cd.md(the page's skeleton): the pipeline's ASCII with gates → strategy (blue-green, reason: rollback <15 s) → the release's canonical order (migrate expand → deploy → smoke → switch → contract) → rollback policy → exercise 4's two NEVERs → monthly DORA. One page the 3 a.m. you reads top to bottom without stops.
Professor's summary
- The merge gate is green and deterministic; the slow runs after: the pipeline protects the feedback cycle as much as the code.
- The digest is the release's artifact: commit→image→environment auditable in one command; the deploy deploys, never rebuilds.
- Blue-green + expand-contract + tested rollback: the pair turning deployment into a boring event — and boring is the goal.