Module 9 · Deployment and operations

Lesson 41 — CI/CD

Pipelines with tests, analysis and automatic deployment on merge.

Published
In this lesson
  1. Exercise 1 — The PR pipeline
  2. Exercise 2 — The build with digest
  3. Exercise 3 — The deployment
  4. Exercise 4 — Migrations in the release
  5. Exercise 5 — The circle
  6. Submit

TicketFlow's pipeline. No solutions.md before submitting.

Exercise 1 — The PR pipeline

  1. Write the ci.yml with the 6 gate jobs (lint, unit, integration, e2e, build+scan, check-deploy) with their parallelism: which run in parallel and which in sequence? Which is the pipeline's critical path (the job determining total duration)?
  2. The cache: add the pip cache (requirements.lock's hash as key) and the Docker image cache by registry. Measure: first run vs second — paste the times per job.
  3. The flaky test in the gate: inject a test failing 1 out of 5 times (random) and watch the effect over 5 simulated PRs. How many re-runs were needed? Document the project's policy: manual re-run + mandatory issue (35).

Exercise 2 — The build with digest

  1. Configure the push with tag :git-<sha> and store the push's digest in the job's output. Verify: docker pull of the digest works and the digest's docker inspect matches the local one.
  2. The scan gate: trivy with --exit-code 1 --severity CRITICAL over an image with a simulated critical CVE (an old base pinned on purpose): does the pipeline go red? And the job doesn't block for the report's MEDIUMs?
  3. The auditable chain: write the command that, given a prod environment, returns the exact commit it runs (digest → git-sha tag → commit). How many commands did it cost you? (ideal answer: 1).

Exercise 3 — The deployment

  1. Implement the 5 deploy scripts (deploy.sh, wait_healthy.sh, switch_traffic.sh, smoke.sh, rollback.sh) for a simulated local deployment (two "blue" and "green" composes with nginx in front). Paste the hot-switch log with no 500s.
  2. The tested rollback: make the smoke fail on purpose (healthcheck green but the listing smoke red): did the automatic rollback fire and traffic return to blue? Measure the total rollback time (target: <15 s).
  3. The strategy decision: with your 36 metrics (p95, errors) and your infra cost, write the decision paragraph: blue-green or canary for TicketFlow, and which 46 metric would fire the automatic rollback if you moved to canary tomorrow?

Exercise 4 — Migrations in the release

  1. Write the release-migrate job running manage.py migrate from the release image against staging's DB, and the guard: the job fails if a pending migration is NOT an expand (11's migration linter in the pipeline).
  2. Simulate the full release in staging: migrate expand → deploy green → smoke → switch → (contract pending). Document the intermediate state: does the old app (blue) work against the expanded schema? (11's test proving it).
  3. The two NEVERs: (a) try deploying the new app against the unmigrated schema: which error and how many 500s until the rollback? (b) do a deploy with migrate && runserver in the container: what does it complicate in the rollback? Document both experiences as runbook warnings.

Exercise 5 — The circle

  1. Add the deploy annotations to the monitoring (45's structured log or Grafana's API if you have it): deployment_started/completed with commit+digest+actor. Verify the p95 graph shows the deploy's vertical line.
  2. The project's 4 DORA metrics: compute by hand (or from the GH Actions log) lead time, frequency, failure rate, MTTR of the last 10 deploys. Which is healthy and which is the team's shame?
  3. The docs/ci-cd.md document: the pipeline's diagram (ASCII), the gates, the deployment strategy with its reason, and the rollback policy. Max 1 page — the document the next you reads at 3 a.m. with a broken deploy.

Submit

Paste the ci.yml with its times, the auditable digest evidence, the switch + rollback log, and the 4 DORA metrics. Next: Lesson 42 — Cloud.