TicketFlow's pipeline. No solutions.md before submitting.
Exercise 1 — The PR pipeline
- Write the
ci.ymlwith the 6 gate jobs (lint, unit, integration, e2e, build+scan, check-deploy) with their parallelism: which run in parallel and which in sequence? Which is the pipeline's critical path (the job determining total duration)? - The cache: add the pip cache (requirements.lock's hash as key) and the Docker image cache by registry. Measure: first run vs second — paste the times per job.
- The flaky test in the gate: inject a test failing 1 out of 5 times (random) and watch the effect over 5 simulated PRs. How many re-runs were needed? Document the project's policy: manual re-run + mandatory issue (35).
Exercise 2 — The build with digest
- Configure the push with tag
:git-<sha>and store the push's digest in the job's output. Verify:docker pullof the digest works and the digest'sdocker inspectmatches the local one. - The scan gate: trivy with
--exit-code 1 --severity CRITICALover an image with a simulated critical CVE (an old base pinned on purpose): does the pipeline go red? And the job doesn't block for the report's MEDIUMs? - The auditable chain: write the command that, given a prod environment, returns the exact commit it runs (digest → git-sha tag → commit). How many commands did it cost you? (ideal answer: 1).
Exercise 3 — The deployment
- Implement the 5 deploy scripts (deploy.sh, wait_healthy.sh, switch_traffic.sh, smoke.sh, rollback.sh) for a simulated local deployment (two "blue" and "green" composes with nginx in front). Paste the hot-switch log with no 500s.
- The tested rollback: make the smoke fail on purpose (healthcheck green but the listing smoke red): did the automatic rollback fire and traffic return to blue? Measure the total rollback time (target: <15 s).
- The strategy decision: with your 36 metrics (p95, errors) and your infra cost, write the decision paragraph: blue-green or canary for TicketFlow, and which 46 metric would fire the automatic rollback if you moved to canary tomorrow?
Exercise 4 — Migrations in the release
- Write the
release-migratejob runningmanage.py migratefrom the release image against staging's DB, and the guard: the job fails if a pending migration is NOT an expand (11's migration linter in the pipeline). - Simulate the full release in staging: migrate expand → deploy green → smoke → switch → (contract pending). Document the intermediate state: does the old app (blue) work against the expanded schema? (11's test proving it).
- The two NEVERs: (a) try deploying the new app against the unmigrated schema: which error and how many 500s until the rollback? (b) do a deploy with
migrate && runserverin the container: what does it complicate in the rollback? Document both experiences as runbook warnings.
Exercise 5 — The circle
- Add the deploy annotations to the monitoring (45's structured log or Grafana's API if you have it): deployment_started/completed with commit+digest+actor. Verify the p95 graph shows the deploy's vertical line.
- The project's 4 DORA metrics: compute by hand (or from the GH Actions log) lead time, frequency, failure rate, MTTR of the last 10 deploys. Which is healthy and which is the team's shame?
- The
docs/ci-cd.mddocument: the pipeline's diagram (ASCII), the gates, the deployment strategy with its reason, and the rollback policy. Max 1 page — the document the next you reads at 3 a.m. with a broken deploy.
Submit
Paste the ci.yml with its times, the auditable digest evidence, the switch + rollback log, and the 4 DORA metrics. Next: Lesson 42 — Cloud.