Module 9 · Deployment and operations

Lesson 41 — CI/CD

Pipelines with tests, analysis and automatic deployment on merge.

Published
In this lesson
  1. Objectives
  2. 1. The pipeline: the merge gates
  3. 2. The build: the digest image and the scan
  4. 3. The deployment strategy: the one the risk asks for
  5. 4. Migrations in the pipeline (11's partner)
  6. 5. The full circle: observing the deploy
  7. Self-assessment

Stack: GitHub Actions · Project: TicketFlow Status: Published — pipelines with tests, analysis and deploy-on-merge Prerequisite: Lesson 40 — Docker


Objectives

  1. Write TicketFlow's pipeline: the jobs, their order, their parallelism and what gates the merge (and what doesn't).
  2. Decide the deployment strategy (deploy-on-merge, blue-green, canary) by the release's risk.
  3. Close the loop with the pipeline's own observability: a deploy nobody knows ran deploys nothing.

1. The pipeline: the merge gates

A PR's pipeline runs the CHEAP and the DETERMINISTIC; the slow runs after the merge or on a schedule. The jobs:

yaml
# .github/workflows/ci.yml (the project's summary)
jobs:
  lint:       ruff + black --check + mypy (services/, <30 s)      # gates
  unit:       pytest tests/unit (cached, <30 s)                 # gates
  integration: pytest tests/integration (PG/Redis services on the runner, <3 min)  # gates
  e2e:        compose up + pytest -m e2e (~4 min)                 # gates with image cache
  build:      docker build + trivy + push :sha (parallel, gates on blocking CVEs)
  check-deploy: python manage.py check --deploy (22/27)           # gates
  load:       k6 smoke (1-min plateau) — nightly or manual ONLY   # does NOT gate

The merge gate: lint + unit + integration + e2e + build + check-deploy green → the merge button unlocks. load does NOT gate (a 5-min plateau per PR would be the tax killing the feedback cycle; it runs nightly and on tags). The pipeline's discipline: everything gating must be green and deterministic — a flaky test in the gate is a door that opens with a re-run (the "retry until it passes" culture, 35 forbids it).

2. The build: the digest image and the scan

The build job uses the registry's cache (cache-from: type=registry): 40's builder reuses deps layers across PRs (the 90 s build drops to 15 s). The push carries the commit SHA as tag (ticketflow:git-a1b2c3d) and the resulting digest is THE release artifact: the deploy does not rebuild — it deploys THE digest that passed the tests (the commit→image→environment chain is auditable end to end). The scan (trivy) gates in "fail on exploitable CRITICAL CVEs" mode (40's triage: the OTHERS go into the report, they don't block) — the pipeline is where the security policy gets enforced without the 3 a.m. meeting.

3. The deployment strategy: the one the risk asks for

Three options, in order of sophistication: direct redeploy (stop, deploy, boot: 30 s of downtime — fine up to ~10 internal users, not for TicketFlow), blue-green (two full environments; the proxy's switch flips blue to green: rollback = re-switch in 5 s; cost: 2× infra), canary (the new version receives 5% → 25% → 100% of traffic with 46's metrics deciding: automatic rollback if the p95 or error rate spikes; cost: routing complexity). TicketFlow picks: blue-green for the app + expand-contract migrations (11) — the exact pair: the EXPAND migration is backward compatible (the old version keeps working against the new schema), the green deploy coexists with blue during the window, and the next release's CONTRACT closes the cycle.

yaml
deploy:
  steps:
    - run: docker pull ticketflow:git-a1b2c3d
    - run: ./deploy.sh green        # green boots with the environment's config (27)
    - run: ./wait_healthy.sh green  # 40's healthcheck decides
    - run: ./switch_traffic.sh blue → green   # the proxy flips hot
    - run: ./smoke.sh && ./rollback.sh if_smoke_fails

The post-deploy smoke (35's 3 checks: listing, reservation, healthz) decides the final switch; the rollback is the sibling script: a deployment without a tested rollback isn't a deployment, it's a bet.

4. Migrations in the pipeline (11's partner)

TicketFlow's release canonical order: (1) CI green; (2) migrate EXPAND against prod's DB (the job runs from the release image: docker run ticketflow:git-a1b2c3d python manage.py migrate — migrations travel INSIDE the image, not in a magic branch); (3) blue-green deploy (the new app on the expanded schema); (4) smoke + switch; (5) the CONTRACT is left for the next release (the weekly cleanup job runs it). NEVER: migrating AFTER the deploy (the new app hits the old schema: 500s in a chain) or migrating at container boot (40). 11's ADR lives in the pipeline: CI is where decisions become code.

5. The full circle: observing the deploy

The deploy is a system EVENT (45/46): the pipeline annotates deployment_started/completed with commit, digest and actor into the monitoring — the p95 graphs carry the deploy's vertical line (47 asks "what changed at 14:20?" and the annotation answers). And the release loop: the cadence (deploy on merge or a weekly train?) is decided by the rollback's cost: blue-green with a 5 s rollback → deploy on merge is reasonable (TicketFlow); a painful rollback → train and a longer staging. The pipeline's own metrics (reduced DORA): lead time PR→prod, deployment frequency, deploy failure rate, MTTR — four numbers the team of 1 reviews monthly, printed on 46's dashboard.


Self-assessment

  1. What gates the merge on TicketFlow and what runs afterwards/on a schedule? Why doesn't load gate?
  2. The digest as artifact: which auditable chain does it build, and why does the deploy NOT rebuild?
  3. Blue-green vs canary: what does each demand from your metrics, and why does TicketFlow pick blue-green?
  4. List the release's canonical order with migrations and the two NEVERs of §4.
  5. Why is the deploy a monitoring event, and which four DORA metrics does the team review?

Continue with the exercises. The solutions only after trying it yourself.