Stack: Django · Project: TicketFlow Status: Published — closing the observability module Prerequisite: Lesson 46 — Metrics, traces and alerts
Objectives
- Run the diagnostic protocol of a real incident: confirm → scope → locate → mitigate → explain.
- Reproduce the failure outside prod with the incident's data (anonymized, 23) and 34/37's tools.
- Write the blameless postmortem: the timeline, the root causes, and the actions that get closed (or nothing gets closed).
1. The protocol: the order that avoids the three classic mistakes
The badly-handled incident's mistakes: fixing before understanding (the blind fix that adds one more variable), the intrusive detective (the heavy debugger/profiler in prod: 37 forbids it), and the hive (5 people editing the same file). TicketFlow's protocol, in order:
1. CONFIRM: is it real? (46's dash: the symptom the alert claims?) — 2 min
2. SCOPE: who/where? (one endpoint? one tenant? since when? — 45's trace/log) — 5 min
3. MITIGATE whatever mitigation exists: 41's rollback, 27's feature flag, 23's rate limit — the user first
4. LOCATE the cause with the tools (trace → log → 37's py-spy dump) WITHOUT changing anything else
5. EXPLAIN: the postmortem (§4) and the actionsThe order matters: mitigate BEFORE locating (the user does not wait for your root cause); locate BEFORE fixing (the fix without a cause is one more bet). The on-call's line: "first stabilize, then diagnose, finally explain — and no line of code enters between 1 and 2 if it can be avoided".
2. The live-diagnosis toolbox
What the on-call has at hand in prod (without taking anything down): (1) the dash (46: the symptom and its emergence — gradual or after the 14:20 deploy? the annotation answers); (2) the trace of an affected request (45/46's trace_id: the broken request's timeline); (3) the log filtered by tid/tenant/error (45's query); (4) the py-spy dump of the slow worker (37: the dump saying where it is hung RIGHT NOW); (5) the DB queries: pg_stat_activity (who holds 10's locks?), pg_locks, pg_stat_statements (which query leads?); (6) the feature/deploy diff query: git diff a1b2c3d..d4e5f6a --stat (what changed between the last good and the first bad? — 06's bisect at the deploy level). The course's case: "checkouts from 14:20 onward take 3 s":
Dash: checkout p95 2.8 s since 14:22 — annotation: deploy a1b2c3d at 14:20 ← suspect
Trace: the new span "seat_availability_prefetch" 2.1 s (it did not exist in yesterday's trace)
Log: WARNING rate_limited × 4000 (23: the prefetch hits Redis: its own rate limit)
pg_stat_statements: the prefetch's query, top 1 by 14× more than yesterday
git diff: the "availability optimization" PR (irony: the badly-placed cache, 38!)
Mitigation: FEATURE_SEAT_PREFETCH=false (27's flag) → p95 back in 4 min, no deployThe full cycle in 20 minutes: dash→trace→log→diff→flag. The rollback was plan B (the flag was more surgical).
3. Reproduce outside prod: the incident's laboratory
The honest reproduction: (1) capture the incident's state (the slow query, the problematic payload, the concurrency pattern — with PII anonymized: 23 defines the test-dataset method); (2) reproduce in 34/36's environment: the integration test failing with the SAME signature (the lock's concurrency? the volume?); (3) the fix with its regression test (33: the test IS the bug's memory). The typical finding: the concurrency incident (the seat lock) does NOT reproduce sequentially — the reproduction needs 34's tools (two real transactions) and 36's (the plateau). If it does not reproduce: the cause is documented as a hypothesis with evidence (the postmortem marks it: "probable cause, not reproducing: the flag/monitor watches it").
4. The blameless postmortem
The postmortem (blameless) is the incident's document: what happened (the timeline in UTC), the impact (46's users/money/SLO), the root causes (contributors: the bug is THE cause; the deploy without canary, the alert without runbook and the missing flag are contributors), and the ACTIONS with owner and date. The causes' rule: write the system, not the person ("the deploy did not go through canary", not "Juan deployed without canary" — blame culture silences reports: the unreported incident repeats). The template's actions:
## Actions (each with issue, owner, date)
- [ ] FEATURE_SEAT_PREFETCH with default false in prod (done in the incident, codified)
- [ ] Regression test: prefetch under rate limit (34) — P1, this week
- [ ] The "p95 checkout +2× after deploy" alert (46): 41's automatic canary — P2
- [ ] Runbook for mass rate_limited WARNINGs — P3And the postmortem's quality test: in 6 months, does a new engineer read it and understand the incident without asking? The postmortem depending on "I was on duty that day" is not a document.
5. The mindset: the sustainable on-call
Debugging in prod is a cool-head sport: the protocol (§1) exists because panic skips steps. TicketFlow's on-call rules (team of 1: you): the supporting MOC (the provider's hotline, the community colleague: "the one who doesn't know the system but asks what panic can't see"); the change blackout (nothing deploys during an active incident except the mitigation); and recovery time (the on-call who sleeps 2 h after the incident does not review the postmortem that night — 31 schedules it). And the maturity metric: recent incidents' MTTR (41/46): the downward trend is the proof that the diagnostic system (alerts → runbooks → reproduction) improves.
Self-assessment
- List the 5-step protocol and explain why mitigate comes BEFORE locate and locate BEFORE fix.
- In §2's case: which tool produced each finding, and why did the flag beat the rollback as mitigation?
- What does the honest reproduction of a concurrency incident require, and what happens if it does not reproduce?
- What distinguishes the blameless postmortem, and why is the "bug" not the only root cause?
- Which three rules sustain a team-of-1 on-call, and which metric measures diagnostic maturity?
Continue with the exercises. The solutions only after trying it yourself.