Module 10 · Observability and debugging

Lesson 47 — Debugging in production

Reproducing failures and writing postmortems without pointing fingers.

Published
In this lesson
  1. Objectives
  2. 1. The protocol: the order that avoids the three classic mistakes
  3. 2. The live-diagnosis toolbox
  4. 3. Reproduce outside prod: the incident's laboratory
  5. 4. The blameless postmortem
  6. 5. The mindset: the sustainable on-call
  7. Self-assessment

Stack: Django · Project: TicketFlow Status: Published — closing the observability module Prerequisite: Lesson 46 — Metrics, traces and alerts


Objectives

  1. Run the diagnostic protocol of a real incident: confirm → scope → locate → mitigate → explain.
  2. Reproduce the failure outside prod with the incident's data (anonymized, 23) and 34/37's tools.
  3. Write the blameless postmortem: the timeline, the root causes, and the actions that get closed (or nothing gets closed).

1. The protocol: the order that avoids the three classic mistakes

The badly-handled incident's mistakes: fixing before understanding (the blind fix that adds one more variable), the intrusive detective (the heavy debugger/profiler in prod: 37 forbids it), and the hive (5 people editing the same file). TicketFlow's protocol, in order:

1. CONFIRM: is it real? (46's dash: the symptom the alert claims?) — 2 min
2. SCOPE: who/where? (one endpoint? one tenant? since when? — 45's trace/log) — 5 min
3. MITIGATE whatever mitigation exists: 41's rollback, 27's feature flag, 23's rate limit — the user first
4. LOCATE the cause with the tools (trace → log → 37's py-spy dump) WITHOUT changing anything else
5. EXPLAIN: the postmortem (§4) and the actions

The order matters: mitigate BEFORE locating (the user does not wait for your root cause); locate BEFORE fixing (the fix without a cause is one more bet). The on-call's line: "first stabilize, then diagnose, finally explain — and no line of code enters between 1 and 2 if it can be avoided".

2. The live-diagnosis toolbox

What the on-call has at hand in prod (without taking anything down): (1) the dash (46: the symptom and its emergence — gradual or after the 14:20 deploy? the annotation answers); (2) the trace of an affected request (45/46's trace_id: the broken request's timeline); (3) the log filtered by tid/tenant/error (45's query); (4) the py-spy dump of the slow worker (37: the dump saying where it is hung RIGHT NOW); (5) the DB queries: pg_stat_activity (who holds 10's locks?), pg_locks, pg_stat_statements (which query leads?); (6) the feature/deploy diff query: git diff a1b2c3d..d4e5f6a --stat (what changed between the last good and the first bad? — 06's bisect at the deploy level). The course's case: "checkouts from 14:20 onward take 3 s":

Dash: checkout p95 2.8 s since 14:22 — annotation: deploy a1b2c3d at 14:20  ← suspect
Trace: the new span "seat_availability_prefetch" 2.1 s (it did not exist in yesterday's trace)
Log: WARNING rate_limited × 4000 (23: the prefetch hits Redis: its own rate limit)
pg_stat_statements: the prefetch's query, top 1 by 14× more than yesterday
git diff: the "availability optimization" PR (irony: the badly-placed cache, 38!)
Mitigation: FEATURE_SEAT_PREFETCH=false (27's flag) → p95 back in 4 min, no deploy

The full cycle in 20 minutes: dash→trace→log→diff→flag. The rollback was plan B (the flag was more surgical).

3. Reproduce outside prod: the incident's laboratory

The honest reproduction: (1) capture the incident's state (the slow query, the problematic payload, the concurrency pattern — with PII anonymized: 23 defines the test-dataset method); (2) reproduce in 34/36's environment: the integration test failing with the SAME signature (the lock's concurrency? the volume?); (3) the fix with its regression test (33: the test IS the bug's memory). The typical finding: the concurrency incident (the seat lock) does NOT reproduce sequentially — the reproduction needs 34's tools (two real transactions) and 36's (the plateau). If it does not reproduce: the cause is documented as a hypothesis with evidence (the postmortem marks it: "probable cause, not reproducing: the flag/monitor watches it").

4. The blameless postmortem

The postmortem (blameless) is the incident's document: what happened (the timeline in UTC), the impact (46's users/money/SLO), the root causes (contributors: the bug is THE cause; the deploy without canary, the alert without runbook and the missing flag are contributors), and the ACTIONS with owner and date. The causes' rule: write the system, not the person ("the deploy did not go through canary", not "Juan deployed without canary" — blame culture silences reports: the unreported incident repeats). The template's actions:

markdown
## Actions (each with issue, owner, date)
- [ ] FEATURE_SEAT_PREFETCH with default false in prod (done in the incident, codified)
- [ ] Regression test: prefetch under rate limit (34) — P1, this week
- [ ] The "p95 checkout +2× after deploy" alert (46): 41's automatic canary — P2
- [ ] Runbook for mass rate_limited WARNINGs — P3

And the postmortem's quality test: in 6 months, does a new engineer read it and understand the incident without asking? The postmortem depending on "I was on duty that day" is not a document.

5. The mindset: the sustainable on-call

Debugging in prod is a cool-head sport: the protocol (§1) exists because panic skips steps. TicketFlow's on-call rules (team of 1: you): the supporting MOC (the provider's hotline, the community colleague: "the one who doesn't know the system but asks what panic can't see"); the change blackout (nothing deploys during an active incident except the mitigation); and recovery time (the on-call who sleeps 2 h after the incident does not review the postmortem that night — 31 schedules it). And the maturity metric: recent incidents' MTTR (41/46): the downward trend is the proof that the diagnostic system (alerts → runbooks → reproduction) improves.


Self-assessment

  1. List the 5-step protocol and explain why mitigate comes BEFORE locate and locate BEFORE fix.
  2. In §2's case: which tool produced each finding, and why did the flag beat the rollback as mitigation?
  3. What does the honest reproduction of a concurrency incident require, and what happens if it does not reproduce?
  4. What distinguishes the blameless postmortem, and why is the "bug" not the only root cause?
  5. Which three rules sustain a team-of-1 on-call, and which metric measures diagnostic maturity?

Continue with the exercises. The solutions only after trying it yourself.