Module 10 · Observability and debugging

Lesson 47 — Debugging in production

Reproducing failures and writing postmortems without pointing fingers.

Published
In this lesson
  1. Exercise 1 — The cold protocol
  2. Exercise 2 — The incident
  3. Exercise 3 — The reproduction
  4. Exercise 4 — The postmortem
  5. Exercise 5 — The sustainable on-call
  6. Professor's summary

Exercise 1 — The cold protocol

  1. The checklist (oncall.md excerpt):
1. CONFIRM (2 min):  46's dash → the symptom? since when? after a deploy?
2. SCOPE (5 min):     trace (46) + log (45) → endpoint/tenant? how many affected?
3. MITIGATE:          flag (27) → rollback (41) → scale (55)  [in this order]
4. LOCATE:            trace → log → py-spy dump (37) → pg_stat_* → git diff
5. EXPLAIN:           blameless postmortem (§4) with actions
  1. The drill: 31's heavy job raises CPU to 90% for 4 min; step 1 asks "does the USER suffer it?" → endpoints' p95 flat, error rate flat → NO page (the CPU alert does not exist: 46 made it a ticket by symptom-vs-cause). The criterion: page the USER's SYMPTOM; the cause goes to ticket with a dash. The "CPU 90%!" panic dies at step 1.
  1. The incident's 5 answers: (who?) all tenants' checkouts since 14:22 (the per-tenant log confirms: not just one); (where?) the checkout endpoint (the dash separates it); (since when?) 14:22 (the deploy's annotation); (what changed?) the prefetch PR (the diff); (how much impact?) p95 2.8 s, ~180 checkouts affected, 46's budget burning at burn 9 — the quantified impact is what prioritizes mitigation.

Exercise 2 — The incident

  1. The measured mitigations: (a) flag: FEATURE_SEAT_PREFETCH=false + settings-cache reload (30 s) → base p95 in ~4 min; NOTHING is lost (the flag only turns the prefetch off). (b) rollback: 11 s switch (41) + the blue's re-deploy + the deploy's other 2 features OUT of prod (the rollback is the hammer: it reverts everything). (c) scale-out: 3 more replicas → the p95 drops to 1.9 s (disguise: the contention at Redis remains; the bill goes up forever). The decision: flag for being surgical and reversible; rollback plan B if the flag does not exist (lesson: EVERY risky feature is born with its flag, 27/41).
  1. The regression guard: 50's review with the checklist "does the optimization touch a limited resource? what happens under rate limit/contention?" would have asked "does the prefetch respect the rate limiter?" — and 36's test with the mix under the limit would have measured it. The postmortem's lesson: "optimizations" enter through the review's door with the same rigor as features (38's badly-placed cache is the most common performance antipattern).

Exercise 3 — The reproduction

  1. The anonymized dataset: 200 request tokens with synthetic UUIDs (same arrival pattern: 14:22's rate), 1 event with 20k seats, the rate limiter's exhausted token. No real email/name/IP (23: the pattern is the data; the identity is not).
  1. The red test (integration + miniature load):
python
@pytest.mark.django_db(transaction=True)
def test_prefetch_bajo_rate_limit_degrada_checkout(rate_limiter, evento, meseta_local):
    rate_limiter.agota("prefetch")          # the incident's state
    with self.assertLess(  # the p95 of 30 checkouts
        p95(lambda: checkout(client, evento)), 0.8):
        ...
# FAILED: p95 = 2.71 s (the incident's signature)
  1. The chosen fix: the ASYNC prefetch (29: the "recently seen" availability refreshes in the background, the request serves from the consistent state) + 38's cache with single-flight for the sync case. Green test (p95 0.42 s) + 36's plateau green. The regression test: test_prefetch_no_bloquea_checkout_bajo_rate_limit — the name IS the miniature postmortem.

Exercise 4 — The postmortem

  1. The document (excerpt):
markdown
# Postmortem: checkout degradation — 2027-09-28 14:22 UTC
Impact: 180 checkouts at p95 2.8 s (48 min), 46's budget: -12% of the month. No money lost.
Timeline (UTC):
  14:20  deploy a1b2c3d (PR "availability optimization": synchronous prefetch)
  14:22  checkout p95 crosses 800 ms; burn-rate page alert
  14:31  trace+log locate the seat_availability_prefetch span under rate limit
  14:36  FEATURE_SEAT_PREFETCH=false → base p95 at 14:40
Causes: root — the synchronous prefetch competes for Redis's rate limit under contention.
Contributors: (1) the PR entered as an "optimization" without a load test; (2) the feature
without a flag in the first deploy; (3) the review without a limited-resources checklist.
Actions: [§4's 4 with issue/owner/date]
  1. The blameless test: the first version said "the junior deployed without a flag" → rewritten: "the deployment does not require a flag for new features (the pipeline does not ask for it)" — the system changes (41's policy), the person is not marked. The tone changes: the document goes from judgment to engineering.
  1. The actions: P1 the regression test (protects TODAY), P2 41's automatic canary (protects the next one), P3 the rate_limited runbook (protects the on-call). The no-issue one is a wish: all 4 ended with issue, owner and date — the postmortem without closed actions is a confession without penance.

Exercise 5 — The sustainable on-call

  1. The maturity metric: incident A (§2's): MTTR 18 min (manual diagnosis, no runbook); game day B (Redis): MTTR 11 min (46's runbook); finding C (44's drift): 4 min (the plan hunts it). The downward trend EXPLAINS the investment: module 10 bought MTTR minutes with logs (45), metrics (46) and the protocol (47) — MTTR is the metric that justifies observability to the ones billing the time.
  1. The final oncall.md (2 pages): the 5-step protocol with commands → the 5 page alerts with runbook (checkout burn, sagas, outbox, heartbeat, 5xx) → the postmortem checklist with the blameless rule → the on-call-of-1 rules (MOC, change blackout during an incident, scheduled recovery). The document's test: the 3 a.m. you follows it without thinking — and the week-30 you is grateful because the system learned from every incident.

Professor's summary

  • Confirm → scope → mitigate → locate → explain: the user first, the cause after, the ego never.
  • The incident resolves with the dash→trace→log→diff→flag chain; the surgical mitigation (flag) beats the hammer (rollback) if the feature was born with a flag.
  • The blameless postmortem with closed actions is the only mechanism through which the system learns; the downward MTTR is the proof.