The simulated incident, from the alert to the postmortem. No solutions.md before submitting.
Exercise 1 — The cold protocol
- Write the 5-step protocol (§1) as the on-call's printable checklist (the
docs/oncall.mdfile): with each step's commands (dash, trace, log, diff, rollback/flag). - The confirmation drill: generate a false symptom (a CPU spike with no user impact: a heavy 31 job). With your checklist: does step 1 (is it real?) save you from the unnecessary page? Document the criterion that distinguishes a real symptom from noise.
- The scoping: for §2's real incident, write the 5 questions of step 2 (who? where? since when? what changed? how much impact?) with each answer using 45/46's tools.
Exercise 2 — The course's incident
- Reproduce §2's incident in staging: the prefetch under rate limit, the p95 at 2.8 s, and the full chronology (dash with deploy annotation → trace with the new span → log with the WARNINGs → git diff → flag).
- Surgical mitigation vs the hammer: time each mitigation: (a) the flag (how long to apply and return to base p95?); (b) the full rollback (41: how long? what does the rollback lose — the deploy's other features?); (c) scale-out (55: does it help this incident or disguise it?). Document the decision.
- The irony's finding: the PR was an "optimization". Write the guard preventing the regression: which 50 review would have caught the prefetch under rate limit? Which 36 test would have measured it?
Exercise 3 — The reproduction
- Capture the incident's state: the prefetch's query, the rate-limit pattern, and the anonymized dataset (23: the identifiers and the pattern, no PII). Paste the reproduction dataset.
- The reproduction test: with 34's tools (integration) or 36's (load): the test that FAILS with the incident's signature. Paste it red.
- The fix with its regression: apply the correct fix (38's cache with TTL and single-flight? the async prefetch? you decide) and verify: green test + 36's plateau green. The regression test enters the suite with the name documenting it.
Exercise 4 — The postmortem
- Write the incident's full postmortem (§4) with the template: UTC timeline, impact (how many users, how much of the budget's SLO?), root causes + contributors, actions with owner and date.
- The blameless test: review your postmortem: does any human name/role appear as the cause? Rewrite it in terms of the system. What changes in the document's tone?
- The actions that get closed: mark §4's 4 actions as real issues (or their equivalent). The rule: an action without an issue with owner and date is a wish. Which of the 4 is P1 and why?
Exercise 5 — The sustainable on-call
- The complete 46 game day (Redis turned off) taken to postmortem: write the document with the exercise's real chronology, and the actions that emerged (the missing outbox alert? the heartbeat runbook).
- The maturity metric: list the course's incidents (§2's, the game day, 42-46's findings) with their estimated MTTR: the trend? Which of this module's investments (45/46) explains it?
- The final
docs/oncall.md: the protocol, the 5 page alerts with their runbook, the postmortem checklist, and the on-call-of-1 rules (MOC, blackout, recovery). Maximum 2 pages: the manual the 3 a.m. you follows without thinking.
Submit
Paste the oncall.md, the incident's chronology with the chosen mitigation, the reproduction test and the blameless postmortem. Next: Lesson 48 — Documentation and ADRs (module 11).