TicketFlow's observability. No solutions.md before submitting.
Exercise 1 — The RED metrics
- Implement the metrics middleware (§1) with the template
routelabel and the latency histogram. Verify the/metricsendpoint and paste one endpoint's series after 20 requests. - Cardinality: break the label on purpose (use the raw path with the UUID) and generate 50 requests: how many series does it create? Document the explosion and the fix.
- The business metrics: add
reservations_created_total(counter),sagas_in_unknown(gauge with 32's query),outbox_pending(gauge),seat_conflicts_total(counter in 26's handler). The dash: one panel per metric with the deploy's vertical line (41).
Exercise 2 — The traces
- Instrument OTel (Django auto-instrumentation + manual on
gateway.chargeas in §2) and export to the local stack (Grafana Tempo/Jaeger or the JSON log with spans). Paste the checkout's trace with its spans and durations. - The sampling: configure
traceidratio=0.1and verify (50 requests): how many traces arrived? Does the ERROR request's trace always arrive (tail/importance sampling: does the error force the keep)? - The full thread: the endpoint's trace → the Celery task's span (context propagation through the queue) → the poller's SQL span. 45's trace_id in the trace: does the incident query join log and trace with one id?
Exercise 3 — The SLOs
- Write §3's SLO table as real PromQL queries: checkout p95 (
histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket{route="checkout"}[5m])) by (le))), availability, and saga convergence. - The error budget: with SLO 99.5% (30 d), compute: how many slow requests can you spend per month? If the current burn is 3.2, in how many days does the budget burn out? Write the math.
- The honest SLI: measure your local environment (36's k6) and fill the real table: does your system MEET the proposed SLOs? Which doesn't, and which 37-39 improvement would move it?
Exercise 4 — The alerts
- Implement §4's alert rules (multi-window burn-rate PromQL for the checkout; the gauges for sagas/outbox/dlq; 29's heartbeat). Paste 3 rules' YAML (one page, one ticket, the heartbeat).
- The false positive hunted: simulate a 2-min p95 spike (a short k6) and verify the multi-window rule does NOT page (the 5-min window doesn't confirm). Then: the sustained failure → page. The difference is the anti-noise policy.
- The runbook: write the
sagas_in_unknownalert's runbook (the confirming dash, 32's affected-saga query, the 3 likely causes with their fix, and the rollback link). How long does a new on-call take to reach the cause with your runbook? (have someone else read it: the real test).
Exercise 5 — The game day
- The exercise: in staging, kill the Redis (cache and 29's broker). Which alerts fire and when? Does the expiration heartbeat page when the worker dies? Document the game day's chronology (minute by minute).
- The gap found: which failure did NOT fire any alert (outbox_pending with the poller alive but no Redis? the browse p95 rising with the cold cache?)? Add the missing alert.
- The final dashboard: assemble the one-screen dash (checkout RED, business, infra, the deploy annotations) and its URL in the runbook. The on-call test: with the dash + runbook, a third party diagnoses the game day without your help: did they make it in <15 min?
Submit
Paste the /metrics series, the checkout's trace, the 3 alert rules and the game day's chronology. Next: Lesson 47 — Debugging in production.