Module 10 · Observability and debugging

Lesson 46 — Metrics, traces and alerts

Prometheus, Grafana, OpenTelemetry and SLO-based alerts, not panic-based ones.

Published
In this lesson
  1. Exercise 1 — The RED metrics
  2. Exercise 2 — The traces
  3. Exercise 3 — The SLOs
  4. Exercise 4 — The alerts
  5. Exercise 5 — The game day
  6. Submit

TicketFlow's observability. No solutions.md before submitting.

Exercise 1 — The RED metrics

  1. Implement the metrics middleware (§1) with the template route label and the latency histogram. Verify the /metrics endpoint and paste one endpoint's series after 20 requests.
  2. Cardinality: break the label on purpose (use the raw path with the UUID) and generate 50 requests: how many series does it create? Document the explosion and the fix.
  3. The business metrics: add reservations_created_total (counter), sagas_in_unknown (gauge with 32's query), outbox_pending (gauge), seat_conflicts_total (counter in 26's handler). The dash: one panel per metric with the deploy's vertical line (41).

Exercise 2 — The traces

  1. Instrument OTel (Django auto-instrumentation + manual on gateway.charge as in §2) and export to the local stack (Grafana Tempo/Jaeger or the JSON log with spans). Paste the checkout's trace with its spans and durations.
  2. The sampling: configure traceidratio=0.1 and verify (50 requests): how many traces arrived? Does the ERROR request's trace always arrive (tail/importance sampling: does the error force the keep)?
  3. The full thread: the endpoint's trace → the Celery task's span (context propagation through the queue) → the poller's SQL span. 45's trace_id in the trace: does the incident query join log and trace with one id?

Exercise 3 — The SLOs

  1. Write §3's SLO table as real PromQL queries: checkout p95 (histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket{route="checkout"}[5m])) by (le))), availability, and saga convergence.
  2. The error budget: with SLO 99.5% (30 d), compute: how many slow requests can you spend per month? If the current burn is 3.2, in how many days does the budget burn out? Write the math.
  3. The honest SLI: measure your local environment (36's k6) and fill the real table: does your system MEET the proposed SLOs? Which doesn't, and which 37-39 improvement would move it?

Exercise 4 — The alerts

  1. Implement §4's alert rules (multi-window burn-rate PromQL for the checkout; the gauges for sagas/outbox/dlq; 29's heartbeat). Paste 3 rules' YAML (one page, one ticket, the heartbeat).
  2. The false positive hunted: simulate a 2-min p95 spike (a short k6) and verify the multi-window rule does NOT page (the 5-min window doesn't confirm). Then: the sustained failure → page. The difference is the anti-noise policy.
  3. The runbook: write the sagas_in_unknown alert's runbook (the confirming dash, 32's affected-saga query, the 3 likely causes with their fix, and the rollback link). How long does a new on-call take to reach the cause with your runbook? (have someone else read it: the real test).

Exercise 5 — The game day

  1. The exercise: in staging, kill the Redis (cache and 29's broker). Which alerts fire and when? Does the expiration heartbeat page when the worker dies? Document the game day's chronology (minute by minute).
  2. The gap found: which failure did NOT fire any alert (outbox_pending with the poller alive but no Redis? the browse p95 rising with the cold cache?)? Add the missing alert.
  3. The final dashboard: assemble the one-screen dash (checkout RED, business, infra, the deploy annotations) and its URL in the runbook. The on-call test: with the dash + runbook, a third party diagnoses the game day without your help: did they make it in <15 min?

Submit

Paste the /metrics series, the checkout's trace, the 3 alert rules and the game day's chronology. Next: Lesson 47 — Debugging in production.