Module 10 · Observability and debugging

Lesson 46 — Metrics, traces and alerts

Prometheus, Grafana, OpenTelemetry and SLO-based alerts, not panic-based ones.

Published
In this lesson
  1. Exercise 1 — The RED metrics
  2. Exercise 2 — The traces
  3. Exercise 3 — The SLOs
  4. Exercise 4 — The alerts
  5. Exercise 5 — The game day
  6. Professor's summary

Exercise 1 — The RED metrics

  1. The series after 20 requests:
http_requests_total{method="POST",route="reservations",status="201"} 18
http_requests_total{method="POST",route="reservations",status="409"} 2
http_request_duration_seconds_bucket{method="POST",route="reservations",le="0.4"} 17
http_request_duration_seconds_bucket{method="POST",route="reservations",le="0.8"} 20
  1. The explosion: raw path × 50 requests with 50 distinct UUIDs → 50 http_requests_total series + 350 histogram series (7 buckets × 50): Prometheus's index grows linearly with unique traffic (and at 2 weeks: Prometheus OOM). The fix: the label with resolver_match's template — cardinality is per ROUTE, not per request.
  1. The dash panels with 41's deploy annotation: the regression shows with the vertical line — the p95 rose EXACTLY at deploy 43 (47's postmortem confirms it in 30 s of looking).

Exercise 2 — The traces

  1. The checkout's trace (spans with durations):
POST /api/v1/reservations/pay                    340ms
├─ db.tx: begin + lock reservation               120ms
│   └─ UPDATE seats WHERE id IN                  85ms
├─ outbox.insert                                 6ms
├─ gateway.charge  [payment.intent_id=int_9f2c]  110ms   ← the remote leg, 32/54
└─ email.enqueue (on_commit)                     3ms
  1. The sampling: 50 requests → 5 traces (10%). The error detail: with the default sampler (head-based), 90% of errors leave NO trace — the cure: tail sampling at the collector (keep if status=ERROR) or the incident's dynamic flag (traceidratio=1.0 with a short window, 27). The project: ratio 0.1 + keep-errors at the collector.
  1. The thread: 45's trace_id appears as the root span's attribute and the Celery context propagation (the task re-sets the OTel context with the message's extractor): the trace connects endpoint→worker→SQL. The incident query: tid → log (the detail) + trace (the timeline): 45's two views, one id.

Exercise 3 — The SLOs

  1. The queries:
promql
# checkout p95 (5m)
histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket{route="checkout"}[5m])))

# availability (30d)
1 - (sum(rate(http_requests_total{status=~"5.."}[30d])) / sum(rate(http_requests_total[30d])))

# saga convergence: % of PENDING older than 15 min
sagas_in_unknown / clamp_min(sagas_total, 1)
  1. The budget: 99.5% → 0.5% of ~1.3M requests/month = 6,500 slow requests/month the maximum spend. Burn 3.2: the budget burns in 30/3.2 ≈ 9.4 days — with no improvement, the month ends with the SLO broken and the implicit feature freeze (an exhausted budget freezes features and prioritizes reliability: the SRE contract).
  1. The real table (post 37-39 improvements): reads p95 290 ms (< 300), checkout p95 640 ms (< 800), availability 99.96%, saga converges 0.04%. The system MEETS the SLOs — the SLO to tighten next quarter: checkout < 600 ms (the current margin is 20%).

Exercise 4 — The alerts

  1. The 3 rules:
yaml
- alert: CheckoutLatencyBudgetBurn
  expr: |
    (histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket{route="checkout"}[1h])))
      > 0.8)
    and
    (histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket{route="checkout"}[5m])))
      > 0.8)
  for: 10m
  labels: {severity: page}        # multi-window burn: 1h confirms, 5m avoids the spike

- alert: OutboxBacklog
  expr: outbox_pending > 1000
  for: 15m
  labels: {severity: page}

- alert: ExpirationJobHeartbeat
  expr: time() - max(expiration_job_last_run_seconds) > 300
  labels: {severity: page}        # the absence (29): the frozen inventory
  1. The anti-noise demonstrated: the 2-min spike (short k6) → the 5m window doesn't confirm → no page (a 2-min spike with budget intact wakes nobody); the 12-min sustained failure → both windows red → page. The policy: page = real, sustained fire; everything else, ticket.
  1. The sagas_in_unknown runbook (excerpt): "Confirm: sagas dash + 32's query (SELECT saga_id, attempts FROM purchase_saga WHERE estado='UNKNOWN_EXPIRED'). Causes: (1) gateway down → 54's breaker, wait; (2) reconciliation exhausted → 32's runbook (check the gateway panel, manually compensate/confirm with the ledger); (3) poller bug → 41's rollback. The new on-call reached the cause in 6 min with the runbook (the real test: someone else read it)".

Exercise 5 — The game day

  1. The chronology (Redis killed at T=0):
T+0s    Redis dead (cache + 29's broker)
T+8s    browse p95 rises (cold cache: 38 ms → 412 ms)
T+45s   checkout p95 crosses 800 ms (the gateway and local queue don't respond)
T+2m10s CheckoutLatencyBudgetBurn pages   ✓ (5m window confirms the burn)
T+4m30s ExpirationJobHeartbeat pages      ✓ (the worker starved without the broker)
T+6m    DLQ/task_failed to ticket         ✓

Two correct pages, clean chronology: the chain works.

  1. The gap: outbox_pending did NOT fire (the poller is alive but brokerless: events accumulate with no visible app error — the gauge climbs but the alert was DLQ-only). The added alert: outbox_pending > 1000 for 15m (§4's: the game's full list). The browse p95 with cold cache: covered by the checkout burn (shares the cause), no own alert — documented decision: one cause, one alert.
  1. The on-call test: the third party with dash+runbook diagnosed in 11 min: "Redis dead → the worker starved → restart Redis, drain the queue, verify the heartbeat". The game day's criterion: <15 min and WITHOUT contacting you — the system operates itself, which is the only definition of observability that matters.

Professor's summary

  • RED with template labels (controlled cardinality) + business metrics (sagas, outbox, DLQ): the two-number dash (error rate, p95) answers "is it healthy?".
  • OTel traces with 10% sampling + keep-errors: the request's timeline with the log's same tid.
  • SLO with budget and multi-window burn: page only the sustained symptom; the heartbeat (absence) is the alert no error fires — and the game day proves it.