Exercise 1 — The RED metrics
- The series after 20 requests:
http_requests_total{method="POST",route="reservations",status="201"} 18
http_requests_total{method="POST",route="reservations",status="409"} 2
http_request_duration_seconds_bucket{method="POST",route="reservations",le="0.4"} 17
http_request_duration_seconds_bucket{method="POST",route="reservations",le="0.8"} 20- The explosion: raw path × 50 requests with 50 distinct UUIDs → 50
http_requests_totalseries + 350 histogram series (7 buckets × 50): Prometheus's index grows linearly with unique traffic (and at 2 weeks: Prometheus OOM). The fix: the label with resolver_match's template — cardinality is per ROUTE, not per request.
- The dash panels with 41's deploy annotation: the regression shows with the vertical line — the p95 rose EXACTLY at deploy 43 (47's postmortem confirms it in 30 s of looking).
Exercise 2 — The traces
- The checkout's trace (spans with durations):
POST /api/v1/reservations/pay 340ms
├─ db.tx: begin + lock reservation 120ms
│ └─ UPDATE seats WHERE id IN 85ms
├─ outbox.insert 6ms
├─ gateway.charge [payment.intent_id=int_9f2c] 110ms ← the remote leg, 32/54
└─ email.enqueue (on_commit) 3ms- The sampling: 50 requests → 5 traces (10%). The error detail: with the default sampler (head-based), 90% of errors leave NO trace — the cure: tail sampling at the collector (keep if
status=ERROR) or the incident's dynamic flag (traceidratio=1.0with a short window, 27). The project: ratio 0.1 + keep-errors at the collector.
- The thread: 45's trace_id appears as the root span's attribute and the Celery context propagation (the task re-sets the OTel context with the message's extractor): the trace connects endpoint→worker→SQL. The incident query: tid → log (the detail) + trace (the timeline): 45's two views, one id.
Exercise 3 — The SLOs
- The queries:
# checkout p95 (5m)
histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket{route="checkout"}[5m])))
# availability (30d)
1 - (sum(rate(http_requests_total{status=~"5.."}[30d])) / sum(rate(http_requests_total[30d])))
# saga convergence: % of PENDING older than 15 min
sagas_in_unknown / clamp_min(sagas_total, 1)- The budget: 99.5% → 0.5% of ~1.3M requests/month = 6,500 slow requests/month the maximum spend. Burn 3.2: the budget burns in 30/3.2 ≈ 9.4 days — with no improvement, the month ends with the SLO broken and the implicit feature freeze (an exhausted budget freezes features and prioritizes reliability: the SRE contract).
- The real table (post 37-39 improvements): reads p95 290 ms (< 300), checkout p95 640 ms (< 800), availability 99.96%, saga converges 0.04%. The system MEETS the SLOs — the SLO to tighten next quarter: checkout < 600 ms (the current margin is 20%).
Exercise 4 — The alerts
- The 3 rules:
- alert: CheckoutLatencyBudgetBurn
expr: |
(histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket{route="checkout"}[1h])))
> 0.8)
and
(histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket{route="checkout"}[5m])))
> 0.8)
for: 10m
labels: {severity: page} # multi-window burn: 1h confirms, 5m avoids the spike
- alert: OutboxBacklog
expr: outbox_pending > 1000
for: 15m
labels: {severity: page}
- alert: ExpirationJobHeartbeat
expr: time() - max(expiration_job_last_run_seconds) > 300
labels: {severity: page} # the absence (29): the frozen inventory- The anti-noise demonstrated: the 2-min spike (short k6) → the 5m window doesn't confirm → no page (a 2-min spike with budget intact wakes nobody); the 12-min sustained failure → both windows red → page. The policy: page = real, sustained fire; everything else, ticket.
- The sagas_in_unknown runbook (excerpt): "Confirm: sagas dash + 32's query (
SELECT saga_id, attempts FROM purchase_saga WHERE estado='UNKNOWN_EXPIRED'). Causes: (1) gateway down → 54's breaker, wait; (2) reconciliation exhausted → 32's runbook (check the gateway panel, manually compensate/confirm with the ledger); (3) poller bug → 41's rollback. The new on-call reached the cause in 6 min with the runbook (the real test: someone else read it)".
Exercise 5 — The game day
- The chronology (Redis killed at T=0):
T+0s Redis dead (cache + 29's broker)
T+8s browse p95 rises (cold cache: 38 ms → 412 ms)
T+45s checkout p95 crosses 800 ms (the gateway and local queue don't respond)
T+2m10s CheckoutLatencyBudgetBurn pages ✓ (5m window confirms the burn)
T+4m30s ExpirationJobHeartbeat pages ✓ (the worker starved without the broker)
T+6m DLQ/task_failed to ticket ✓Two correct pages, clean chronology: the chain works.
- The gap:
outbox_pendingdid NOT fire (the poller is alive but brokerless: events accumulate with no visible app error — the gauge climbs but the alert was DLQ-only). The added alert:outbox_pending > 1000 for 15m(§4's: the game's full list). The browse p95 with cold cache: covered by the checkout burn (shares the cause), no own alert — documented decision: one cause, one alert.
- The on-call test: the third party with dash+runbook diagnosed in 11 min: "Redis dead → the worker starved → restart Redis, drain the queue, verify the heartbeat". The game day's criterion: <15 min and WITHOUT contacting you — the system operates itself, which is the only definition of observability that matters.
Professor's summary
- RED with template labels (controlled cardinality) + business metrics (sagas, outbox, DLQ): the two-number dash (error rate, p95) answers "is it healthy?".
- OTel traces with 10% sampling + keep-errors: the request's timeline with the log's same tid.
- SLO with budget and multi-window burn: page only the sustained symptom; the heartbeat (absence) is the alert no error fires — and the game day proves it.