Stack: Prometheus + Grafana + OpenTelemetry · Project: TicketFlow Status: Published — SLOs and alerts by SLO, not by panic Prerequisite: Lesson 45 — Structured logging
Objectives
- Instrument TicketFlow with the RED metrics (rate, errors, duration) per endpoint and the business ones (reservations/min, sagas in UNKNOWN).
- Add distributed traces (OpenTelemetry) connecting endpoint → service → DB → queue with 45's trace_id.
- Define SLOs and alerts that wake the on-call for what matters: the error budget's burn rate, not every CPU spike.
1. The RED metrics and the business ones
RED per endpoint (what the user experiences): Rate (req/s), Errors (5xx and domain-4xx separated: 26 decided the 409 is business, not error), Duration (a latency histogram: p50/p95/p99). In Django: django-prometheus exposes the per-route counters, or your own middleware:
REQUESTS = Counter("http_requests_total", "HTTP requests", ["method", "route", "status"])
LATENCY = Histogram("http_request_duration_seconds", "Latency", ["method", "route"],
buckets=(0.05, 0.1, 0.2, 0.4, 0.8, 1.6, 3.2))
class MetricsMiddleware:
def __call__(self, request):
t0 = time.monotonic()
response = self.get_response(request)
route = request.resolver_match.route if request.resolver_match else "unmatched"
REQUESTS.labels(request.method, route, str(response.status_code)).inc()
LATENCY.labels(request.method, route).observe(time.monotonic() - t0)
return responseThe route label's nuance (the template /api/v1/reservations/{ref}, not the full path): the raw path generates a series per UUID (the cardinality explosion that takes Prometheus down). And the BUSINESS metrics — the ones on the CEO's dashboard: reservations_created_total, sagas_in_unknown (gauge: 32), outbox_pending (gauge: the delayed poller), dlq_depth (30), seat_conflicts_total (the 409s: contention is data). TicketFlow's dash: a RED row per endpoint + a business row + an infra row (39's connections, 29's queue) — "is the system healthy?" gets answered by looking at TWO numbers: the error rate and the p95.
2. The traces: a request's full stack
The distributed trace is 45's trace_id with muscle: each leg (endpoint → service → SQL → gateway → queue → worker) is a span with duration and context. OpenTelemetry with Django/psycopg/redis/celery auto-instrumentation:
# core/otel.py
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
resource = Resource.create({"service.name": "ticketflow-web", "deployment.environment": "prod"})
tracer = trace.get_tracer("ticketflow")
# the manual leg where it matters (the gateway, 32):
with tracer.start_as_current_span("gateway.charge") as span:
span.set_attribute("payment.intent_id", intent_id)
out = gateway.cobrar(intent_id, amount)
span.set_attribute("payment.status", out.status)37's checkout trace reads like the flamegraph with a timeline: POST /reservations (total 340ms) → db.tx reserva (120) → UPDATE seats (85) → outbox.insert (6) → gateway.charge (110, remote) → email.enqueue (3). The sampling: parentbased_traceidratio=0.1 (10% of traces: full volume is pure cost; 100% ONLY during the incident — 47's dynamic flag). And the log connection: the span annotates the trace_id (45): the incident query goes from log to trace with the same id — two views of the same thread.
3. The SLO: the measurable promise
The SLO turns "it should go well" into a number with a window: 99.5% of requests in < 400 ms (p95) measured over 30 days and 99.9% availability (5xx < 0.1%). The error budget is 1−SLO: the 0.5% of slow requests/month the team CAN spend. The burn rate (how fast it's being spent): current ratio / allowed ratio — burn 1 = spending exactly the budget; burn 14 over 1 h = the month's budget burns in ~2 days. TicketFlow's SLO table:
| SLI | SLO | Window |
|---|---|---|
| Read endpoints' p95 latency | < 300 ms | 30 d |
| Checkout p95 latency | < 800 ms | 30 d |
| Availability (non-5xx) | 99.9% | 30 d |
| Saga convergence (PENDING > 15 min) | < 0.1% | 30 d |
The business SLO that isn't HTTP: the saga converging ("the user ends up confirmed in <15 min" is 32's real promise).
4. The alerts: they wake you by burn, not by spike
The maturity hierarchy: the spike alert (CPU > 80%: unnecessary paging — the system can be HEALTHY at 80%) gets replaced by the burn-rate one (Google SRE, multi-window): a page alert (wakes someone) if burn > 14.4 over 1 h AND > 6 over 5 min (avoids the 2-min spike's false alarm); a ticket alert (business hours) if burn > 6 over 6 h. Symptom alerts (the user suffers it: p95, error rate, sagas UNKNOWN) vs cause alerts (CPU, disk): causes go to ticket/dash, symptoms to page. TicketFlow's minimal catalog:
| Alert | Threshold | Severity |
|---|---|---|
| checkout p95 | burn 14.4×1h | page |
| http_5xx rate | > 1% × 5 min | page |
| sagas_in_unknown > 5 | 10 min | page (money in limbo) |
| outbox_pending > 1000 | 15 min | page (events not leaving) |
| dlq_depth > 50 | 15 min | ticket (30) |
| task_failed_total rate | > 5/min × 10 min | ticket |
| expiration heartbeat > 5 min | — | page (29: the job that doesn't run) |
The heartbeat (absence) is the alert only 46 can give: the dead expiration job fires NO errors — the inventory freezes silently until the customer notices.
5. The runbook and the complete circuit
Every page alert carries a runbook: the URL of the confirming dash, the diagnostic command (45/§2's trace), the 3 most likely causes and their fix, and the link to 41's rollback. An alert without a runbook is a pager without instructions. The incident's full circuit: page alert → runbook → dash (confirms the symptom) → trace (locates the leg) → log (the detail) → fix/rollback → postmortem (47) → the new or adjusted alert. And the chain's test: the game day (47's exercise): kill Redis in staging → does the expected page fire in <5 min? Did the runbook guide the on-call? Unproven observability is theory.
Self-assessment
- RED per endpoint: which three metrics, and why does the route label never carry the raw path (cardinality)?
- What does the OTel trace add to 45's trace_id, and what gets sampled (and why not 100%)?
- Define the checkout's SLO and its error budget: what does burn rate 14.4 mean, and why does it page?
- Symptom vs cause: which alerts page and which go to ticket? Why doesn't high CPU page?
- Why is the expiration heartbeat a page, and which test validates the whole chain (game day)?
Continue with the exercises. The solutions only after trying it yourself.