Module 10 · Observability and debugging

Lesson 46 — Metrics, traces and alerts

Prometheus, Grafana, OpenTelemetry and SLO-based alerts, not panic-based ones.

Published
In this lesson
  1. Objectives
  2. 1. The RED metrics and the business ones
  3. 2. The traces: a request's full stack
  4. 3. The SLO: the measurable promise
  5. 4. The alerts: they wake you by burn, not by spike
  6. 5. The runbook and the complete circuit
  7. Self-assessment

Stack: Prometheus + Grafana + OpenTelemetry · Project: TicketFlow Status: Published — SLOs and alerts by SLO, not by panic Prerequisite: Lesson 45 — Structured logging


Objectives

  1. Instrument TicketFlow with the RED metrics (rate, errors, duration) per endpoint and the business ones (reservations/min, sagas in UNKNOWN).
  2. Add distributed traces (OpenTelemetry) connecting endpoint → service → DB → queue with 45's trace_id.
  3. Define SLOs and alerts that wake the on-call for what matters: the error budget's burn rate, not every CPU spike.

1. The RED metrics and the business ones

RED per endpoint (what the user experiences): Rate (req/s), Errors (5xx and domain-4xx separated: 26 decided the 409 is business, not error), Duration (a latency histogram: p50/p95/p99). In Django: django-prometheus exposes the per-route counters, or your own middleware:

python
REQUESTS = Counter("http_requests_total", "HTTP requests", ["method", "route", "status"])
LATENCY = Histogram("http_request_duration_seconds", "Latency", ["method", "route"],
                    buckets=(0.05, 0.1, 0.2, 0.4, 0.8, 1.6, 3.2))

class MetricsMiddleware:
    def __call__(self, request):
        t0 = time.monotonic()
        response = self.get_response(request)
        route = request.resolver_match.route if request.resolver_match else "unmatched"
        REQUESTS.labels(request.method, route, str(response.status_code)).inc()
        LATENCY.labels(request.method, route).observe(time.monotonic() - t0)
        return response

The route label's nuance (the template /api/v1/reservations/{ref}, not the full path): the raw path generates a series per UUID (the cardinality explosion that takes Prometheus down). And the BUSINESS metrics — the ones on the CEO's dashboard: reservations_created_total, sagas_in_unknown (gauge: 32), outbox_pending (gauge: the delayed poller), dlq_depth (30), seat_conflicts_total (the 409s: contention is data). TicketFlow's dash: a RED row per endpoint + a business row + an infra row (39's connections, 29's queue) — "is the system healthy?" gets answered by looking at TWO numbers: the error rate and the p95.

2. The traces: a request's full stack

The distributed trace is 45's trace_id with muscle: each leg (endpoint → service → SQL → gateway → queue → worker) is a span with duration and context. OpenTelemetry with Django/psycopg/redis/celery auto-instrumentation:

python
# core/otel.py
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter

resource = Resource.create({"service.name": "ticketflow-web", "deployment.environment": "prod"})
tracer = trace.get_tracer("ticketflow")

# the manual leg where it matters (the gateway, 32):
with tracer.start_as_current_span("gateway.charge") as span:
    span.set_attribute("payment.intent_id", intent_id)
    out = gateway.cobrar(intent_id, amount)
    span.set_attribute("payment.status", out.status)

37's checkout trace reads like the flamegraph with a timeline: POST /reservations (total 340ms) → db.tx reserva (120) → UPDATE seats (85) → outbox.insert (6) → gateway.charge (110, remote) → email.enqueue (3). The sampling: parentbased_traceidratio=0.1 (10% of traces: full volume is pure cost; 100% ONLY during the incident — 47's dynamic flag). And the log connection: the span annotates the trace_id (45): the incident query goes from log to trace with the same id — two views of the same thread.

3. The SLO: the measurable promise

The SLO turns "it should go well" into a number with a window: 99.5% of requests in < 400 ms (p95) measured over 30 days and 99.9% availability (5xx < 0.1%). The error budget is 1−SLO: the 0.5% of slow requests/month the team CAN spend. The burn rate (how fast it's being spent): current ratio / allowed ratio — burn 1 = spending exactly the budget; burn 14 over 1 h = the month's budget burns in ~2 days. TicketFlow's SLO table:

SLISLOWindow
Read endpoints' p95 latency< 300 ms30 d
Checkout p95 latency< 800 ms30 d
Availability (non-5xx)99.9%30 d
Saga convergence (PENDING > 15 min)< 0.1%30 d

The business SLO that isn't HTTP: the saga converging ("the user ends up confirmed in <15 min" is 32's real promise).

4. The alerts: they wake you by burn, not by spike

The maturity hierarchy: the spike alert (CPU > 80%: unnecessary paging — the system can be HEALTHY at 80%) gets replaced by the burn-rate one (Google SRE, multi-window): a page alert (wakes someone) if burn > 14.4 over 1 h AND > 6 over 5 min (avoids the 2-min spike's false alarm); a ticket alert (business hours) if burn > 6 over 6 h. Symptom alerts (the user suffers it: p95, error rate, sagas UNKNOWN) vs cause alerts (CPU, disk): causes go to ticket/dash, symptoms to page. TicketFlow's minimal catalog:

AlertThresholdSeverity
checkout p95burn 14.4×1hpage
http_5xx rate> 1% × 5 minpage
sagas_in_unknown > 510 minpage (money in limbo)
outbox_pending > 100015 minpage (events not leaving)
dlq_depth > 5015 minticket (30)
task_failed_total rate> 5/min × 10 minticket
expiration heartbeat > 5 min—page (29: the job that doesn't run)

The heartbeat (absence) is the alert only 46 can give: the dead expiration job fires NO errors — the inventory freezes silently until the customer notices.

5. The runbook and the complete circuit

Every page alert carries a runbook: the URL of the confirming dash, the diagnostic command (45/§2's trace), the 3 most likely causes and their fix, and the link to 41's rollback. An alert without a runbook is a pager without instructions. The incident's full circuit: page alert → runbook → dash (confirms the symptom) → trace (locates the leg) → log (the detail) → fix/rollback → postmortem (47) → the new or adjusted alert. And the chain's test: the game day (47's exercise): kill Redis in staging → does the expected page fire in <5 min? Did the runbook guide the on-call? Unproven observability is theory.


Self-assessment

  1. RED per endpoint: which three metrics, and why does the route label never carry the raw path (cardinality)?
  2. What does the OTel trace add to 45's trace_id, and what gets sampled (and why not 100%)?
  3. Define the checkout's SLO and its error budget: what does burn rate 14.4 mean, and why does it page?
  4. Symptom vs cause: which alerts page and which go to ticket? Why doesn't high CPU page?
  5. Why is the expiration heartbeat a page, and which test validates the whole chain (game day)?

Continue with the exercises. The solutions only after trying it yourself.