Stack: Django/DRF · Project: TicketFlow Status: Published — timeouts, retries and breakers between services Prerequisite: Lesson 53 — Modular monolith vs microservices
Objectives
- Armor every remote boundary of TicketFlow with per-operation timeouts (the no-timeout default is the chain collapse).
- Apply retries with backoff + jitter ONLY to what is idempotent (14/29), and the circuit breaker to stop hammering the fallen.
- Degrade with grace: the per-feature fallback (the listing without cache, the checkout without suggestions) that sells "less" instead of "nothing".
1. The timeout: the piece without which everything else is theory
The rule: every remote call carries a timeout — without it, the thread/connection waits forever, the pool exhausts (39), and your system dies by the slow third party (the chain collapse: the cascading failure where YOU are the one falling from someone else's fall). TicketFlow's timeouts, per operation (not global — the charge tolerates more than availability):
# core/http.py — the client with the project's timeouts
import httpx
class GatewayClient:
TIMEOUTS = httpx.Timeout(connect=2.0, read=8.0, write=2.0, pool=1.0) # the charge (32)
def cobrar(self, intent_id, amount):
return self._client.post("/charges", json=..., timeout=self.TIMEOUTS)The timeout's analysis: it must exceed the HEALTHY operation's p99 (the charge: p99 400 ms → read 8 s = ×20: margin for the slow 2%) and be LESS than the system's patience (42's edge: 60 s; the user's request: the front's spinner ~10 s). And the connection timeout (2 s) separate from the read one: "does not connect" (down) differs from "does not respond" (overloaded) — the retry differs (§2). The timeout's result is 32's UNKNOWN: the saga handles it (reconciliation) — the timeout is not the end, it is the entry to the don't-know protocol.
2. The retry: backoff, jitter and idempotency
The unconditional retry is the self-inflicted DDoS: hammering the fallen service finishes it off. The project's rules: (1) only the idempotent (14: the GET, the charge with Idempotency-Key, the sending with 29's dedup — the non-idempotent POST's retry DUPLICATES the effect); (2) exponential backoff + jitter (29: 1s/2s/4s ± random — the jitter prevents 50 clients from retrying at once when the service returns: the retry's thundering herd); (3) 3-5 attempts maximum (after that: DLQ/compensation — the infinite retry is 32's "wait and pray"); (4) the connect retry vs the read retry: the failed connect (never arrived) is safe to repeat; the read-timeout (arrived, we don't know) demands idempotency — the rule separating the safe from the dangerous.
def with_retry(fn, *, retries=3, base=1.0, retryable=(ConnectError, TimeoutError)):
for attempt in range(retries):
try:
return fn()
except retryable as exc:
if attempt == retries - 1:
raise
sleep(base * 2 ** attempt + random.uniform(0, 0.5)) # backoff + jitterThe fine piece: retryable — the 5xx and the timeout get retried; the 4xx (26's business rejection) does NOT (the card's decline does not improve with the attempt: it is compensation, not retry — 32).
3. The circuit breaker: stop hammering the fallen
The breaker (the fuse): after N consecutive failures, calls ABORT immediately (fail-fast, without waiting for the timeout ×N) during the open window; then half-open (ONE probe tries) and it closes if healthy:
class CircuitBreaker:
def __init__(self, umbral=5, ventana=30):
self.fallos, self.abierto_hasta = 0, 0.0
def call(self, fn, *args, **kw):
if self.abierto: # OPEN: fail-fast without touching the network
raise BreakerOpen("pasarela en cuarentena 30 s")
try:
out = fn(*args, **kw)
self.fallos = 0 # HALF-OPEN→CLOSED: healthy
return out
except RetryableError:
self.fallos += 1
if self.fallos >= self.umbral:
self.abierto_hasta = time.monotonic() + self.ventana
raiseThe measurable benefit: without the breaker, every request of the peak waits the 8 s timeout against the fallen gateway (39's pool exhausts: THE WHOLE CHECKOUT dies for the payment); with the breaker, requests fail in 1 ms with a clear 503 and the pool breathes — the breaker LOCALIZES the fall (only payment fails; browse continues: failure by compartment). And the breaker's state is a METRIC (46): breaker_state{servicio} with the alert on OPEN — the breaker that opens and nobody knows is the silent failure.
4. The graceful degradation: selling "less" instead of "nothing"
The per-feature fallback: every remote dependency carries its degraded mode defined BEFOREHAND: (1) 38's cache falls → the origin serves (slow but correct: the FLUSHALL test already proved it); (2) the gateway falls → the checkout gets queued "we'll notify you" (32's saga: the user does NOT lose the seat — the charge's UNKNOWN waits); (3) the suggestions service (52) falls → the manual seat picker keeps working (27's feature flag in reverse: turn off the nice, not the essential); (4) the email (29) falls → the queue accumulates (30's convergence). The degradation's hierarchy: first what the user does not notice (cache), then what they notice but can live with (suggestions), never the inventory nor the money (10's invariants do not degrade: the duplicated seat is not "degradation", it is dead business).
The design rule: every endpoint of 13 carries its degradation response in the contract (26's 503 with Retry-After, the checkout's manual mode documented in 47's runbook) — degradation is designed, not improvised the day of the fall.
5. Controlled chaos: the proof of resilience
Resilience without proof is a hypothesis: the game day (46/47) systematized: (1) turn off the sandbox gateway → does the breaker open in <30 s? does the pool breathe? does checkout show the degraded mode? does the saga resume when it returns? (2) add latency (the proxy with a 5 s delay) → does the timeout cut at 8 s? does 42's edge not kill first? (3) turn off Redis → 38's FLUSHALL at scale. Each exercise produces: the chronology (47), the gaps (the missing alert, the nonexistent fallback) and the fix with its test. The chaos game in an environment of 1: staging + calendar (31: the quarterly game day) — resilience is measured by what SURVIVES the chaos, not by what the code says.
Self-assessment
- Which two collapses does the timeout prevent, and how is it sized (p99 ×20? the edge?)? Why are connection and read different timeouts?
- Which are the 4 retry rules, and why does the read-timeout demand idempotency while the connect does not?
- What does the breaker do in OPEN and which concrete collapse does it prevent (39's pool?)? Why is its state a metric?
- List TicketFlow's degradation hierarchy: what DOES degrade and what NEVER does (10's invariant)?
- Which three chaos games test resilience, and what does each produce?
Continue with the exercises. The solutions only after trying it yourself.