Module 12 · Enterprise level (capstone)

Lesson 54 — Resilience

Timeouts, retries with backoff and circuit breakers between services.

Published
In this lesson
  1. Objectives
  2. 1. The timeout: the piece without which everything else is theory
  3. 2. The retry: backoff, jitter and idempotency
  4. 3. The circuit breaker: stop hammering the fallen
  5. 4. The graceful degradation: selling "less" instead of "nothing"
  6. 5. Controlled chaos: the proof of resilience
  7. Self-assessment

Stack: Django/DRF · Project: TicketFlow Status: Published — timeouts, retries and breakers between services Prerequisite: Lesson 53 — Modular monolith vs microservices


Objectives

  1. Armor every remote boundary of TicketFlow with per-operation timeouts (the no-timeout default is the chain collapse).
  2. Apply retries with backoff + jitter ONLY to what is idempotent (14/29), and the circuit breaker to stop hammering the fallen.
  3. Degrade with grace: the per-feature fallback (the listing without cache, the checkout without suggestions) that sells "less" instead of "nothing".

1. The timeout: the piece without which everything else is theory

The rule: every remote call carries a timeout — without it, the thread/connection waits forever, the pool exhausts (39), and your system dies by the slow third party (the chain collapse: the cascading failure where YOU are the one falling from someone else's fall). TicketFlow's timeouts, per operation (not global — the charge tolerates more than availability):

python
# core/http.py — the client with the project's timeouts
import httpx

class GatewayClient:
    TIMEOUTS = httpx.Timeout(connect=2.0, read=8.0, write=2.0, pool=1.0)   # the charge (32)

    def cobrar(self, intent_id, amount):
        return self._client.post("/charges", json=..., timeout=self.TIMEOUTS)

The timeout's analysis: it must exceed the HEALTHY operation's p99 (the charge: p99 400 ms → read 8 s = ×20: margin for the slow 2%) and be LESS than the system's patience (42's edge: 60 s; the user's request: the front's spinner ~10 s). And the connection timeout (2 s) separate from the read one: "does not connect" (down) differs from "does not respond" (overloaded) — the retry differs (§2). The timeout's result is 32's UNKNOWN: the saga handles it (reconciliation) — the timeout is not the end, it is the entry to the don't-know protocol.

2. The retry: backoff, jitter and idempotency

The unconditional retry is the self-inflicted DDoS: hammering the fallen service finishes it off. The project's rules: (1) only the idempotent (14: the GET, the charge with Idempotency-Key, the sending with 29's dedup — the non-idempotent POST's retry DUPLICATES the effect); (2) exponential backoff + jitter (29: 1s/2s/4s ± random — the jitter prevents 50 clients from retrying at once when the service returns: the retry's thundering herd); (3) 3-5 attempts maximum (after that: DLQ/compensation — the infinite retry is 32's "wait and pray"); (4) the connect retry vs the read retry: the failed connect (never arrived) is safe to repeat; the read-timeout (arrived, we don't know) demands idempotency — the rule separating the safe from the dangerous.

python
def with_retry(fn, *, retries=3, base=1.0, retryable=(ConnectError, TimeoutError)):
    for attempt in range(retries):
        try:
            return fn()
        except retryable as exc:
            if attempt == retries - 1:
                raise
            sleep(base * 2 ** attempt + random.uniform(0, 0.5))   # backoff + jitter

The fine piece: retryable — the 5xx and the timeout get retried; the 4xx (26's business rejection) does NOT (the card's decline does not improve with the attempt: it is compensation, not retry — 32).

3. The circuit breaker: stop hammering the fallen

The breaker (the fuse): after N consecutive failures, calls ABORT immediately (fail-fast, without waiting for the timeout ×N) during the open window; then half-open (ONE probe tries) and it closes if healthy:

python
class CircuitBreaker:
    def __init__(self, umbral=5, ventana=30):
        self.fallos, self.abierto_hasta = 0, 0.0

    def call(self, fn, *args, **kw):
        if self.abierto:                                   # OPEN: fail-fast without touching the network
            raise BreakerOpen("pasarela en cuarentena 30 s")
        try:
            out = fn(*args, **kw)
            self.fallos = 0                                 # HALF-OPEN→CLOSED: healthy
            return out
        except RetryableError:
            self.fallos += 1
            if self.fallos >= self.umbral:
                self.abierto_hasta = time.monotonic() + self.ventana
            raise

The measurable benefit: without the breaker, every request of the peak waits the 8 s timeout against the fallen gateway (39's pool exhausts: THE WHOLE CHECKOUT dies for the payment); with the breaker, requests fail in 1 ms with a clear 503 and the pool breathes — the breaker LOCALIZES the fall (only payment fails; browse continues: failure by compartment). And the breaker's state is a METRIC (46): breaker_state{servicio} with the alert on OPEN — the breaker that opens and nobody knows is the silent failure.

4. The graceful degradation: selling "less" instead of "nothing"

The per-feature fallback: every remote dependency carries its degraded mode defined BEFOREHAND: (1) 38's cache falls → the origin serves (slow but correct: the FLUSHALL test already proved it); (2) the gateway falls → the checkout gets queued "we'll notify you" (32's saga: the user does NOT lose the seat — the charge's UNKNOWN waits); (3) the suggestions service (52) falls → the manual seat picker keeps working (27's feature flag in reverse: turn off the nice, not the essential); (4) the email (29) falls → the queue accumulates (30's convergence). The degradation's hierarchy: first what the user does not notice (cache), then what they notice but can live with (suggestions), never the inventory nor the money (10's invariants do not degrade: the duplicated seat is not "degradation", it is dead business).

The design rule: every endpoint of 13 carries its degradation response in the contract (26's 503 with Retry-After, the checkout's manual mode documented in 47's runbook) — degradation is designed, not improvised the day of the fall.

5. Controlled chaos: the proof of resilience

Resilience without proof is a hypothesis: the game day (46/47) systematized: (1) turn off the sandbox gateway → does the breaker open in <30 s? does the pool breathe? does checkout show the degraded mode? does the saga resume when it returns? (2) add latency (the proxy with a 5 s delay) → does the timeout cut at 8 s? does 42's edge not kill first? (3) turn off Redis → 38's FLUSHALL at scale. Each exercise produces: the chronology (47), the gaps (the missing alert, the nonexistent fallback) and the fix with its test. The chaos game in an environment of 1: staging + calendar (31: the quarterly game day) — resilience is measured by what SURVIVES the chaos, not by what the code says.


Self-assessment

  1. Which two collapses does the timeout prevent, and how is it sized (p99 ×20? the edge?)? Why are connection and read different timeouts?
  2. Which are the 4 retry rules, and why does the read-timeout demand idempotency while the connect does not?
  3. What does the breaker do in OPEN and which concrete collapse does it prevent (39's pool?)? Why is its state a metric?
  4. List TicketFlow's degradation hierarchy: what DOES degrade and what NEVER does (10's invariant)?
  5. Which three chaos games test resilience, and what does each produce?

Continue with the exercises. The solutions only after trying it yourself.