Module 12 · Enterprise level (capstone)

Lesson 54 — Resilience

Timeouts, retries with backoff and circuit breakers between services.

Published
In this lesson
  1. Exercise 1 — The timeouts
  2. Exercise 2 — The retries
  3. Exercise 3 — The breaker
  4. Exercise 4 — The degradation
  5. Exercise 5 — The game day
  6. Professor's summary

Exercise 1 — The timeouts

  1. The documented collapse: 20 requests without timeout against the slow one → 20 threads/workers blocked for 60 s (gunicorn with 4 workers: the WHOLE service answers only 4 more requests in the minute), PG's pool with no free connections (the threads don't return them), the browse (which never touches the gateway) at p95 4 s for lack of workers. ONE dependency's failure topples EVERYTHING: the definition of chain collapse. With the 8 s timeout: the damage stays in the checkout (§3: the breaker limits it further).
  1. The table:
CallconnectreadCriterion
Gateway: charge2 s8 sp99 400 ms ×20; < edge 60 s
Gateway: intent query2 s3 slight operation
SMTP: sending2 s5 sasync (29): tolerates more
Redis: cache0.2 s0.5 sif it takes >0.5 s the cache defeats its purpose (38)
Postgres—statement_timeout 10 s09's linter kills the infinite query

Redis with the SHORT timeout is the cache's rule: caching must be faster than not caching — the cache taking 2 s is worse than the miss.

  1. The test with the real timeout: the gateway fake raises ReadTimeout (not a simulated GatewayTimeout) → the saga enters UNKNOWN → 32's reconciliation (GET with the intent) → CONFIRMED. The full chain: timeout → don't-know → protocol: the money's resilience is the whole flow, not the parameter.

Exercise 2 — The retries

  1. The flaky's evidence: attempt 1 fails, 2 fails, 3 passes:
attempt 1: fail (ConnectError) → sleep 1.0 + 0.31 (jitter) = 1.31 s
attempt 2: fail (TimeoutError) → sleep 2.0 + 0.07           = 2.07 s
attempt 3: OK
  1. The distinction: the 402 (decline) → 1 call (no retry: the business rejection does not improve by repeating: it is 32's compensation); the ConnectError → 3 with backoff. The rule in the client: retryable=(ConnectError, ReadTimeout) — the 4xx never.
  1. The herd: 50 clients without jitter when the service returns: a peak of 50 requests in <100 ms (the returning service gets hit with double force: the second fall); with jitter (±0.5 s): the peak spreads over ~500 ms (the delays' histogram shows the dispersion). The jitter is free and is the difference between "the service returns" and "the service returns and falls again".

Exercise 3 — The breaker

  1. The 3 states' tests:
python
def test_abre_tras_umbral():
    for _ in range(5):                       # 5 consecutive failures
        with pytest.raises(RetryableError): breaker.call(fallar)
    with pytest.raises(BreakerOpen):         # OPEN: fail-fast, no network touched
        breaker.call(fallar)
    assert gw.llamadas == 5                  # the 6th NEVER went out

def test_half_open_sana_o_reabre():
    ... open for the 30 s ...
    breaker.call(exito)   # half-open probe: CLOSED again, failures to 0
    # the failed probe: abierto_hasta += ventana (re-OPEN)
  1. The numbers: WITHOUT breaker (fallen gateway, 50 VUs): checkout p95 8.1 s (everyone waits the timeout), 39's pool exhausted, browse p95 2.4 s (collateral), browse errors +12%. WITH breaker: the payment requests fail in ~2 ms with 503+Retry-After, browse p95 280 ms (intact), the pool breathing. The breaker LOCALIZES: payment falls, the rest never notices.
  1. The observable breaker: breaker_state{service="gateway"} → gauge 0/1/2 (closed/open/half) → the alert breaker_open > 2 min (page: the gateway has been in quarantine for 2 min — or the breaker is badly tuned and hammers in vain). The breaker without a metric is a fuse in the basement: it blows and nobody knows.

Exercise 4 — The degradation

  1. The degraded mode's full flow:
Gateway down → breaker OPEN → POST /reservations/{ref}/pay → 503
  {"type": ".../payment-unavailable", "Retry-After": 30,
   "detail": "We're having payment trouble. Your seat stays held;
              retry in a few minutes or we'll notify you."}
→ the PENDING saga holds the seat (the reservation's TTL extends: the business's policy)
→ the gateway returns → 29's retry re-enqueues the charge → reconciliation → CONFIRMED → email

The user did NOT lose the seat: the retention with extension is the product decision (51: the peak's conversion is worth the retention) — the designed fallback, not the improvised one.

  1. The suggestions' degraded mode: FEATURE_SUGGESTIONS=false (27's flag in reverse) → the front shows the manual picker (00b's same seatmap) → 35's E2E in degraded mode: the full purchase WITHOUT suggestions passes (the alternative mode's test: the "pretty" feature dies, the business lives).
  1. The doc (excerpt): "Degrades: cache→origin (slow-correct), suggestions→manual, email→accumulated queue (converges), gateway→503+retention+retry. NEVER degrades: seat uniqueness (10: duplicating the inventory is not degrading, it is dying), payment validity (the money), authentication (18: never 'degraded auth')". 15 lines the on-call reads before improvising.

Exercise 5 — The game day

  1. Game 1's chronology:
T+0    sandbox down (docker stop)
T+2s   3 payment requests waiting the timeout
T+8s   3 timeouts → 3 breaker failures
T+15s  5th failure → breaker OPEN (metric breaker_state=1)
T+16s  checkout 503+Retry-After in ~2 ms (the rest of the system: p95 intact)
T+2m   page alert: breaker_open > 2 min ✓
T+5m   sandbox returns → half-open probe OK → CLOSED → 29's retry drains the backlog
T+7m   the held sagas confirm; zero seats lost
  1. Game 2 (latency 5 s): the gateway's timeout (8 s) does NOT cut (5 s < 8 s: the "healthy but slow" operation passes with p95 5.1 s) — BUT 42's edge (60 s) neither, and the spinner's user (~10 s) does see the pain: the finding: the OPERATION's timeout does not protect from the functional-slow (46's SLO does: the burn alert fires). The lesson: the timeout protects from the down; the SLO from the slow; the breaker from the hammering — three tools, three different pains.
  1. The report (excerpt): gaps: (1) the extended retention did not tell the user its limit (the 503's detail now says it); (2) the breaker's alert did not exist before the game (created); (3) the degradation runbook (exercise 4's doc) got written AFTER being needed (the P1 action). Actions with issue/owner/date. The final resilience: what survived the chaos: the browse, the pool, the inventory; what did not: the notice's UX (fixed).

Professor's summary

  • Every remote call carries a per-operation timeout (p99 ×20, less than the edge): connect and read are different pains with different retries.
  • Retry: only the idempotent, backoff+jitter, 3-5 attempts, the 4xx never — the breaker stops hammering and LOCALIZES the fall (the rest breathes).
  • Degradation is designed (what falls first? what never falls?) and proven with controlled chaos: resilience is what survives the game, not what the code says.