Exercise 1 — The timeouts
- The documented collapse: 20 requests without timeout against the slow one → 20 threads/workers blocked for 60 s (gunicorn with 4 workers: the WHOLE service answers only 4 more requests in the minute), PG's pool with no free connections (the threads don't return them), the browse (which never touches the gateway) at p95 4 s for lack of workers. ONE dependency's failure topples EVERYTHING: the definition of chain collapse. With the 8 s timeout: the damage stays in the checkout (§3: the breaker limits it further).
- The table:
| Call | connect | read | Criterion |
|---|---|---|---|
| Gateway: charge | 2 s | 8 s | p99 400 ms ×20; < edge 60 s |
| Gateway: intent query | 2 s | 3 s | light operation |
| SMTP: sending | 2 s | 5 s | async (29): tolerates more |
| Redis: cache | 0.2 s | 0.5 s | if it takes >0.5 s the cache defeats its purpose (38) |
| Postgres | — | statement_timeout 10 s | 09's linter kills the infinite query |
Redis with the SHORT timeout is the cache's rule: caching must be faster than not caching — the cache taking 2 s is worse than the miss.
- The test with the real timeout: the gateway fake raises
ReadTimeout(not a simulatedGatewayTimeout) → the saga enters UNKNOWN → 32's reconciliation (GET with the intent) → CONFIRMED. The full chain: timeout → don't-know → protocol: the money's resilience is the whole flow, not the parameter.
Exercise 2 — The retries
- The flaky's evidence: attempt 1 fails, 2 fails, 3 passes:
attempt 1: fail (ConnectError) → sleep 1.0 + 0.31 (jitter) = 1.31 s
attempt 2: fail (TimeoutError) → sleep 2.0 + 0.07 = 2.07 s
attempt 3: OK- The distinction: the 402 (decline) → 1 call (no retry: the business rejection does not improve by repeating: it is 32's compensation); the ConnectError → 3 with backoff. The rule in the client:
retryable=(ConnectError, ReadTimeout)— the 4xx never.
- The herd: 50 clients without jitter when the service returns: a peak of 50 requests in <100 ms (the returning service gets hit with double force: the second fall); with jitter (±0.5 s): the peak spreads over ~500 ms (the delays' histogram shows the dispersion). The jitter is free and is the difference between "the service returns" and "the service returns and falls again".
Exercise 3 — The breaker
- The 3 states' tests:
def test_abre_tras_umbral():
for _ in range(5): # 5 consecutive failures
with pytest.raises(RetryableError): breaker.call(fallar)
with pytest.raises(BreakerOpen): # OPEN: fail-fast, no network touched
breaker.call(fallar)
assert gw.llamadas == 5 # the 6th NEVER went out
def test_half_open_sana_o_reabre():
... open for the 30 s ...
breaker.call(exito) # half-open probe: CLOSED again, failures to 0
# the failed probe: abierto_hasta += ventana (re-OPEN)- The numbers: WITHOUT breaker (fallen gateway, 50 VUs): checkout p95 8.1 s (everyone waits the timeout), 39's pool exhausted, browse p95 2.4 s (collateral), browse errors +12%. WITH breaker: the payment requests fail in ~2 ms with 503+Retry-After, browse p95 280 ms (intact), the pool breathing. The breaker LOCALIZES: payment falls, the rest never notices.
- The observable breaker:
breaker_state{service="gateway"}→ gauge 0/1/2 (closed/open/half) → the alertbreaker_open > 2 min(page: the gateway has been in quarantine for 2 min — or the breaker is badly tuned and hammers in vain). The breaker without a metric is a fuse in the basement: it blows and nobody knows.
Exercise 4 — The degradation
- The degraded mode's full flow:
Gateway down → breaker OPEN → POST /reservations/{ref}/pay → 503
{"type": ".../payment-unavailable", "Retry-After": 30,
"detail": "We're having payment trouble. Your seat stays held;
retry in a few minutes or we'll notify you."}
→ the PENDING saga holds the seat (the reservation's TTL extends: the business's policy)
→ the gateway returns → 29's retry re-enqueues the charge → reconciliation → CONFIRMED → emailThe user did NOT lose the seat: the retention with extension is the product decision (51: the peak's conversion is worth the retention) — the designed fallback, not the improvised one.
- The suggestions' degraded mode:
FEATURE_SUGGESTIONS=false(27's flag in reverse) → the front shows the manual picker (00b's same seatmap) → 35's E2E in degraded mode: the full purchase WITHOUT suggestions passes (the alternative mode's test: the "pretty" feature dies, the business lives).
- The doc (excerpt): "Degrades: cache→origin (slow-correct), suggestions→manual, email→accumulated queue (converges), gateway→503+retention+retry. NEVER degrades: seat uniqueness (10: duplicating the inventory is not degrading, it is dying), payment validity (the money), authentication (18: never 'degraded auth')". 15 lines the on-call reads before improvising.
Exercise 5 — The game day
- Game 1's chronology:
T+0 sandbox down (docker stop)
T+2s 3 payment requests waiting the timeout
T+8s 3 timeouts → 3 breaker failures
T+15s 5th failure → breaker OPEN (metric breaker_state=1)
T+16s checkout 503+Retry-After in ~2 ms (the rest of the system: p95 intact)
T+2m page alert: breaker_open > 2 min ✓
T+5m sandbox returns → half-open probe OK → CLOSED → 29's retry drains the backlog
T+7m the held sagas confirm; zero seats lost- Game 2 (latency 5 s): the gateway's timeout (8 s) does NOT cut (5 s < 8 s: the "healthy but slow" operation passes with p95 5.1 s) — BUT 42's edge (60 s) neither, and the spinner's user (~10 s) does see the pain: the finding: the OPERATION's timeout does not protect from the functional-slow (46's SLO does: the burn alert fires). The lesson: the timeout protects from the down; the SLO from the slow; the breaker from the hammering — three tools, three different pains.
- The report (excerpt): gaps: (1) the extended retention did not tell the user its limit (the 503's detail now says it); (2) the breaker's alert did not exist before the game (created); (3) the degradation runbook (exercise 4's doc) got written AFTER being needed (the P1 action). Actions with issue/owner/date. The final resilience: what survived the chaos: the browse, the pool, the inventory; what did not: the notice's UX (fixed).
Professor's summary
- Every remote call carries a per-operation timeout (p99 ×20, less than the edge): connect and read are different pains with different retries.
- Retry: only the idempotent, backoff+jitter, 3-5 attempts, the 4xx never — the breaker stops hammering and LOCALIZES the fall (the rest breathes).
- Degradation is designed (what falls first? what never falls?) and proven with controlled chaos: resilience is what survives the game, not what the code says.