TicketFlow's fuses, tested. No solutions.md before submitting.
Exercise 1 — The timeouts
- Invent the no-timeout: an HTTP client without a timeout against an endpoint taking 60 s (simulated): fire 20 concurrent requests and watch 39's pool (how many connections wait? the p95 of the REST of the system?). Document the chain collapse.
- The per-operation timeouts: the table of EVERY remote call of TicketFlow (gateway, SMTP, Redis, the DB) with its connect/read/write and §1's criterion (p99 ×20, less than the edge). Paste the table.
- The directed UNKNOWN: the charge's timeout fires 32's protocol (reconciliation): write the test: charge with read-timeout → saga UNKNOWN → reconciliation → CONFIRMED (32's exercise 3 with the REAL timeout, not a simulated one).
Exercise 2 — The retries
- Implement
with_retry(§2) with backoff+jitter and apply it to the gateway's availability GET and the email sending (29's dedup already makes it idempotent). Tests: the flaky that fails twice and passes on the third → how many attempts and which delays (is the jitter visible in the logs?). - The 4xx does not retry: the test: the card's decline (26's 402) does NOT fire a retry (the mock counts the calls: 1, not 3) — and the connect-error DOES (3). The distinction documented in the client.
- The retry's thundering herd: 50 clients with backoff WITHOUT jitter when the service returns: which peak does it generate? Repeat with jitter: does the peak flatten? (36's test in miniature or the delays' histogram).
Exercise 3 — The breaker
- Implement the
CircuitBreaker(§3) with CLOSED/OPEN/HALF-OPEN states and the tests of the 3 states: 5 failures → OPEN; after 30 s the probe (half-open) succeeds → CLOSED; the probe fails → back to OPEN. - The measured benefit: with the simulated fallen gateway, 50 VUs of 36 against the checkout WITHOUT breaker vs WITH breaker: paste both numbers (does the browse's p95 survive in the second? does 39's pool breathe?). The breaker that localizes the fall.
- The observable breaker: the metric
breaker_state{servicio}(46) with the alert on OPEN > 2 min. Test: the breaker opens → the metric says it → the page alert (46's exercise applied to the breaker).
Exercise 4 — The degradation
- The checkout's degraded mode: the gateway down → the endpoint returns 503+Retry-After (26) and the saga holds the seat (32: "we'll notify you") — the full test: fall → 503 → the gateway returns → the saga resumes → CONFIRMED. Did the user lose the seat? (the answer is NO and why).
- The degradable feature: 52's suggestions down → the flag in reverse (turn off the nice): the manual picker continues. The degraded mode's test: the suggestions endpoint 503 → the checkout works WITHOUT suggestions (what does 35's E2E prove in this mode?).
- The documented hierarchy: the endpoint→degradation→what NEVER degrades table (10's invariant: the duplicated seat is not degradation) in
docs/reference/degradacion.md(48). Max 15 lines: the fall's runbook.
Exercise 5 — The resilience game day
- Game 1 (gateway down): turn off the sandbox, time it: does the breaker open? how fast? does the checkout degrade? does the saga resume when it returns? The minute-by-minute chronology (47's format).
- Game 2 (the malignant latency): the proxy with a 5 s delay on the gateway: does the timeout cut? where (42's edge before yours? — 42's finding applied)? Adjust the timeouts if the game contradicts them.
- The game day's report: the 2 games with chronology, gaps (the alert? the fallback? the runbook?) and the actions (issue/owner/date). The resilience test: what SURVIVED the chaos and what did NOT.
Submit
Paste the timeouts' table, the breaker's tests with the metrics, the checkout's degraded mode and the game day's report. Next: Lesson 55 — Scaling.