Exercise 1 — The filter
| Data | Stale ok? | Invalidatable? | Decision |
|---|---|---|---|
| Published listing | yes (5 min) | yes (by event) | cache-aside + HTTP |
| Event detail | yes (60 s) | yes | cache-aside + ETag |
| Seat availability | NO — decides purchase | — | NO (12 informs, TX decides) |
| Org tariffs | yes (15 min) | yes (on edit) | cache-aside |
| Feature config (27) | yes (30 s) | yes (deploy) | short cache-aside |
| User profile | yes (60 s) | yes (+ GDPR purge) | cache-aside + PII policy |
| Active reservation | NO | — | NO (it is transactional state) |
| Search | yes (2 min) | per query (key = params) | cache-aside with canonical key |
- The classic violation found: the listing's availability cached in the checkout view "to save the query" — the user sees A1 free, pays, 409. The cache INFORMS (the listing), the TX DECIDES (10): the fix is removing the cache from the checkout, not refining it. The second: staff permissions cached without a TTL — 21's role escalation would take until... never: without a TTL, the safety valve doesn't exist.
- The FLUSHALL in staging: the system loses speed (browse rebuilds in ~2-4 s per city) but nothing breaks EXCEPT the leaderboard's visit counter (12) written with INCR without persisting — a write-behind in disguise. Fix: accept the loss (non-critical counters, documented) or persist them to the outbox. The decision gets documented, not suffered.
Exercise 2 — Single-flight
- The disaster's evidence:
WITHOUT protection: 50 threads / cold key → 50 listing queries (4.1 s)
WITH single-flight: 50 threads → 2 queries (1 rebuilder + 1 loser with no stale) (0.3 s)The test: ThreadPoolExecutor(50) + CaptureQueriesContext(connection) after cache.clear() — the unprotected disaster is Friday's 10:00 peak reproduced in a test.
- The shape migration:
:v3keeps being served while it exists (correct for the old shape),:v4rebuilds on the first miss. Zero invalidation deploys: key versioning IS the migration — the same technique as §5's ETag and static assets' cache busting.
Exercise 3 — HTTP caching
- The 304 test:
def test_listado_304(self):
r1 = self.client.get("/api/v1/events?city=Madrid")
r2 = self.client.get("/api/v1/events?city=Madrid",
HTTP_IF_NONE_MATCH=r1["ETag"])
self.assertEqual(r2.status_code, 304)
self.assertEqual(r2.content, b"")- The NEVER-public ones: everything requiring Authorization and being personal:
Cache-Control: private, no-store(the intermediary must NOT store it; 40/43 respects it). The test on the 3 endpoints (reservations, profile, payment) is 23's contract guard: a misplaced header on a proxy exposes one user's purchase history to another.
- The ADR (excerpt): "CDN for public listings. Estimated gain: 92% hit rate on browse (k6/36: 70% of traffic). Invalidation: ETag by max_updated_at + max-age 60 — no active purge needed. Risk: 60 s staleness accepted by the business. Review: if the business demands <60 s, CDN API purge only on publications (a rare event)".
Exercise 4 — Invalidation by events
- The live invalidation test:
def test_event_updated_invalida(self):
cache.set("events:Madrid:v3", [DATO_VIEJO], 300)
publicar_evento(OutboxEvent(event_type="EventUpdated", payload={"ciudad": "Madrid"}))
drenar_consumidores()
self.assertEqual(cache.get("events:Madrid:v3"), None) # or the new data if the consumer rewrites it- The guarantee: "if invalidation fails, the old data lives at most TTL (300 s) and the system CONVERGES on its own" — the TTL is not an optimization: it is 32's valve. The runbook line: "stale data > TTL in support = a consumer bug, not a cache bug" (35/47's dump diagnoses it).
- The GDPR purge: the export (23) fires
cache.delete_pattern(f"profile:{user_id}:")and the following GET rebuilds WITHOUT the deleted user.delete_pattern(SCAN in Redis, not KEYS) over a bounded namespace is cheap; over a globalit is a latency suicide — the per-user key prefix makes the purge possible.
Exercise 5 — The measurement
- The before/after (200 plateau, 70% browse mix):
WITHOUT cache: browse p95 412 ms listing: 91k total queries pg total_time 6.2e5 ms
WITH cache: browse p95 38 ms listing: 1.2k queries pg total_time 8.4e3 ms
hit rate: 98.7% (the 1.3% miss = rebuilds every 5 min + single-flight)- The cost: 214 MB of Redis for 50k events × per-city keys (the key holds a compressed dict). Policy:
maxmemory-policy allkeys-lru— the cache is expendable (the FLUSHALL test proved it). The documented risk: if Redis evicts thelock:key during a single-flight, two rebuilds enter at once (partial stampede: 2 queries, not 50 — the LRU's cost is acceptable; 12's critical lock does NOT live in the same LRU Redis: it lives in 31's durable lock).
docs/caching.md(excerpt): "DO NOT cache: availability for purchase decisions (TX, 10), active reservation (state), permissions without a TTL (21), tokens (18), the ledger (08: money). Single reason: the cost of one broken invalidation ALWAYS exceeds the latency saved."
Professor's summary
- The two-question filter (can it be stale? can you invalidate it?) decides; profiling (37) only says what is slow.
- Cache-aside + single-flight + stale + jitter beats the stampede; the versioned key (
:v3) is the shape migration without deploys. - The TTL is the convergence valve (32): event-driven invalidation lives and the TTL bounds its failure; whatever fails the filter gets written into docs/caching.md forever.