Module 8 · Performance and caching

Lesson 38 — Caching strategies

What to cache, how to invalidate and what to do about stale data.

Published
In this lesson
  1. Exercise 1 — The filter
  2. Exercise 2 — Single-flight
  3. Exercise 3 — HTTP caching
  4. Exercise 4 — Invalidation by events
  5. Exercise 5 — The measurement
  6. Professor's summary

Exercise 1 — The filter

DataStale ok?Invalidatable?Decision
Published listingyes (5 min)yes (by event)cache-aside + HTTP
Event detailyes (60 s)yescache-aside + ETag
Seat availabilityNO — decides purchase—NO (12 informs, TX decides)
Org tariffsyes (15 min)yes (on edit)cache-aside
Feature config (27)yes (30 s)yes (deploy)short cache-aside
User profileyes (60 s)yes (+ GDPR purge)cache-aside + PII policy
Active reservationNO—NO (it is transactional state)
Searchyes (2 min)per query (key = params)cache-aside with canonical key
  1. The classic violation found: the listing's availability cached in the checkout view "to save the query" — the user sees A1 free, pays, 409. The cache INFORMS (the listing), the TX DECIDES (10): the fix is removing the cache from the checkout, not refining it. The second: staff permissions cached without a TTL — 21's role escalation would take until... never: without a TTL, the safety valve doesn't exist.
  1. The FLUSHALL in staging: the system loses speed (browse rebuilds in ~2-4 s per city) but nothing breaks EXCEPT the leaderboard's visit counter (12) written with INCR without persisting — a write-behind in disguise. Fix: accept the loss (non-critical counters, documented) or persist them to the outbox. The decision gets documented, not suffered.

Exercise 2 — Single-flight

  1. The disaster's evidence:
WITHOUT protection:  50 threads / cold key  →  50 listing queries  (4.1 s)
WITH single-flight:  50 threads             →  2 queries (1 rebuilder + 1 loser with no stale)  (0.3 s)

The test: ThreadPoolExecutor(50) + CaptureQueriesContext(connection) after cache.clear() — the unprotected disaster is Friday's 10:00 peak reproduced in a test.

  1. The shape migration: :v3 keeps being served while it exists (correct for the old shape), :v4 rebuilds on the first miss. Zero invalidation deploys: key versioning IS the migration — the same technique as §5's ETag and static assets' cache busting.

Exercise 3 — HTTP caching

  1. The 304 test:
python
def test_listado_304(self):
    r1 = self.client.get("/api/v1/events?city=Madrid")
    r2 = self.client.get("/api/v1/events?city=Madrid",
                         HTTP_IF_NONE_MATCH=r1["ETag"])
    self.assertEqual(r2.status_code, 304)
    self.assertEqual(r2.content, b"")
  1. The NEVER-public ones: everything requiring Authorization and being personal: Cache-Control: private, no-store (the intermediary must NOT store it; 40/43 respects it). The test on the 3 endpoints (reservations, profile, payment) is 23's contract guard: a misplaced header on a proxy exposes one user's purchase history to another.
  1. The ADR (excerpt): "CDN for public listings. Estimated gain: 92% hit rate on browse (k6/36: 70% of traffic). Invalidation: ETag by max_updated_at + max-age 60 — no active purge needed. Risk: 60 s staleness accepted by the business. Review: if the business demands <60 s, CDN API purge only on publications (a rare event)".

Exercise 4 — Invalidation by events

  1. The live invalidation test:
python
def test_event_updated_invalida(self):
    cache.set("events:Madrid:v3", [DATO_VIEJO], 300)
    publicar_evento(OutboxEvent(event_type="EventUpdated", payload={"ciudad": "Madrid"}))
    drenar_consumidores()
    self.assertEqual(cache.get("events:Madrid:v3"), None)   # or the new data if the consumer rewrites it
  1. The guarantee: "if invalidation fails, the old data lives at most TTL (300 s) and the system CONVERGES on its own" — the TTL is not an optimization: it is 32's valve. The runbook line: "stale data > TTL in support = a consumer bug, not a cache bug" (35/47's dump diagnoses it).
  1. The GDPR purge: the export (23) fires cache.delete_pattern(f"profile:{user_id}:") and the following GET rebuilds WITHOUT the deleted user. delete_pattern (SCAN in Redis, not KEYS) over a bounded namespace is cheap; over a global it is a latency suicide — the per-user key prefix makes the purge possible.

Exercise 5 — The measurement

  1. The before/after (200 plateau, 70% browse mix):
WITHOUT cache:  browse p95 412 ms   listing: 91k total queries  pg total_time 6.2e5 ms
WITH cache:     browse p95  38 ms   listing:  1.2k queries      pg total_time 8.4e3 ms
hit rate: 98.7% (the 1.3% miss = rebuilds every 5 min + single-flight)
  1. The cost: 214 MB of Redis for 50k events × per-city keys (the key holds a compressed dict). Policy: maxmemory-policy allkeys-lru — the cache is expendable (the FLUSHALL test proved it). The documented risk: if Redis evicts the lock: key during a single-flight, two rebuilds enter at once (partial stampede: 2 queries, not 50 — the LRU's cost is acceptable; 12's critical lock does NOT live in the same LRU Redis: it lives in 31's durable lock).
  1. docs/caching.md (excerpt): "DO NOT cache: availability for purchase decisions (TX, 10), active reservation (state), permissions without a TTL (21), tokens (18), the ledger (08: money). Single reason: the cost of one broken invalidation ALWAYS exceeds the latency saved."

Professor's summary

  • The two-question filter (can it be stale? can you invalidate it?) decides; profiling (37) only says what is slow.
  • Cache-aside + single-flight + stale + jitter beats the stampede; the versioned key (:v3) is the shape migration without deploys.
  • The TTL is the convergence valve (32): event-driven invalidation lives and the TTL bounds its failure; whatever fails the filter gets written into docs/caching.md forever.