Module 8 · Performance and caching

Lesson 38 — Caching strategies

What to cache, how to invalidate and what to do about stale data.

Published
In this lesson
  1. Objectives
  2. 1. The two-question filter
  3. 2. The three strategies and their coupling
  4. 3. The key and the TTL: the design almost nobody does
  5. 4. The stampede: the noon disaster
  6. 5. HTTP caching and invalidation by event
  7. Self-assessment

Stack: Django/DRF + Redis · Project: TicketFlow Status: Published — what to cache, how to invalidate and what to do about staleness Prerequisite: Lesson 37 — Profiling


Objectives

  1. Decide WHAT to cache with the invalidation-cost criterion, not enthusiasm: the filter's two questions.
  2. Pick the strategy (cache-aside, write-through, HTTP caching) by data type and accepted coupling.
  3. Armor against the three classic disasters: stampede (12), broken invalidation and the cache as the store of truth.

1. The two-question filter

Profiling (37) says what is slow; the filter says what CAN be cached without buying a problem: (1) can the data be a few seconds old without breaking the business? The events listing: yes (a seat sold 2 s ago misleads nobody while browsing). The availability that decides the purchase: NO (32: the cache informs, the TX decides). (2) Do you know how to invalidate it? If the answer is "it depends on who touches what and it isn't clear", the cache gives you ACCIDENTAL hits with surprises: broken invalidation is worse than slowness (the user sees the deleted event, or the old price — wrong data is infinitely more expensive than slow data). If both answers are yes, cache it. If only the first: fix invalidation first or don't cache.

2. The three strategies and their coupling

Cache-aside (the 90% case): the app asks the cache, on a miss it reads the DB and writes the cache with a TTL. DB and cache get written at different moments: inconsistency is temporary by design.

python
def listado_publicado(ciudad: str) -> list[dict]:
    key = f"events:published:{ciudad}:v3"
    data = cache.get(key)
    if data is None:                          # miss: 1 concurrent rebuilds (see §4)
        data = [evento_to_dict(e) for e in Event.objects.publicados().filter(ciudad=ciudad)]
        cache.set(key, data, timeout=300)
    return data

Write-through: the app writes to DB and cache in the same operation (the cache never misses on hot reads); cost: cold writes nobody may read, and cache failure complicates the write. On TicketFlow only for hot, stable reads (the org's pricing config). Write-behind (cache as buffer, async flush): write gain, loss risk — NOT for money; valid for non-critical counters (analytics). And the fourth home that isn't Redis: HTTP caching (below): the CDN/client cache is the cheapest and the most forgotten.

3. The key and the TTL: the design almost nobody does

The key carries the shape's version (:v3): a dict's contract change → bump the version and the old keys expire on their own (mass invalidation without FLUSHDB, which knocks over everything else). The TTL is not a lucky number: it is the stale-while-revalidate the business accepts — listings 5 min (the business accepts 5 min of staleness), event profile 60 s, sessions/capabilities are NOT cached (21). And personal data: a cache holding PII carries 23's policy (short TTL, encryption at rest if the Redis is shared, and the right to erasure means cached PII must be purgeable per user — cache.delete_pattern(f"profile:{user_id}:*")).

The anti-cache-as-store rule: the cache is always rebuildable from the DB. If wiping the Redis (FLUSHALL) breaks the business, there is real data in the cache — write-behind without a durable flush — and that is a bug waiting for the restart. The project's test: flushall in staging → the system degrades slow but correct.

4. The stampede: the noon disaster

When the hot key's cache expires and 200 requests arrive at once: the 200 all MISS and the 200 hit the DB with the same query (12's thundering herd). Three defenses in order of preference: (1) single-flight: only ONE rebuilds, the rest wait briefly or get served the old value:

python
def listado_protected(ciudad: str) -> list[dict]:
    key = f"events:{ciudad}:v3"
    lock_key = f"lock:{key}"
    data = cache.get(key)
    if data is not None:
        return data
    if cache.add(lock_key, "1", 10):          # NX: only one rebuilder
        try:
            data = reconstruir(ciudad)
            cache.set(key, data, timeout=300)
        finally:
            cache.delete(lock_key)
        return data
    stale = cache.get(f"stale:{key}")          # the others: the old one (if any) or a short wait
    return stale if stale is not None else reconstruir(ciudad)

(2) the double stale: serve the expired value while revalidating (the stale: key with a long TTL and the logic above); (3) jitter in the TTLs (12): timeout=300 + random(0, 30) so keys don't ALL expire on the hour. The test validating the armor: simulate 50 threads over the cold key and count the DB queries in CaptureQueriesContext: expected 1-3, not 50.

5. HTTP caching and invalidation by event

The layer that pays off: Cache-Control: public, max-age=60 on the public listing + an ETag by content: the CDN and the browser cache what your Redis doesn't even touch:

python
from django.views.decorators.http import condition

def etag_listado(request, ciudad=None):
    return f'events-{ciudad or "all"}-{Event.objects.max_updated().strftime("%s")}-v3'

@condition(etag_func=etag_listado)
def listar_eventos(request): ...

The ETag with the set's max updated_at: the 304 (no body) saves 90% of the listing's bandwidth and the query computing it is pure index (09). Redis invalidation: by domain event (25/30) — EventUpdated/EventPublished fire cache.delete(f"events:{ciudad}:v3") from the consumer; live invalidation by event, with the TTL as a safety net ALWAYS (if event invalidation fails, the TTL is the valve bounding the disaster to minutes). Never: manual invalidation "when support says so".


Self-assessment

  1. The two-question filter: which TicketFlow datum fails the 1st, which the 2nd, and which passes both?
  2. Cache-aside vs write-through: what temporary inconsistency does each accept, and where does write-behind NOT fit in this project?
  3. Why does the key carry a version (:v3), and which disaster does that avoid versus FLUSHDB or manual invalidation?
  4. Reproduce the stampede: which exact sequence triggers it, and what are the three defenses with their cost?
  5. What would the ETag/Cache-Control cache, and how does it interact with invalidation by domain events (25)?

Continue with the exercises. The solutions only after trying it yourself.