Module 12 · Enterprise level (capstone)

Lesson 55 — Scaling

Read replicas, partitioning and stateless services.

Published
In this lesson
  1. Objectives
  2. 1. The zero law: stateless
  3. 2. Read replicas: the browse scales, the writing does not
  4. 3. The write bottleneck: the plan before the rung
  5. 4. The partition: when contention is over a resource
  6. 5. The complete ladder: rung, cost and the next one
  7. Self-assessment

Stack: Django/DRF · Project: TicketFlow Status: Published — replicas, partitioning and stateless services Prerequisite: Lesson 54 — Resilience


Objectives

  1. Apply the three laws of scaling: state outside (stateless), read replicas for the browse, and the real bottleneck (the write DB) with its plan.
  2. Partition what the index cannot fix: the shard by event and the partition's limit.
  3. Know the complete ladder (vertical → replicas → cache → partition → services) with each rung's cost before climbing it.

1. The zero law: stateless

The STATELESS service is the prerequisite of all horizontal scaling: any replica can serve any request — state (session, data, files) lives OUTSIDE (10's DB, 12's Redis, 42's bucket). TicketFlow's audit: (1) 18's session is already a token (no in-memory session store: replicas share no memory); (2) the export files (23) already live in the bucket; (3) dev's local-memory cache does NOT go to prod (every replica would have its divergent cache: 38); (4) the process state (54's in-memory breaker) is acceptable (per replica, tolerated divergence) — but the run lock (31) already lives in Postgres/Redis. The zero law's proof: kill one replica IN THE MIDDLE of a checkout with another replica alive: the user does not notice (the token's session continues) — the load balancer (43) distributes and the state is external.

The hidden-state trap: Python's lru_cache in a module (27's config cached per process: replicas diverge until restart — tolerable if the deploy's TTL bounds it, a bug if not), mutable globals (12's in-memory counter: per replica, NOT global), and the assumed shared file-system (replica A's /tmp is not replica B's).

2. Read replicas: the browse scales, the writing does not

Postgres with read replicas (42: managed gives them away): reads (36's browse: 70% of traffic) go to replicas, writes to the primary. Django's rule:

python
DATABASE_ROUTERS = ["core.routers.ReadWriteRouter"]

class ReadWriteRouter:
    def db_for_read(self, model, **hints):
        return "replica" if not hints.get("for_write") else "default"
    def db_for_write(self, model, **hints):
        return "default"

The problem the replica introduces (the rung's price): replication lag — the buyer creates the reservation (primary) and the "my reservations" listing (replica) does NOT see it during the ~50-500 ms of lag: the "I just bought and it doesn't show" support ticket. The cure: read-your-writes — same-user post-login operations go to the primary (the router's per-request hint: the view marks for_write or the checkout cookie forces the primary for 10 s). The assignment rule: the replica for public-anonymous (browse, catalog), the primary for personal-written and what the user JUST wrote. 36's number: the browse (290 ms p95) drops to ~120 ms with dedicated replicas and the primary breathes — but the checkout's bottleneck (the seat lock) is a WRITE: the replicas never touch it (36's lesson closing here).

3. The write bottleneck: the plan before the rung

Writing scales worse: the primary is ONE (multi-primary is §5's expensive, dangerous rung). TicketFlow's write-bottleneck plan (the seat lock): (1) reduce the work per write: 36's bulk UPDATE (measured: −40%), the minimal index on the hot table (09: every index is an extra write); (2) queue the non-critical: email/analytics are already queued (29) — the primary only does the business; (3) shorten the locks: the short TX (32: the gateway outside the TX), skip_locked (10); (4) partition (§4) if the contention is over the SAME resource; (5) vertical scaling (more CPU/RAM to the primary: the cheap rung until its limit — the next is the shard). The deciding metric: pg_stat_activity at the peak (36) + the primary's lock-wait: the plan activates by numbers, not by fear.

4. The partition: when contention is over a resource

36's case: the arena with 20k seats and 500 simultaneous buyers: the lock serializes the SAME event's buyers while ANOTHER event's wait behind (resource contention: the plan's node). The partition (by event at the gateway/logic, by range/hash in the DB):

sql
-- partition by event (list) or by date range: the course's case
CREATE TABLE reservations (...) PARTITION BY LIST (event_id);
CREATE TABLE reservations_event_01H... PARTITION OF reservations FOR VALUES IN ('01H...');
-- each partition: its index (09), its vacuum, its independent lock-wait

The benefit: event A's contention does not touch B (the lock-wait spreads per partition); the cost: queries WITHOUT the partition filter scan ALL of them (09's plan demands the WHERE event_id — the partition's contract: every query names the partition), and partitions get created/maintained (31's job creates the next event's). The rung's rule: partition when the contention is OVER A RESOURCE and the index no longer fixes it — the index first (09), the partition after: the premature partition is 51's over-engineering.

5. The complete ladder: rung, cost and the next one

The backend's ladder (bottom up, each rung with its cost):

RungWhat it buysCostTrigger
Vertical (more CPU/RAM)A bit of everythingMoney, physical limitthe first one: always
Read replicasThe browse (reads)Replication lag (§2)reads > 70% of traffic
Cache (38)Hot readsInvalidationthe listing dominates the p95
Queues (29)Non-critical out of the requestEventualitythe email in the p95
Partition (§4)Resource contentionThe queries' contractlock-wait per resource
Sharding (multi-primary)Global writingComplexity ×10the primary at its REAL limit
Services (53)Diverging teams/rhythmsThe network and the sagaConway

The anxious scaler's anti-pattern: climbing rungs without a trigger (the sharding of the 50k-sales app: 51 flags it). The rung's metric: each one activates by ITS number (the lock-wait, the lag, 46's dash's p95) and gets re-measured after climbing (36): the ladder climbs by measuring, not by guessing.


Self-assessment

  1. Which three forms of hidden state kill the zero law, and what is the proof (the dead replica mid-checkout)?
  2. What does the read replica buy, which problem does it introduce (the lag?) and what is the read-your-writes cure?
  3. List 4 reductions of the write bottleneck BEFORE partitioning/sharding. Which metric decides?
  4. When partitioning, and which contract does it impose (the plan's WHERE)? Why the index first?
  5. Walk the ladder: what does each rung buy, cost and fire on? Which is the anxious scaler's anti-pattern?

Continue with the exercises. The solutions only after trying it yourself.