Module 12 · Enterprise level (capstone)

Lesson 55 — Scaling

Read replicas, partitioning and stateless services.

Published
In this lesson
  1. Exercise 1 — The zero law
  2. Exercise 2 — The replicas
  3. Exercise 3 — The write bottleneck
  4. Exercise 4 — The partition
  5. Exercise 5 — The complete ladder
  6. Submit

The rungs, measured. No solutions.md before submitting.

Exercise 1 — The zero law

  1. The state audit: list ALL your app's state (memory, files, local cache, globals): what violates the zero law? What is tolerable (the per-replica breaker?) and what is not (the local cache in prod?)?
  2. The proof: with 2 replicas (compose scale or 2 processes), kill replica A mid-way through a continuous checkout (k6 or loop): does the user notice? does the token session (18) continue? Document the chronology.
  3. The hunted hidden state: search lru_cache, @cache, mutable globals and /tmp in your code: each finding with its verdict (tolerable with TTL? bug on replicas?).

Exercise 2 — The replicas

  1. Set up the read replica (42's managed or compose: pg_basebackup in dev) and §2's DATABASE_ROUTERS. Verify: the browse goes to the replica (pg_stat_activity per DB), the POST to the primary.
  2. The lag: write the read-your-writes test: create the reservation → query "my reservations" IMMEDIATELY → does it see it? Without the cure: does the simulated lag (pg_sleep on the replica or the fake router's delay) hide it? With the cure (the request's hint): does it appear?
  3. The measurement: k6 (36) with the browse at 70% over: primary only vs primary+replica: paste both p95s and the primary's usage. How much does the primary breathe?

Exercise 3 — The write bottleneck

  1. §3's 4 reductions on YOUR checkout: verify each (the bulk UPDATE? the gateway outside the TX? the minimal indexes? skip_locked?) and measure with pg_stat_statements (36): how much of the total write time remains?
  2. The lock-wait as metric: instrument pg_locks/pg_stat_activity (the lock's wait_event) during the plateau: how many wait and for how long? The lock_wait_ms metric on 46's dash: the one deciding partitioning.
  3. The written plan: with exercises 2-3's numbers: at which volume (how many reservations/min?) would you fire each next rung (partition? shard?)? The ladder's document with ITS measurable triggers.

Exercise 4 — The partition

  1. The course's case: reproduce the resource contention: 2 big events (20k seats each) with 100 simultaneous buyers on the SAME event and 100 spread: does the spread one suffer for the same-event (the shared lock-wait?)? Paste the numbers.
  2. Implement the partition by event (§4's LIST) and repeat the test: does the spread one free up? Does the same-event stay the same (the contention is PER event: the partition does not touch it — the same-event cure is 36's turnstile)?
  3. The partition's contract: the test demanding WHERE event_id in the reservations' queries (09's plan with the prune?) and 31's job creating the next event's partition (with the proof: the event without a partition: does the insert fail or go to default? decide and document).

Exercise 5 — The complete ladder

  1. The document docs/scaling.md: §5's ladder with YOUR numbers (each rung's measurable triggers per exercises 2-4). The document 53's CTO asks for.
  2. The hunted anti-pattern: the imaginary CTO asks for "sharding just in case": the 51/53 answer with YOUR real number (how much does the primary hold TODAY? is the trigger ×10 away from that?). Max 8 lines.
  3. The module's close: the final map of TicketFlow scaled (ASCII): replicas, cache, queues, partition, 42's edge and 54's breaker — where every piece of the course lives in the scaled system. This map is the whole course's snapshot.

Submit

Paste the state audit, the replicas' p95s, the lock-wait test with the metric and the final scaled map. Next: Lesson 56 — Compliance.