Module 12 · Enterprise level (capstone)

Lesson 55 — Scaling

Read replicas, partitioning and stateless services.

Published
In this lesson
  1. Exercise 1 — The zero law
  2. Exercise 2 — The replicas
  3. Exercise 3 — The write bottleneck
  4. Exercise 4 — The partition
  5. Exercise 5 — The complete ladder
  6. Professor's summary

Exercise 1 — The zero law

  1. The audit: dev's locmem cache (violation in prod: replica A serving the old listing that B updated — 38 already forbade it: REDIS_URL mandatory outside dev), the visits counter in memory (per replica: the dashboard sums 3 times — tolerable if the dashboard knows, a bug if not), 54's breaker per replica (tolerable: each replica fuses on its own, the divergence bounded to 30 s), the export's partial CSVs in /tmp (23: bug — goes to 42's bucket or to 31's job in a single worker).
  2. The proof: replica A dead mid-way through the continuous checkout: the LB pulls it (43's healthcheck) in ~10 s, in-flight requests fail (the client's retry or 26's 503), the user retries and B serves them: the token (18) stays valid (no in-memory session). Chronology: 10 s of degradation, 0 of lost session, 0 of lost state. The zero law verified.
  3. The hunted: @lru_cache on tarifas_activas() (08): divergent per replica with changing data → migrated to Redis with TTL (38); the old exercise's global CONTADOR = 0: deleted (Redis INCR); /tmp/export.csv: to the bucket. All three: the honest state inventory.

Exercise 2 — The replicas

  1. The read-your-writes test:
python
@pytest.mark.django_db(transaction=True)
def test_read_your_writes(con_lag_replica):
    ref = reservar(user, event, ["A1"], clock=...).public_ref
    # without the cure (the request goes to the lagging replica, 500 ms lag):
    with lag_replica(0.5):
        assert "mis reservas" DOES_NOT_SEE(ref)         # the "just bought, doesn't show"
    # with the cure (the request's hint: the user's view forces the primary):
    with lag_replica(0.5), force_primary_for_request():
        assert "mis reservas" SEES(ref)

The cure in the router: the middleware marks the post-write request (the cookie/flag after the POST) and the router sends those reads to the primary for 10 s — the price: the primary receives the personal reads (the browse keeps going to the replica: 70% of traffic).

  1. The measurement: primary only: browse p95 290 ms, primary at 78% at plateau; with replica: browse 118 ms, primary at 41% (the margin for the peak's writing). The number justifying the rung: the primary breathes → the seat lock (the critical write) has air.

Exercise 3 — The write bottleneck

  1. The 4 reductions verified: bulk UPDATE (36), gateway outside TX (32: the checkout's TX 120 ms, not 1.2 s), minimal indexes (09's index unused in writing: retired), skip_locked. pg_stat_statements: the checkout's total write time: -47% versus 36's state.
  2. The lock-wait metric: at 36's plateau with 2 big events: wait_event='Lock' with ~14 sessions waiting and lock-wait p95 of 380 ms concentrated on the seat's UPDATE. The lock_wait_ms metric (the histogram on 46's dash) is the one deciding §4: the written threshold: "partition if lock-wait p95 > 500 ms sustained".
  3. The plan (excerpt): "replica for the browse if reads >70% ( done), partition by event if lock-wait p95 > 500 ms (exercise 4 measures it), multi-primary shard if the primary exceeds 80% sustained CPU+IO after partitioning (the REAL limit: ~200k reservations/min with this schema — the trigger is ×10 from the present)".

Exercise 4 — The partition

  1. The resource contention demonstrated: 100 buyers on the same event: p95 1.9 s (the lock's queue); 100 spread across 2 events: p95 1.7 s — the spread one does NOT free up: the seat's lock is per ROW (10: select_for_update per seat), but the checkout's mass UPDATE (36's bulk) takes locks over the event's common index: the contention still goes by event. The lesson: row-lock does not fix everything: the write pattern (the bulk by event) contends over a resource.
  2. With the partition: the spread one p95 0.9 s (the two events: independent partitions: the lock-wait spreads); the same-event 1.8 s (the contention PER event persists: 36's turnstile is the cure, not the partition). The partition separates events; it does not reduce one's internal contention.
  3. The contract: the test with CaptureQueriesContext verifies the reservations' queries carry event_id (the plan's prune: EXPLAIN shows the single partition); 31's job creates reservations_event_<uuid> at publish (the event without a partition: goes to the DEFAULT partition with a log warning (45) — the insert does not fail: the business does not stop for the infra).

Exercise 5 — The complete ladder

  1. The doc (excerpt):
markdown
# TicketFlow's ladder (measurable triggers)
1. Vertical: active by default (42/44's tier).
2. Replicas: reads > 70% of traffic → DONE (browse 118 ms).
3. Cache: the dominant listing → DONE (38, hit 98.7%).
4. Queues: the non-critical → DONE (29/31).
5. Partition by event: lock-wait p95 > 500 ms sustained → MEASURED: 380 ms. Under observation.
6. Multi-primary shard: primary > 80% sustained CPU+IO after partition → trigger at ×10.
7. Services (53): the PCI/volume of ADR-0007.
  1. The answer to "sharding just in case": "The primary is at 41% with the replica (exercise 2): the shard's trigger is ×10 from the present and costs complexity ×10 (53). The plan: partitioning at the 500 ms lock-wait (measured today: 380 ms). Sharding gets decided by ITS number, not by fear: the document has it."
  1. The final map (ASCII):
        [42-38's edge/CDN]  ETag + 23's WAF
                │
   [43's LB]───┴─[web ×N stateless replicas]──[PG replica (browse)]
                │        │                        [primary PG (writing, partition by event)]
        [38 Redis cache]│[54's breaker]                 │
                │        │                        [outbox→stream 30]
        [29 Celery workers]────[DLQ 30]           │
                │                                 │
        [beat/31 locks]                    [42's bucket: exports]

Every box of the map is a lesson of the course: the complete system's final snapshot.


Professor's summary

  • Zero law: state outside the process (the local cache and /tmp are replica bugs); the proof is the dead replica mid-flow.
  • The replica frees the browse (with read-your-writes for the "just bought"); the write bottleneck shrinks BEFORE partitioning, and the partition separates resources — not one's internal contention.
  • The ladder climbs by ITS number (lock-wait, lag, p95) with the trigger written: sharding out of fear is the anxious scaler's tax.