Exercise 1 — The script
- The two p95s (the course environment's reference figures): runserver ~850 ms (a single synchronous process, no pooling, debug-ish mode: measuring the dev server), gunicorn 4 workers ~210 ms. The first measurement wasn't of the system: it was of the development toy. runserver exists to iterate on HTML, not to sustain traffic — a load test against it is data about nothing.
- The correctness check:
check(detalle, { "detail schema": (r) => r.status === 200
&& typeof r.json().seats_available === "number" });Without correctness, an endpoint answering 500 quickly "wins" the load test: thresholds mix speed and correctness, and checks>0.999 forces the system to answer WELL under pressure.
Exercise 2 — The checkout
1-2. The plateau (200 VUs, 5 min, gunicorn+PG):
http_req_duration: p95=380ms p99=720ms ✓ thresholds
http_req_failed: 0.02% ✓
rate_409: 1.8% of reservation attempts ← normal contention (two clickers, same seat)
checks: 99.98% ✓The 409s (3.6 of every 200 attempts) are the system deciding cleanly; k6's check accepts them and the Counter records them — the 409 rate is CONTENTION data (if it spikes to 30%, the mix is unrealistic or there is an availability bug: seats shown as free are no longer free, 12/32).
- Arena vs small event: reserve p95 in arenas 2.1 s vs 290 ms in small ones. The cause: the UPDATE of the same event's seats contends (10's lock serializes the same event's rows) and the planner (09) scans 20k seat rows. Conclusion: capacity is NOT one number — it is per event. Partitioning by event (55) and the
(event_id, status)index are the answer; the global "N concurrent users" is the sum of per-event capacities with the mix.
Exercise 3 — The bottleneck
- pg_stat_statements' top on the plateau:
total_time calls query
1.8e6 ms 4200 UPDATE seats SET status=$1 WHERE id IN (...) AND status='FREE'
4.2e5 ms 91000 SELECT ... FROM events WHERE uuid=$1
...The UPDATE explains 74% of DB time: it is the bottleneck. The listing's queries (cached, 12) don't even appear: the intuition "the listing is the slow part" was internal marketing.
- The pool:
pg_stat_activityshows 38idle in transactionconnections at the peak (the "open transaction, call the gateway, commit" pattern of 10/32 holds the lock AND the connection during 1.2 s of remote latency) — the checkout's p95 coincides with the connection-wait moments. The finding: separate the local TX from the remote (the saga already does it: charge outside the TX, 32) and tune CONN_MAX_AGE/pool (39). Two findings for the price of one SELECT.
- The 3 plateaus with gateway delay: 0/200/500 ms → checkout p99 720/1100/1600 ms. The increase is ~1:1 with the delay: the gateway adds its latency WITHOUT saturating your CPU. Honest conclusion: you can commit to the p95 of YOUR part (DB+app), not of the whole chain — that is why the SLO is written per component and 54's breaker protects the user from the slow third party.
Exercise 4 — The curve
- The 400 peak: graceful degradation if: p95 rises to 610 ms (not to 5 s), rate_409 rises to 4.2%, real errors <0.3%, CPU <85%. Collapse would be: chained timeouts (the exhausted pool blocks everything), container OOM, growing 500s. 55's design threshold: the latency curve's slope — graceful degradation has a gentle slope; the hard bottleneck (pool/lock) has a vertical one.
- The ramp-down: baseline returned in 40 s, but the worker's log showed 1.2k Celery tasks accumulated at the peak (the peak's reservations' emails/webhooks) draining for 6 more minutes. The finding: the queue ABSORBS the peak (good) but the expiration job (31) running during the drain contends with new reservations (skip_locked (10) resolved it without corruption). The queue's and the worker's capacity are PART of the capacity number: "supports 200" includes emails going out <5 min post-peak.
- The capacity document (
load/results/2027-09-28.md):
# TicketFlow capacity — 2027-09-28 (commit a1b2c3d)
Hardware: staging 2 vCPU/4GB app, PG 16 2 vCPU, shared dev Redis.
Seed: 50k events (200 arenas with 20k seats), 100k users.
Scenario: 70/20/9/1 mix, 0.5-2s think, 200 VU × 5m plateau.
Results: p95 380ms / p99 720ms global; reserve p95 2.1s in arenas.
Declared capacity: 200 mixed concurrent ≈ 12k reservations/min.
Peak 400: graceful degradation (409 4.2%, no 500s). Recovery: 40s.
Findings: bulk seats UPDATE (74% of DB time); idle-in-tx with the gateway.
Measured improvements: bulk UPDATE → arenas reserve p95 1.4s. Next: partition by event.15 lines worth more than an hour's meeting: hardware, seed, scenario, numbers, findings, next step.
Exercise 5 — The measured improvement
- The before/after:
BEFORE: UPDATE ... WHERE id IN (12 ids) × 4200 calls → 1.8e6 ms total
AFTER: bulk_update in a single statement + (event_id, status) index → 5.4e5 ms
arenas reserve p95: 2.1s → 1.4s (plateau 200) global p95: 380 → 290 ms- Post-improvement capacity: the clean plateau rises to ~280 VUs (+40%): the improvement hit the EXACT bottleneck (measured, not guessed). Next: partition by event (55) because the remaining bottleneck is per-EVENT contention (the arenas), not global — exercise 2's data points to it.
- The product decision: with 280 measured concurrent and a forecast peak of 400, two paths: (a) admit by turns to the checkout (23: the per-event rate limiter as the presale "gate") — product cost: waiting UX; infra cost: ~0; (b) scale to 3 replicas (55): cost ~3× infra, BUT the per-event lock does NOT distribute (contention is over the same event's rows): scaling replicas helps browse, NOT the full-arena checkout. The correct answer: (a) for the arena (the bottleneck is the row), replicas for browse — the decision's mix comes from the numbers, not from fear.
Professor's summary
- Define the SLOs first (p95/p99/errors/correctness) and count the 409 as success: without prior criteria, load testing is theater.
- The bottleneck is localized with pg_stat_statements + pg_stat_activity + per-endpoint breakdown; remote latency is measured separately and not fixed with servers.
- Capacity is a dated document (hardware, seed, scenario) feeding product and infra decisions — and every improvement is re-measured with the same run.