The presale peak under control. No solutions.md before submitting.
Exercise 1 — The script
- Install k6 and write
load/browse.js: 50 VUs, ramp 1m/3m/1m, GET /events and /events/:id with 0.5-2 s think time. Thresholds: p95<300 ms, error rate<0.1%. - Run it against
manage.py runserverand then against gunicorn (4 workers): paste both p95s. The difference IS the lesson (what were you measuring the first time?). - Add the correctness check: the detail returns
seats_available >= 0and the listing's schema satisfies the contract (35). Load without correctness only measures the speed of chaos.
Exercise 2 — The checkout under pressure
- Write
load/checkout.js(the lesson's one): 70/20/9/1 mix, p95/p99 thresholds, and authentication with a test user (18). Run the 200 × 5 min plateau against the honest environment (gunicorn+real PG+2M-seat seed). - The 409 as success: add
rate_409as your own metric (Counter) and verify that on the plateau the 409s do NOT count as failures but DO get recorded (how many? what percentage of total attempts?). - The realistic seed: 50k events, 2M seats with distributions (small events with 500 seats vs arenas with 20k). Run the checkout ONLY against arenas and ONLY against small events: does the p95 differ? Why (the big event's lock, 10?)?
Exercise 3 — The bottleneck, localized
- With
pg_stat_statements: on the 200 plateau, list the top 5 queries by total time. Which is it and how much of the p95 does it explain? - The connection pool (39): measure with
SELECT * FROM pg_stat_activityduring the run: how many connectionsidle in transaction? Does the slow endpoint's p95 coincide with "waiting for connection"? Document the finding. - The remote latency: the fake gateway with configurable delay (0/200/500 ms). Run 3 plateaus: how much of the checkout's p99 is the gateway? Which capacity conclusions CAN you draw without controlling the real network? (54 will cover timeout/breaker).
Exercise 4 — The full curve
- Run the complete scenario (ramp/plateau/peak 400/ramp-down) and graph it (or read the JSON): does the peak degrade gracefully (409/429 grow, latency stable) or collapse (chained timeouts)? Paste the figures.
- The ramp-down: after the peak, how many seconds until the p95 returns to baseline? If it takes >60 s: what accumulated (a pending Celery queue, invalidated cache, zombie connections)? Document the finding.
- The capacity number: with a clean plateau, write
load/results/2027-09-28.mdwith hardware, seed, scenario, results and conclusions in ≤15 lines. It is the document the business understands.
Exercise 5 — The measured improvement
- Take exercise 3's bottleneck (the bulk seats UPDATE) and apply the improvement: a single UPDATE with
CASE/bulk_updatein batches (09). Re-run the SAME plateau: paste the p95 andpg_stat_statementsbefore/after. - Did capacity rise? How many VUs does the clean plateau hold now? Which improvement would you apply NEXT and why (partition by event, 55, or read replicas)?
- The presale plan: with the measured capacity number, write the product decision: do I admit N users per turn to the checkout (23/32) or scale horizontally (55)? Justify with the run's numbers.
Submit
Paste the runserver vs gunicorn p95s, the plateau's result with 409s, the bottleneck finding with pg_stat_statements and the improvement's before/after. Next: Lesson 37 — Profiling (module 8).