Module 7 · Testing

Lesson 36 — Load testing (k6)

Find the bottleneck before the concert queue does.

Published
In this lesson
  1. Exercise 1 — The script
  2. Exercise 2 — The checkout under pressure
  3. Exercise 3 — The bottleneck, localized
  4. Exercise 4 — The full curve
  5. Exercise 5 — The measured improvement
  6. Submit

The presale peak under control. No solutions.md before submitting.

Exercise 1 — The script

  1. Install k6 and write load/browse.js: 50 VUs, ramp 1m/3m/1m, GET /events and /events/:id with 0.5-2 s think time. Thresholds: p95<300 ms, error rate<0.1%.
  2. Run it against manage.py runserver and then against gunicorn (4 workers): paste both p95s. The difference IS the lesson (what were you measuring the first time?).
  3. Add the correctness check: the detail returns seats_available >= 0 and the listing's schema satisfies the contract (35). Load without correctness only measures the speed of chaos.

Exercise 2 — The checkout under pressure

  1. Write load/checkout.js (the lesson's one): 70/20/9/1 mix, p95/p99 thresholds, and authentication with a test user (18). Run the 200 × 5 min plateau against the honest environment (gunicorn+real PG+2M-seat seed).
  2. The 409 as success: add rate_409 as your own metric (Counter) and verify that on the plateau the 409s do NOT count as failures but DO get recorded (how many? what percentage of total attempts?).
  3. The realistic seed: 50k events, 2M seats with distributions (small events with 500 seats vs arenas with 20k). Run the checkout ONLY against arenas and ONLY against small events: does the p95 differ? Why (the big event's lock, 10?)?

Exercise 3 — The bottleneck, localized

  1. With pg_stat_statements: on the 200 plateau, list the top 5 queries by total time. Which is it and how much of the p95 does it explain?
  2. The connection pool (39): measure with SELECT * FROM pg_stat_activity during the run: how many connections idle in transaction? Does the slow endpoint's p95 coincide with "waiting for connection"? Document the finding.
  3. The remote latency: the fake gateway with configurable delay (0/200/500 ms). Run 3 plateaus: how much of the checkout's p99 is the gateway? Which capacity conclusions CAN you draw without controlling the real network? (54 will cover timeout/breaker).

Exercise 4 — The full curve

  1. Run the complete scenario (ramp/plateau/peak 400/ramp-down) and graph it (or read the JSON): does the peak degrade gracefully (409/429 grow, latency stable) or collapse (chained timeouts)? Paste the figures.
  2. The ramp-down: after the peak, how many seconds until the p95 returns to baseline? If it takes >60 s: what accumulated (a pending Celery queue, invalidated cache, zombie connections)? Document the finding.
  3. The capacity number: with a clean plateau, write load/results/2027-09-28.md with hardware, seed, scenario, results and conclusions in ≤15 lines. It is the document the business understands.

Exercise 5 — The measured improvement

  1. Take exercise 3's bottleneck (the bulk seats UPDATE) and apply the improvement: a single UPDATE with CASE/bulk_update in batches (09). Re-run the SAME plateau: paste the p95 and pg_stat_statements before/after.
  2. Did capacity rise? How many VUs does the clean plateau hold now? Which improvement would you apply NEXT and why (partition by event, 55, or read replicas)?
  3. The presale plan: with the measured capacity number, write the product decision: do I admit N users per turn to the checkout (23/32) or scale horizontally (55)? Justify with the run's numbers.

Submit

Paste the runserver vs gunicorn p95s, the plateau's result with 409s, the bottleneck finding with pg_stat_statements and the improvement's before/after. Next: Lesson 37 — Profiling (module 8).