Module 7 · Testing

Lesson 36 — Load testing (k6)

Find the bottleneck before the concert queue does.

Published
In this lesson
  1. Objectives
  2. 1. What a load test asks (and what it doesn't)
  3. 2. The test environment: honest and isolated
  4. 3. Finding the bottleneck: layer by layer
  5. 4. The peak and the ramp-down: what the curve reveals
  6. 5. From the number to the plan
  7. Self-assessment

Stack: k6 + Django/DRF · Project: TicketFlow Status: Published — closing the testing module Prerequisite: Lesson 35 — E2E and contract testing


Objectives

  1. Write a k6 scenario for TicketFlow's presale: user ramp-up, p95/p99 thresholds and correctness checks under pressure.
  2. Find the bottleneck with profiling discipline (37 goes deeper): measure layer by layer, don't guess.
  3. Decide the honest capacity: how many concurrent users the system supports and on what hardware — the number the business needs.

1. What a load test asks (and what it doesn't)

The right question: how many concurrent users does the checkout support without breaking the SLOs (46)? — not "how much can it take?" (no objective, no criterion). The criteria are defined BEFORE: p95 of POST /reservations < 400 ms, p99 < 900 ms, error rate < 0.5%, zero double-sales (correctness is not negotiable under pressure). And the REALISTIC scenario is not "GET /events 1000 times" (the cached listing withstands miracles, 12): it is the business's mix — 70% browse (listings), 20% detail, 9% reservations, 1% payments — with Friday presale's 10:00 peak.

javascript
// load/checkout.js — the presale scenario
import http from "k6/http";
import { check, sleep } from "k6";
import { Trend } from "k6/metrics";

const reservaDur = new Trend("reserva_duration");

export const options = {
  stages: [
    { duration: "2m", target: 200 },   // ramp: traffic grows, it doesn't jump
    { duration: "5m", target: 200 },   // plateau: the steady state you measure
    { duration: "1m", target: 400 },   // stress peak: where does it break?
    { duration: "1m", target: 0 },     // ramp-down: does it recover?
  ],
  thresholds: {
    "http_req_duration{scenario:default}": ["p(95)<400"],
    "reserva_duration": ["p(99)<900"],
    "http_req_failed": ["rate<0.005"],
    "checks": ["rate>0.999"],
  },
};

export default function () {
  const evento = http.get(`${BASE}/api/v1/events/01H...`).json();
  check(evento, { "detail ok": (r) => r.status === 200 });
  sleep(Math.random() * 2 + 0.5);                      // think time: humans, not machines
  const res = http.post(`${BASE}/api/v1/reservations`,
    JSON.stringify({ event: evento.id, seats: ["A1"] }),
    { headers: { "Content-Type": "application/json", Authorization: `Bearer ${TOKEN}` } });
  reservaDur.add(res.timings.duration);
  check(res, { "reserve 201 or 409": (r) => [201, 409].includes(r.status) });   // 409 is the system SUCCEEDING
}

Two details separating a good script from a naive one: the think time (humans think between actions; without sleep you measure your server against a merciless robot horde) and the check accepting 409 as success (under contention, the conflict is the system working: the trap is counting it as an error and "fixing" what isn't broken).

2. The test environment: honest and isolated

Load is tested against an environment with the SAME topology as prod (app + PG + Redis + worker, 40) but isolated (a small staging or a replica on your machine for the first numbers). Two classic lies: testing against manage.py runserver (a single synchronous process: you measure the dev server, not gunicorn) and testing WITHOUT data (1M seats make the index (09) and the planner behave differently than with 100 rows). The honest setup: gunicorn with prod's workers (the 2×CPU+1 formula as a starting point), a realistic data seed (50k events, 2M seats, 100k users), and the peak's queries' EXPLAIN already validated (09).

Isolation: never against prod (22/47: a badly launched load test IS a DoS incident), and k6's tokens are test users (23's rate limiter will throttle you if you reuse real credentials — and that is another finding).

3. Finding the bottleneck: layer by layer

The bad number arrives ("p95 2 s"), and the guesser says "more servers" (55). The discipline: measure each layer in the SAME run: (1) the p95 per endpoint (k6 separates them with tags); (2) the server side: gunicorn latency logs, pg_stat_statements for the top queries, the connection pool (39: does the p95 coincide with "waiting for connection"?), CPU usage per container (is a 100% core Python or Postgres?). TicketFlow's typical result:

global p95 2.1s  →  breakdown:
  browse (GET /events)        180 ms   ✓ (cached, 12)
  detail (GET /events/:id)    210 ms   ✓
  reserve (POST /reservations) 2.1 s   ✗ the bottleneck
       └─ pg_stat_statements: UPDATE seats SET status WHERE id IN (...) 1.8 s
          └─ the checkout's bulk UPDATE contends over the same event

The bottleneck is almost never "the framework is slow": it is a query (09), a lock (10), the pool (39) or the remote gateway (32: the checkout's p95 includes 1.2 s of the gateway sandbox — local load can't fix someone else's latency; you separate your own latency from the remote one in the measurement).

4. The peak and the ramp-down: what the curve reveals

The curve's phases tell the story: in the ramp (2m→200), the p95 rises smoothly and stabilizes (enough capacity); in the plateau, the honest number (the SLO is met or not); in the stress peak (400), the question is graceful degradation or collapse? — 429/409 errors grow (good: the system rejects with grace) or chained timeouts and OOM (bad: the whole system falls). And the ramp-down (target 0): does latency return to baseline within 30 s? If not: something accumulated (a queue of tasks not drained (29), zombie connections, the whole cache invalidated (38) — "it limped away after the peak" is a finding as valuable as the peak).

bash
k6 run load/checkout.js --out json=results.json   # or influxdb/grafana (46)
k6 run load/checkout.js --vus 100 --duration 5m   # the quick smoke pre-plateau

The fixed run (plateau 200 for 5 m) is what produces the capacity number; exploratory ramps are diagnosis. Every run is recorded with its commit (44: the load environment's IaC) and its result in the repo (load/results/2027-03-01.md) — capacity is data with a date, not a legend.

5. From the number to the plan

The exercise's output: "TicketFlow supports 200 concurrent users of realistic mix (≈ 12k reservations/min) with p95 380 ms on 2 vCPU/4GB; at 400 it degrades with graceful 409s (no corruption); it recovers baseline in 40 s". That number feeds: autoscaling (43/55), the presale queue (the "enter by turns" of real festivals: limiting checkout admission BEFORE the lock contends, 23), and the business's capacity plan. And the improvement list prioritized by measured impact: (1) the seats UPDATE in bulk with a single CASE instead of N updates (09), (2) partitioning reservations by event (55), (3) read replicas for browse (55). Every improvement is re-measured with the SAME k6 run: before/after with numbers, not opinion.


Self-assessment

  1. Which SLOs do you define BEFORE launching k6 for POST /reservations, and why does the 409 count as system success?
  2. Why are think time and a realistic data seed mandatory, and what does each lie measure (runserver, no data)?
  3. The example's bottleneck is the seats UPDATE: how did you localize it with the server's tools, and what is it NOT?
  4. What does each curve phase reveal (ramp/plateau/peak/ramp-down), and which hidden finding does the ramp-down hold?
  5. How does "200 concurrent users" become product and infrastructure decisions?

Continue with the exercises. The solutions only after trying it yourself.