Module 8 · Performance and caching

Lesson 37 — Profiling

Measure before optimizing: where the time actually goes.

Published
In this lesson
  1. Exercise 1 — The baseline
  2. Exercise 2 — The N+1 in disguise
  3. Exercise 3 — The honest micro-profile
  4. Exercise 4 — py-spy on the worker
  5. Exercise 5 — The profiling report
  6. Submit

Measure before optimizing, on TicketFlow. No solutions.md before submitting.

Exercise 1 — The baseline

  1. Write the local profiling test: 500 reservations of one event with cProfile and the top 12 by cumulative. Paste the complete output.
  2. The hypothesis: read the output and write ONE hypothesis of the form "X% of the time is in Y because of Z". Is it the function you suspected or another one?
  3. The minimal improvement your hypothesis suggests (without touching the DB) and the re-measurement with the SAME test: paste the total time's and the top 3's before/after.

Exercise 2 — The N+1 in disguise

  1. Implement the "my reservations" view with the DELIBERATE N+1: a serializer with a SerializerMethodField querying the tariff per row. Measure with django-debug-toolbar (or silk): how many queries for 40 reservations?
  2. Fix it with select_related/prefetch_related and re-measure. How many queries and how much time?
  3. The second disguise: the same serializer accessing reserva.seats.all() inside the loop (the related manager's lazy evaluation). Why does prefetch_related cure it and what happens if you ONLY do .iterator()? (hint: the N+1 is per row, not per batch).

Exercise 3 — The honest micro-profile

  1. With timeit: compare str(uuid) vs f"{uuid}" vs uuid.hex over 100k iterations. Does it matter? Note the real number (expected: it doesn't; the lesson is learning to say no).
  2. line_profiler on calcular_total() (08): which line concentrates the time? Is it the Decimal or the tariff query? (if it's the query: it's the N+1 again, not a CPU problem).
  3. The snapshot finding: apply §5's denormalization (precompute total_amount at confirmation) and verify with 08's test that the price snapshot holds (the tariff's price may change afterwards; the reservation's total does NOT).

Exercise 4 — py-spy on the worker

  1. Launch the local stack (gunicorn with 4 workers + 36's k6 suite in smoke mode) and py-spy dump a worker DURING the plateau: what is it doing? Paste the 3 workers' dumps.
  2. The flamegraph: py-spy record for 60 s during the plateau and paste the SVG (or the stack top if you export it to text). Which function dominates? Does it match 36's DB bottleneck or did it discover something new?
  3. The stuck case: simulate a stuck worker (a Celery task sleeping 120 s) and use py-spy dump to diagnose WITHOUT restarting: what would you tell the on-call with only the dump?

Exercise 5 — The profiling report

  1. Write the docs/perf/2027-09-checkout.md report with: baseline (k6), findings per tool (cProfile/silk/py-spy), applied improvements and re-measurement, and the decided NON-improvements (what did you see slow and is NOT worth fixing?).
  2. The project's rule: every performance improvement in a PR carries ITS measured before/after. Write the PR template (5 lines) and paste it into CONTRIBUTING.md.

Submit

Paste the cProfile output, the N+1's before/after, the py-spy dump and the report. Next: Lesson 38 — Caching strategies.