Module 7 · Testing

Lesson 33 — Unit testing done well

pytest, mocks without abusing them and tests that survive refactors.

Published
In this lesson
  1. Objectives
  2. 1. What makes a unit test good
  3. 2. pytest in Django: the minimal setup
  4. 3. Test doubles: the taxonomy and the rule
  5. 4. The anti-patterns that kill suites
  6. 5. TicketFlow's suite structure
  7. Self-assessment

Stack: Django/DRF · Project: TicketFlow Status: Published — opening the testing module Prerequisite: Lesson 32 — Eventual consistency and sagas


Objectives

  1. Write unit tests that survive refactoring: test the WHAT (contract), not the HOW (internal structure).
  2. Master pytest with fixtures and parametrization in Django without giving up the DRF ecosystem.
  3. Use test doubles (fake/mock/stub) with judgment: when a fake is available, don't mock what already exists.

1. What makes a unit test good

Three properties (reduced F.I.R.S.T.): fast (milliseconds: no DB, no network, no real clock), deterministic (same input, same output — that is why 24's injected Clock) and testing observable behavior ("reserving a taken seat raises SeatUnavailable"), not structure ("it calls Seat.objects.filter twice" — that test dies with the first legitimate refactor). TicketFlow's suite runs in <10 s: if your unit test takes 500 ms, it is an integration test in disguise and goes to the slow-tests pit.

python
# tests/unit/test_resumen.py — pytest, pure, no Django
def test_resumen_incluye_expiracion(reserva, clock_fijo):
    r = resumen(reserva, clock=clock_fijo)
    assert r.expires_in_seconds == 600

A unit test's quality test: does it fail when you break the behavior and pass when you refactor? If it fails when renaming a private method (with no behavior change), you were testing the HOW.

2. pytest in Django: the minimal setup

pytest + pytest-django: settings via env var, fixtures instead of setUp, and the real power in parametrization:

python
# pytest.ini / pyproject
[tool.pytest.ini_options]
DJANGO_SETTINGS_MODULE = "ticketflow.settings"
python_files = ["test_*.py"]

# conftest.py — shared fixtures, the suite's backbone
import pytest
from model_bakery import baker

@pytest.fixture
def clock_fijo():
    return FakeClock(datetime(2027, 3, 1, 12, 0, tzinfo=dt.UTC))

@pytest.fixture
def evento(db):                       # db: enables DB access (test transactions)
    return baker.make("events.Event", starts_at=..., capacity=100)

@pytest.mark.parametrize("seats,total", [(1, 5000), (3, 13500), (10, 40000)])
def test_total_con_descuento_por_volumen(seats, total, evento, comprador):
    assert calcular_total(evento, seats) == total

Parametrization turns a case table into one test: 3 cases = 3 runs with individual reporting — the report tells you WHICH row failed, not "the function failed". model_bakery for the data bulk (the verbose Event.objects.create(venue=...,...) fixtures) with the fields that DO matter explicit and the rest defaulted.

3. Test doubles: the taxonomy and the rule

The hierarchy (best to worst): real in-memory object (a list) > fake (a simple Protocol implementation: 24's InMemoryReservationRepository, FakeClock, FakeGateway) > stub (returns fixed values) > mock (verifies INTERACTIONS: "X was called with Y"). The course's rule: mock boundaries, fake states — the gateway (network, money) is ALWAYS mocked or faked; the repository already HAS a fake (25), use it instead of patch("...Reservation.objects"); and the interaction mock only for external effects with no return value you want to verify (the email was sent: an assertion over the SMTP fake, not over "send_mail was called with...").

python
def test_pago_rechazado_compensa(gateway_rechaza, mailbox):
    with pytest.raises(PaymentDeclined):
        confirmar_compra(reserva, gateway=gateway_rechaza, clock=clock_fijo)
    assert reserva.refresh_from_db() or reserva.estado == "CANCELLED"    # behavior
    # NO: gateway.cobrar.assert_called_once_with(...)                    # interaction: only if the effect IS the contract

Mock excess is smell #1 (Sullivan): a suite with 30 mocks per test tests YOUR IMAGINATION about internal calls, not the system. 28's test that catches the outbox-after-commit is possible because we do NOT mock the transaction: we use the real one.

4. The anti-patterns that kill suites

The four horsemen: (1) the refactor test: an assertion about internal calls (above); (2) the dependent test: test_02_expira assumes test_01_reserva ran and left data — every test builds ITS world (fixtures) and runs alone (pytest test_x.py::test_02 green); (3) the guesser test: assertions on complete literal error messages (assert "El asiento A12 fue reservado por otra persona mientras" in str(err)) — assert on err.code/type, the text is for humans and changes; (4) the real-time test: time.sleep(1) waiting for a worker — time is injected (Clock/29), never slept.

And the fifth: mental snapshot: updating assertions "just because" when a test fails after a legitimate change (blind update). Every assertion you change by hand must pass through your head: was it the old behavior or the new one? If you don't know, it wasn't a behavior test.

5. TicketFlow's suite structure

tests/
  unit/        # pure services with fakes: <10 s, runs on every save (33)
  integration/ # real Postgres in a container, transactions/locks/outbox (34)
  e2e/         # full HTTP, <20 tests, the safety net (35)
  contract/    # shapes and problem+json (15/26/28)

Conventions: name test_<subject>_<scenario>_<result> (test_reservar_asiento_ocupado_lanza_conflict); one composite assert per concept; no conditional logic inside the test (an if in a test is two unnamed tests — use parametrization). Coverage: the metric that MATTERS is the domain services' (>90%), not the vanity total: the "covered" settings.py at import proves nothing.


Self-assessment

  1. The three properties of a unit test: which test in your suite breaks them today and how do you mend it?
  2. What turns a test into a "refactor test", and what is the golden rule of WHAT vs HOW?
  3. Doubles taxonomy: when fake, when stub, when mock? Why NOT mock Reservation.objects when InMemoryReservationRepository already exists?
  4. The dependent test: what does it break (parallelization, ordering, CI) and which convention eliminates it?
  5. Why does domain-services coverage >90% matter and total coverage not?

Continue with the exercises. The solutions only after trying it yourself.