Module 7 · Testing

Lesson 35 — E2E and contract testing

The full purchase flow under test and contracts that don't break.

Published
In this lesson
  1. Objectives
  2. 1. What E2E means here (and what it doesn't)
  3. 2. The E2E catalog: few and valuable
  4. 3. Stability: flakiness is a bug
  5. 4. Contract tests: the frontend fails first
  6. 5. The E2E engine room
  7. Self-assessment

Stack: Django/DRF · Project: TicketFlow Status: Published — the full purchase flow under test Prerequisite: Lesson 34 — Integration testing


Objectives

  1. Cover the full purchase flow (search → reserve → pay → confirm → tickets) with few, fast and stable E2E tests.
  2. Pin inter-team contracts with consumer tests (pact-style) that fail before the frontend does.
  3. Decide what NOT to test in E2E: the pyramid holds because the dome is small.

1. What E2E means here (and what it doesn't)

E2E = the system deployed as in production (app + real Postgres + Redis), exercised through its real interface (HTTP), with external dependencies doubled ONLY at the business boundary (gateway sandbox, fake SMTP). It is not "open a browser with Selenium to check the button is blue": TicketFlow is an API; its E2E is an HTTP client executing the user's STORY. The difference from integration (34): there you tested ONE seam (the lock); here the whole STORY — the entire value: "I bought a ticket and it arrives confirmed".

python
# tests/e2e/test_compra_completa.py
def test_compra_feliz(client_api, pasarela_sandbox, mailbox):
    r1 = client_api.post("/api/v1/events?city=Madrid&date=2027-03-01")
    assert r1.status_code == 200
    evento = r1.json()["items"][0]

    r2 = client_api.post("/api/v1/reservations", {evento: evento["id"], "seats": ["A1", "A2"]})
    assert r2.status_code == 201
    ref = r2.json()["public_ref"]

    r3 = client_api.post(f"/api/v1/reservations/{ref}/pay", {"method": "card"})
    assert r3.status_code == 200

    drenar_cola()                                     # the worker's eager mode (29)
    r4 = client_api.get(f"/api/v1/reservations/{ref}")
    assert r4.json()["status"] == "CONFIRMED"          # the story ended well
    assert len(mailbox) == 1

drenar_cola() is the honest piece of async in tests: the worker runs eager inside the process; the test waits for the story to CONVERGE (32's eventual guarantee), not for a sleep.

2. The E2E catalog: few and valuable

The budget: <20 E2E tests, the critical business stories: happy purchase, purchase with decline → compensation (32), purchase with a disputed seat (one seat, one winner), reservation expiring during payment, GDPR export (23), the gateway's incoming webhook (17), saga recovery (UNKNOWN). Each one is the policy of a flow that sells the product. What does NOT go up to E2E: validation edge cases (400 with each bad field — those are serializer unit tests), permissions per role (unit/integration of permissions), performance (36). The fat-E2E mistake: 400 E2E tests "because it's the real thing" — a 90-min suite, fragile (every test depends on global state), nobody maintains it: the dome sinks.

python
# the marker that separates and the budget in CI (41)
@pytest.mark.e2e
def test_compra_feliz(...): ...
# pytest -m "not e2e" in pre-push; pytest -m e2e in the pipeline

3. Stability: flakiness is a bug

An E2E test that "sometimes" fails destroys the whole suite: the team learns to re-run, and within three months nobody looks at red. The three causes and their cures: (1) asynchrony: sleep → drain the queue/active wait with timeout (wait_for(lambda: saga.estado == "CONFIRMED") with a 5 s deadline and a failure that dumps state); (2) data ordering: every test creates ITS event/user (baker), never a "globally seeded" fixture in the DB; (3) shared resources: ports, temp files, Redis DB 15 (34). The team rule: a test failing twice with no clear cause gets QUARANTINED (@pytest.mark.flaky in CI or a custom marker) with an open issue — it is neither deleted nor ignored: it gets cured or tamed.

python
def wait_for(cond, timeout=5.0, msg="condition never arrived"):
    fin = time.monotonic() + timeout
    while time.monotonic() < fin:
        if cond():
            return
        time.sleep(0.05)
    raise AssertionError(f"{msg}: dump={dump_estado()}")

Active wait with deadline and dump: the test's failure reads itself (which saga, which state, which pending queue), no print-hunting.

4. Contract tests: the frontend fails first

The contract between TicketFlow's backend and the front (or the mobile app) breaks silently: you rename seat_refs → seats (28) and the red shows up in production, on the buyer's phone. The contract test pins the agreement BEFORE: the CONSUMER describes what it uses and the PROVIDER verifies it. Pact-style format (in its minimal version, native tests):

python
# tests/contract/test_front_puede_comprar.py — the consumer (front) defines:
CONTRATO_FRONT = {
    "POST /api/v1/reservations": {
        "request": {"event": "uuid", "seats": ["str"]},
        "response": {"status": 201, "body": {"public_ref": "str", "seats": ["str"],
                                              "total": "int", "expires_at": "iso8601"}},
    },
    "GET /api/v1/reservations/{ref}": {"response": {"status": 200, "body": {"status": "str"}}},
}

@pytest.mark.contract
@pytest.mark.parametrize("interaccion", CONTRATO_FRONT.items(), ids=lambda kv: kv[0])
def test_el_front_puede(interaccion, client_api, evento_publicado):
    ruta, esperado = interaccion
    ...runs the request per the contract and validates status + body shape...

Golden rule: the contract tests what the CONSUMER USES, not the whole body — a new field doesn't break it (backwards compatible, 14); a used field that changes, does. In a team of 1 (this course), the contract protects you from YOURSELF in 6 months; in a team, it is the executable agreement that replaces the meeting of "has the front adapted the field yet?".

5. The E2E engine room

The infrastructure that makes the dome possible: (1) the isolated test environment (a compose with app+PG+Redis, 40 formalizes it with depends_on: condition: service_healthy); (2) the deterministic draining of queues (eager mode or a real worker with wait_for); (3) the test user and token (18: a fixture that creates the user and gets an access token, no "manual" login); (4) the gateway sandbox (a FakeGateway mounted as an app or the real sandbox if it has a test mode — 14 gives it idempotency); (5) the clean reset between tests (table truncation + Redis flush, not "restarting the world" — 30 s vs 3 min per test).

The E2E health metric: flaky_rate = no-cause failures / runs < 1%. And the engine room's rule: if tomorrow the business asks "is the checkout broken?", the answer is pytest -m e2e -k compra in 8 minutes — not "let's try it in staging and see".


Self-assessment

  1. What distinguishes E2E from integration, and why does E2E double only the BUSINESS dependencies (gateway) and not the infra ones (Postgres)?
  2. The <20 E2E budget: which stories enter TicketFlow's catalog and which tests stay OUT deliberately?
  3. The three flakiness causes and their cure: why is one flaky test contagious to the whole suite?
  4. In a contract test, what does the consumer define and what does the provider do? Why doesn't a NEW field break the contract?
  5. wait_for with deadline and dump: what does it replace and what does it contribute to failure diagnosis?

Continue with the exercises. The solutions only after trying it yourself.