Module 5 · Architecture and maintainable code

Lesson 26 — Centralized error handling

One consistent error format for the whole API, not 40 different responses.

Published
In this lesson
  1. Objectives
  2. 1. The problem: 40 error formats
  3. 2. The exception pyramid
  4. 3. The central handler in DRF
  5. 4. Custom exceptions with extra payload
  6. 5. What to log and what not to
  7. Self-assessment

Stack: Django/DRF · Project: TicketFlow Status: Published — the whole API's error contract Prerequisite: Lesson 25 — Service patterns


Objectives

  1. Unify ALL the API's errors into a single format (RFC 7807 application/problem+json), with one piece of code translating them.
  2. Chain exceptions domain → infrastructure → HTTP without each view repeating try/except.
  3. Log what matters (with correlation IDs, 45) and not log what doesn't (payloads with PII, 23).

1. The problem: 40 error formats

Without a central handler, each view picks its format and within the same ticket you get: {"error": "not found"} on one route, {"detail": "Not found"} (DRF default) on another, a 500 HTML page on the third, and a 200 with {"ok": false} on the worst one. The client cannot program against that. The rule: one format for all errors, decided once. Lesson 13 already chose RFC 7807; this lesson implements it end to end.

json
{
  "type": "https://api.ticketflow.dev/problems/seat-unavailable",
  "title": "The seat is no longer available",
  "status": 409,
  "detail": "Seat A12 was reserved by someone else while you were completing the purchase.",
  "instance": "/api/v1/reservations",
  "trace_id": "7c9e6679-7425-40de-944b-e07fc1f90ae7"
}

type is a stable URI the client can document; title is generic (it doesn't change between requests); detail is specific; instance identifies the request; trace_id hooks into the logs (45). Content-Type: application/problem+json.

2. The exception pyramid

Three levels, each with its vocabulary:

python
# Domain (services/models): violated business rules
class DomainError(Exception):
    code = "domain-error"; status = 400; title = "Invalid request"

class SeatUnavailable(DomainError):
    code = "seat-unavailable"; status = 409; title = "Seat unavailable"

class ReservationExpired(DomainError):
    code = "reservation-expired"; status = 410; title = "The reservation expired"

# Infrastructure: the outside world failed
class GatewayTimeout(Exception): ...
class BrokerUnavailable(Exception): ...

# HTTP: only in the view layer (or the central handler)

The service raises SeatUnavailable (domain). The view does NOT catch it: let it bubble. The central handler translates it to problem+json with its status. The view only catches what it decides to handle differently from the general rule — and that is rare.

3. The central handler in DRF

DRF already centralizes serializer/auth errors with EXCEPTION_HANDLER; extending it covers the rest:

python
# core/exceptions.py
from rest_framework.views import exception_handler as drf_exception_handler

def problem_handler(exc, context):
    response = drf_exception_handler(exc, context)   # covers Http404, PermissionDenied, APIException, validation
    if response is not None:                          # DRF knows it: format as a problem
        response.data = to_problem(exc, response.status_code, response.data)
        return response
    if isinstance(exc, DomainError):                  # domain: the new translator
        return problem_response(exc.status, exc.code, exc.title, str(exc))
    logger.exception("unhandled", extra={"trace_id": get_trace_id()})   # 500: the weird gets logged
    return problem_response(500, "internal-error", "Internal error", None)

With settings.EXCEPTION_HANDLER = "core.exceptions.problem_handler" and settings.DEBUG off, Django also renders 500s as JSON via middleware when the route requires it. Implementation keys: domain 4xxs generate NO traceback nor alert (they are business, not bugs); 500s DO (log with logger.exception + alert, 46); instance is filled with request.path; the trace_id is set by the correlation middleware (45) — the handler only reads it.

4. Custom exceptions with extra payload

Sometimes the client needs more than title/detail: the conflicting seats, the retry after X seconds.

python
class SeatUnavailable(DomainError):
    code, status, title = "seat-unavailable", 409, "Seat unavailable"

    def __init__(self, seat_refs: list[str]):
        super().__init__(f"Seats unavailable: {', '.join(seat_refs)}")
        self.extra = {"conflicting_seats": seat_refs}

# the handler serializes exc.extra inside the problem:
{"type": ".../seat-unavailable", "status": 409, "conflicting_seats": ["A12", "B3"]}

Rule for the extra payload: data the client uses to ACT (retry with other seats), never internal state (table names, SQL, stack). The detail is for humans; type+extra, for the client's code.

5. What to log and what not to

The handler passes through 23/45's filter: the validation detail may contain what the user typed (PII) — the log gets type+trace_id+path, not the payload. The 500s with logger.exception (full stack to the log, never to the response body — the 500 body is generic and leaks nothing). Every 5xx carries the trace_id and enters the http_5xx_total metric (46): the handler is also where error observability is born.

Success signal: greps that no longer find return Response({"error":...}) in views — every error goes through the handler and through DomainError.


Self-assessment

  1. Why can't the same ticket coexist with 3 error formats, and which contract does RFC 7807 fix?
  2. Draw the pyramid: who raises DomainError, who translates it, and which layer must NOT catch it?
  3. What does DRF's default exception_handler cover and what does yours add?
  4. When does an exception carry extra in the problem, and what NEVER goes there?
  5. What gets logged in a 500 and what in a domain 409? Where does the traceback end up?

Continue with the exercises. The solutions only after trying it yourself.