Module 11 · Non-technical skills

Lesson 48 — Documentation and ADRs

Recording design decisions so nobody (including you) has to guess.

Published
In this lesson
  1. Exercise 1 — The docs map
  2. Exercise 2 — The ADRs
  3. Exercise 3 — The living reference
  4. Exercise 4 — The expiry
  5. Exercise 5 — The onboarding
  6. Professor's summary

Exercise 1 — The docs map

  1. The honest audit: a 380-line README mixing installation (tutorial), commands (reference), philosophy (explanation) and 3 buried howtos — nobody navigates the everything-README: the reader looks for THEIR question and gets lost. The typical findings: 47's runbook ( how-to, well done), the.env.example ( reference), 24's philosophy written as a code comment (: it is explanation and belongs in an ADR).
  2. The resulting structure: onboarding.md (tutorial), howtos/{nuevo-endpoint,rotar-secreto}.md, reference/{api,eventos,env}.md, adr/ (12 files + index), postmortems/ (2). Each doc with its review header.
  3. The reader test in the PR: "this doc answers: how do I add an endpoint with problem+json?" — the 60-line how-to serving THAT question; the format's philosophy goes linked to ADR-0002, not mixed in.

Exercise 2 — The ADRs

  1. The model ADR (0009, the orchestrated saga):
markdown
# ADR-0009 — Orchestrated saga for checkout
Date: 2027-09-18 · Status: accepted · Decider: backend
## Context
The charge is remote (32): no distributed transaction. Volume: 50k purchases/day;
the business demands "the user is confirmed in <15 min". Team: 1.
## Decision
Orchestration with the PurchaseSaga table (32): states, reconciliation by intent (14),
idempotent compensations. Choreography for the periphery (emails, analytics).
## Consequences
+ The saga resumes after crashes (the table = truth); UNKNOWN handled (32); auditable.
− One table + states to maintain; the orchestrator is code coupled to the business.
## Review
If 53's multi-service flows exceed 3 orchestrated sagas, evaluate the engine
(temporal/camunda). Trigger: the number of sagas, not fashion.
  1. The index: §3's table with 12 rows + double links (ADR ↔ lesson): the new teammate reads 1 page and knows WHY the system is the way it is.
  1. ADR-0018 and the old one: 0011 stays with Status: superseded by ADR-0018 WITHOUT being deleted or edited (the historical decision was right WITH its then-context: deleting it deletes the learning); 0018 links back. The chain: the decision history is a linked list, not a mutable file.

Exercise 3 — The living reference

  1. The doc↔code test:
python
def test_referencia_de_eventos_coincide():
    eventos_doc = parsear_tabla("docs/reference/eventos.md")          # the doc
    eventos_cod = {e.event_type for e in EVENTOS_REGISTRADOS}         # the code
    assert eventos_cod == eventos_doc, "The events doc lies: sync it"

The doc diverging from code breaks the build: the living reference is the only one that survives the months.

  1. The API schema: manage.py spectacular --file docs/reference/openapi.yaml in the pipeline (41): the API doc is GENERATED, never copied — the PR changing the contract regenerates the yaml and the doc's diff shows up in the PR (50's review sees it).
  1. The env reference: the test reading .env.example and settings.py and verifying every env("X") in the settings has its row in the example (and vice versa): settings↔doc divergence breaks the build. The mechanism (20 lines) buys the config reference's truth forever.

Exercise 4 — The expiry

  1. The header + the check: Last reviewed: 2027-09-28 · Owner: backend in every doc; the CI script:
python
# tools/check_doc_freshness.py (20 lines)
for doc in DOCS.glob("**/*.md"):
    dias = (hoy - parsea_revision(doc)).days
    if dias > 180: warn(doc, "possibly rotten: review or archive")

The warning (not the error: it does not block the merge) shows in the build: explicit expiry beats silence — and the doc reviewed and still valid updates its date in 10 s.

  1. The docs' guard tests: test_listados_sin_total (39) and test_disponibilidad_no_cacheada_en_checkout (38) with the EXACT names linked from the docs: "if this test dies, this doc lies" — the documented statement that is code is the only indestructible one.
  1. The 30-line README: the matrix (tutorial → onboarding.md · how-to → howtos/ · reference → reference/ · why → adr/) + the project's status (compose up, tests, deploy). The door, not the house: the long README is software's most common dead doc.

Exercise 5 — The onboarding

  1. The tutorial (excerpt):
1. git clone && cd ticketflow
2. cp .env.example .env (27) — the dev values work except DATABASE_URL
3. docker compose up -d (40) — wait for the healthchecks
4. docker compose run --rm web python manage.py migrate && seed_demo
5. pytest tests/unit -q → green in <10 s (33)
6. Read docs/adr/README.md (12 decisions, 1 page) and make your introduction PR
  1. The real proof on a clean machine: 3 steps were missing (the node digest for the test front, the seed's TZ, 14's sandbox gateway token). The doc fixed with what was discovered: the tutorial not followed step by step is fiction — the rule: every new doc PR gets followed once.
  1. The 5 that served most (the model answer): 47's runbook (MTTR −40% on incident day), ADR-0009 (the saga: the why that saved 2 debates), 27's.env.example (onboarding without stones), 45's events table (diagnosis by grep), and this proven onboarding (anyone's day-1). The personal rule: when time is tight, document FIRST what the incident or the onboarding will ask tomorrow — the rest can wait for the quarter.

Professor's summary

  • One format per reader (Diátaxis): tutorial, how-to, reference, explanation — the everything-README is the doc nobody reads.
  • The ADR: one decision, context with numbers, honest consequences, review trigger; replaced (Supersede), never edited.
  • Docs-as-code: in the repo, in the PR, generated from code where possible, with explicit expiry and guard tests for statements that can be code.