Module 2 · Databases

Lesson 11 — Versioned migrations

Zero-downtime schema changes: how to evolve TicketFlow with real data inside.

Published
In this lesson
  1. Objectives
  2. 1. History as a contract
  3. 2. Expand-contract: changing without falling over
  4. 3. Data migrations: in Python, not loose SQL
  5. 4. The three dangers (and their lint)
  6. 5. Application: Lesson 08's improvements
  7. Self-assessment

Stack: Django 5 + PostgreSQL · Project: TicketFlow Status: Published Prerequisite: Lesson 10 — Transactions


Objectives

  1. Treat migrations as versioned history (never edit what is deployed).
  2. Apply the expand-contract pattern for zero-downtime changes.
  3. Detect dangerous migrations before they take production down.
  4. Apply Lesson 08's three improvements to TicketFlow with safe migrations.

1. History as a contract

makemigrations generates versioned files in the repo. Rules of the craft:

  • Deployed migrations are never edited: if a migration already ran in production, its fix is a NEW migration (never touch the old file).
  • One migration, one intention: "adds field X" — not "Saturday's reshuffle" (irreversible and unreviewable).
  • migrations.RunPython.noop in the reverse: every data migration defines its way back (even if noop); without a reverse there is no deploy rollback.
  • Squash with judgement: after accumulating hundreds, squashmigrations condenses an app's history — only when the old ones no longer run in any live environment.

2. Expand-contract: changing without falling over

Zero-downtime deploys demand that the old and new versions coexist for a few minutes (N pods deploying in a roll). The pattern:

1. EXPAND    add the new thing (nullable column, table, index) — compatible with the old code
2. DUAL WRITE / migrate data   the new code writes old+new; a job backfills the history
3. SWITCH    the code reads/writes only the new thing
4. CONTRACT  drop the old thing (column/constraint) — when no pod runs the old code

Real example: rename events_event.name → title. Wrong way: a direct rename_field (the old version fails because it can't see the column). Expand-contract way: add title, double write, switch, then RemoveField(name).

The detail that sinks deployments: adding a NOT NULL column without a default or creating a blocking index. On modern PostgreSQL: ADD COLUMN... DEFAULT... NOT NULL is metadata (fast), but ALTER COLUMN TYPE or a FK without a prior index blocks. Rule: CREATE INDEX CONCURRENTLY (doesn't block writes; cannot run inside a transaction → atomic = False on the migration).

3. Data migrations: in Python, not loose SQL

python
def set_default_state(apps, schema_editor):
    Event = apps.get_model("events", "Event")
    Event.objects.filter(state__isnull=True).update(state=EventState.DRAFT)

class Migration(migrations.Migration):
    dependencies = [("events", "0004_previous")]
    operations = [
        migrations.RunPython(set_default_state, migrations.RunPython.noop),
    ]

apps.get_model (not the direct import): the historical version of the model, not the current one. Batch with .iterator() + pagination on big tables (don't load 2M rows into memory). Dangerous data migrations (massive regexes, external calls) → never: a migration is deterministic and local.

4. The three dangers (and their lint)

ChangeRiskSafe path
ADD COLUMN NOT NULL without defaultblocking or failing on old rowsnullable → backfill → NOT NULL
DROP COLUMNin-flight old code uses itfull expand-contract first
ALTER COLUMN TYPEtable rewrite, blockingnew column + dual write + switch
CREATE INDEX (big)blocks writesCONCURRENTLY

Tool: django-migration-linter flags dangerous operations in CI. Your deployment checklist: lint in the pipeline → migrate with --noinput after a backup → smoke test the health check (Lesson 02) → rollback = code revert + migrations designed to be backward-compatible one step.

5. Application: Lesson 08's improvements

The three improvements you noted (NOT NULLs, coherence CHECK, unique slug) are applied expand-contract: each one with its migration, data backfill, and green lint. The exercises walk you through it step by step.


Self-assessment

  1. Why is editing an already-deployed migration sabotage even when "nobody has run it locally yet"?
  2. Explain expand-contract with the name→title rename, step by step, and say what each step breaks if you skip it.
  3. Why does CREATE INDEX CONCURRENTLY demand atomic = False, and what do you lose by leaving the transaction?
  4. Why apps.get_model and not from events.models import Event in a RunPython?
  5. Your linter flags RemoveField as dangerous. What is dangerous about dropping a column if the new code no longer uses it?

Continue with the exercises. The solutions only after trying it yourself.