Module 9 · Deployment and operations

Lesson 44 — Infrastructure as code (Terraform)

Infrastructure is versioned, reviewed and destroyed without tears.

Published
In this lesson
  1. Exercise 1 — The inventory
  2. Exercise 2 — The module
  3. Exercise 3 — The state
  4. Exercise 4 — Secrets and destruction
  5. Exercise 5 — The ADR and the drill
  6. Professor's summary

Exercise 1 — The inventory

  1. The project's honest inventory arriving here: the compose DB (in code: docker-compose.yml), Redis (), the simulated export bucket (: "I created it in the lab console on March 14"), the gateway secrets (in the local.env: per 27, otherwise), the test domain/DNS (: "I don't remember the record"). The s are the pending clickops: each with its rebuild or its import.
  2. The drill: 12 min of rebuild (compose up + migrate + seed) BUT with 3 "hand-recalls": 42's NAT variable, staging's DB user, and the sandbox gateway token. Each recall = a pending.tf or a secret going to the manager.
  3. The rule (excerpt): "Every cloud resource lives in infra/. A resource found outside: imported into code within the week or a justified destroy in a PR. The Terraform plan is the only path for infra changes".

Exercise 2 — The module

  1. The staging vs prod diff (the same code's changes):
~ google_sql_database_instance.pg.settings.tier          "db-f1-micro" → "db-custom-2-7680"
~ google_sql_database_instance.pg.settings.availability  "ZONAL" → "REGIONAL"     (46's SLO)
~ google_cloud_run_service.web.metadata.maxScale         "3" → "8"                 (43/39)
+ google_sql_database_instance.pg.settings.backup_configuration (prod: 7-day PITR)
  1. The replacement: the plan says -/+ google_redis_instance.cache (forces replacement) — it destroys the Redis (cache data dies: acceptable) but the LESSON is the method: changing prod's DB name WOULD be a replacement with data: the correct path is 11's applied to infra: new resource with the new name → migrate data (dump/replica) → switch the app → destroy the old one. The plan is ALWAYS read hunting -/+ over data-bearing resources.

Exercise 3 — The state

  1. The lock's evidence:
Terminal A: terraform apply  → Acquiring state lock. This may take a few moments... OK
Terminal B: terraform apply  → Error: Error acquiring the state lock
             → State lock info: ID=..., Who=buffy@laptop, Created=2027-09-28T18:22:01Z

The second apply FAILS (with the lock's owner and time): without the lock, both applies write the state at once and half the resources end up "created but not in the state" — IaC's silent disaster.

  1. The drift runbook (4 steps): (1) 42's monthly terraform plan → the unexpected diff reveals the clickops ("+ a firewall_rule nobody PR'd"); (2) classify: is it legitimate? (47's emergency change: IMPORT it into the code with a PR — the hotfix existed, now it's code); is it a mistake? (3) revert the clickops (the console returns it to the code's state) or accept it and PR; (4) note the finding in the process postmortem: repeated clickops is a missing permission (can someone not wait for the pipeline?).
  1. The state in clear: state pull shows the sensitive attributes (PG's connection string, secret defaults). The protection: bucket with project-only IAM + bucket versioning (a corrupted state gets restored) + encryption at rest (the provider gives it). The document: "the state is a top-class secret: whoever reads the state holds the keys to the kingdom; the state bucket is shared with nobody".

Exercise 4 — Secrets and destruction

  1. The empty grep: the.tf declares the secret's EXISTENCE; the value gets uploaded by the pipeline (41) with the deployer's identity. At runtime, the app reads it from the manager via its service identity (42) — the chain with no secrets in files or the state.
  1. The ceremony:
$ terraform destroy -var="env=prod"
Error: Instance cannot be destroyed: resource google_sql_database_instance.pg
has lifecycle.prevent_destroy set. Remove it explicitly to allow destruction.

The ephemeral's legitimate destroy: the ephemeral variable distinguishes the data environment (dev/staging: no prevent_destroy on NON-data resources; staging's DB with test data: prevent_destroy stays on anything that could be confused with prod). The principle: the guard belongs in the disaster's path, not the work's path.

  1. The nightly staging: saving ~€25/month (60% of staging's cost × 11 h/day off). What it breaks: staging's data die every night (acceptable if staging seeds from 40's seed, never from prod data — 23 forbids that anyway) and 42's restore drill runs against the evening staging (adjustment: the drill before 21:00 or the drill against the lab). The pattern is worth it if staging is truly stateless.

Exercise 5 — The ADR and the drill

  1. The timed drill:
terraform apply (from scratch): 8m40s   (SQL provisioning: 6m — the slow step, unavoidable)
migrate + seed:                 1m10s
35's smoke:                     2m30s
TOTAL: 12m20s  ✓ (<30 min target)

The slow SQL's mitigation: there is none (provisioning is the provider's) — real DR aims at: infra rebuilds in ~12 min, DATA comes from managed's PITR (42, ~10 min) → TicketFlow's full RTO ≈ 25 min. The ADR's number: "Full-environment RTO: ~25 min (measured 2027-09-28)".

  1. The ADR (excerpt): "The infra is code. GCS backend with lock, state per environment, plan in the PR and apply of the saved plan, secrets in Secret Manager (the.tf declares, the pipeline loads), image_digest as the only deployment input, prevent_destroy on data, measured RTO: 25 min. Review: when the resource count outgrows the triangle (network+data+app) or 53's second service arrives".
  1. The README's diagram:
git commit → CI (41) → digest:sha256
                          │
                          ▼
             terraform plan (PR comment)  ← lock: nobody applies without a plan
                          │ merge
                          ▼
             terraform apply tfplan       ← locks: state lock · prevent_destroy
                          │
                          ▼
             smoke (35) → [env dev|staging|prod]
                          │
                          ▼
             remote state (GCS, per env)  ← lock: versioning + IAM

Professor's summary

  • IaC turns infra into reviewed code: the plan in the PR, the apply of the saved plan, and the -/+ read as the bomb it is.
  • The remote state with per-environment lock is the memory and its lock; clickops drift gets reconciled or destroyed, never ignored.
  • Secrets declared but never written into files; prevent_destroy on data; the measured drill turns DR into an ADR number.