TicketFlow's infra in files. No solutions.md before submitting.
Exercise 1 — The clickops inventory
- Inventory YOUR current infra (40's compose counts as "manual infra"): list every resource (DB, Redis, buckets, domains, secrets) and mark: is it in code? who created it and when? can it be rebuilt from scratch?
- The disaster-recovery test: in a lab environment, delete EVERYTHING (containers, volumes, loose config files) and rebuild using ONLY what's in the repo. Time it. What did you have to remember by hand? (that is the pending clickops).
- Write the project's rule into CONTRIBUTING: every cloud resource carries its.tf; clickops gets reconciled (import or destroy) within the week.
Exercise 2 — TicketFlow's module
- Write the complete
infra/main.tfof §2 for YOUR provider (or for GCP following the example): the private network, managed Postgres (with deletion_protection), Redis, and the app's service withvar.image_digest. - Multi-environment: with
var.env(dev/staging/prod), verify the same module produces 3 different plans: paste the plan diff between staging and prod (what changes? tier, HA, maxScale?). - Reading the plan: introduce a
forces replacementon purpose (change the Redis instance's name) and read the plan: what does it destroy? What would you do to change the name WITHOUT losing data (a new resource + data migration, 11 applied to infra)?
Exercise 3 — The state and the lock
- Configure the remote backend (GCS/S3) with native lock and the state per environment (
env/dev,env/staging,env/prod). Verify the lock: two simultaneousterraform apply(two terminals) → does the second wait or fail with the lock message? - The drift: manually change a resource in the console (or simulate: edit the resource outside Terraform) and run
terraform plan: does the plan detect it? What are the options (import into the code, revert the clickops, ignore)? Document the 4-step drift runbook. - The state and secrets: inspect the state (
terraform state pull | jq): which secrets are in clear? What protection does the state bucket have (IAM, encryption, bucket versioning)? Document the decision.
Exercise 4 — Secrets and destruction
- Move the secrets out of the.tf: the secret exists as a resource (Secret Manager) but the VALUE gets uploaded by 41's pipeline (or a manual
gcloud secrets versions addwith 27's dual rotation). Verify:grep -r "secret_value" infra/empty. - The destroy ceremony: add
prevent_destroyto the DB module and tryterraform destroyin dev: which message do you get? Then: the LEGITIMATE destroy of the ephemeral environment (how do you tell them apart? anephemeral = truevariable relaxing the guard in dev?). - The nightly staging: write 31's job (or the provider's cron) that destroys staging at 21:00 and recreates it at 8:00 with
terraform apply. How much is the monthly saving with 42's numbers? What does this pattern break (staging's data? the restore drill's DB)?
Exercise 5 — The ADR and the final drill
- The closing ADR: "TicketFlow's infra is code" — decisions (state backend, lock, state per environment, secrets in the manager, digest as input) in 12 lines with the review threshold.
- The drill: clean environment →
terraform applyfrom scratch → 35's smoke green. Time the total (§5's target: <30 min). Which part is slowest (SQL provisioning? DNS?) and how is it mitigated? - The module's final diagram: the commit→CI (41)→digest→plan→apply→smoke flow in ASCII with the locks (state lock, deletion_protection, prevent_destroy). This diagram IS the
infra/directory's README.
Submit
Paste the clickops inventory, the multi-environment plan diff, the state-lock evidence and the timed drill. Next: Lesson 45 — Structured logs (module 10).