Stack: Terraform · Project: TicketFlow Status: Published — closing the deployment module Prerequisite: Lesson 43 — Kubernetes and managed services
Objectives
- Declare TicketFlow's infrastructure in Terraform: the VPC, managed Postgres, Redis, the app's service and its IAM.
- Internalize the write→plan→apply flow with the state and the risks of the manual apply (the
terraform planis the infra's code review). - Apply the project's rules: remote state with lock, secrets outside the code (27) and destruction without drama.
1. Clickops and its grave consequences
Without IaC, the infra lives in the provider's console: nobody knows WHAT exists or WHY (was 42's NAT deliberate or left over from an experiment?), staging diverges from prod (27 broken at the infrastructure level), and disaster recovery (47) is "trying to remember how everything was set up". IaC declares the desired state in versioned files: git IS the source of truth, the PR IS the review, the plan IS the test. The project's rule: if clickops created it, it doesn't really exist — every resource carries its.tf file or gets declared orphan (42's hunting, formalized).
2. TicketFlow's Terraform: the real minimal infra
# infra/main.tf (excerpt; providers and backends in their own files)
variable "env" { type = string } # dev | staging | prod (27)
module "network" {
source = "./modules/network"
env = var.env
cidr = "10.0.0.0/16"
}
resource "google_sql_database_instance" "pg" {
name = "ticketflow-${var.env}"
database_version = "POSTGRES_16"
settings {
tier = var.env == "prod" ? "db-custom-2-7680" : "db-f1-micro"
availability_type = var.env == "prod" ? "REGIONAL" : "ZONAL" # 46's SLO decides the cost
ip_configuration {
ipv4_enabled = false
private_network = module.network.id # 42: private VPC only
}
}
deletion_protection = var.env == "prod" # prod's `destroy` demands a double ceremony
}
resource "google_redis_instance" "cache" {
name = "ticketflow-${var.env}"
memory_size_gb = 1
connect_mode = "PRIVATE_SERVICE_ACCESS"
}
resource "google_cloud_run_service" "web" {
name = "ticketflow-web-${var.env}"
location = "europe-west1"
template {
spec {
containers {
image = var.image_digest # 41's artifact, NEVER :latest
env {
name = "DATABASE_URL"
value = "postgres://...${google_sql_database_instance.pg.private_ip_address}..."
}
}
}
}
metadata { annotations = { "autoscaling.knative.dev/maxScale" = "8" } } # 43's HPA in 39's budget
}The critical piece: var.image_digest arrives from the pipeline (41) — the IaC deploys THE tested artifact, it rebuilds nothing. And the per-environment variables (var.env): the SAME code deploys all three environments with different tier/HA (27 carried into infrastructure: 40's image was "one image, three environments"; this is "one module, three environments").
3. The flow: write → plan → apply
The mandatory cycle: (1) terraform plan (the infra's diff: "+1 instance, ~€30/month" — read like a PR's diff); (2) the plan gets saved as an artifact and the apply applies EXACTLY that plan (not one recomputed at midnight); (3) terraform apply with approval (in CI: the PR shows the plan as a comment; the merge fires the apply). Reading the plan is a skill: the forces replacement (the resource that gets DESTROYED and recreated: a DB with a changed name = data gone) versus the safe in-place changes — the accidentally-deleted-DB plan starts with -/+ resource "google_sql_database_instance" and 3 seconds of reading prevent it.
terraform plan -out=tfplan -var="env=prod" -var="image_digest=sha256:9f2c..."
terraform apply tfplan # applies EXACTLY what was reviewed4. The state: the memory and its lock
Terraform's state is the resource→real-ID map: without it, Terraform doesn't know what exists (and the apply of a lost state can CREATE duplicates). The project's rules: (1) remote backend (GCS/S3 with native lock): the state is NEVER local nor committed (it holds the resources' secrets in clear — the state bucket gets project-only access); (2) the state lock (the backend provides it): two simultaneous applies (you and your colleague, or two pipelines) corrupt it — the lock serializes; (3) the state per environment (dev/staging/prod in distinct paths) with the smallest possible blast radius: the prod-state mistake doesn't touch staging. The drift (someone did clickops): terraform plan reveals it as an unexpected diff — 42's monthly review runs it and reconciles (import the orphan resource or destroy it).
5. Secrets, destruction and the final ADR
Secrets do NOT live in the.tf (plan and state would expose them): the.tf references the manager (Secret Manager) and the runtime reads them (27/42):
resource "google_secret_manager_secret" "django_key" { secret_id = "django-secret-key" }
# the value gets loaded by 41's deployment from the manager; the .tf only declares it EXISTSDestruction (terraform destroy): legitimate in ephemeral environments (the nightly staging shutting down: 42's bill thanks it), with locks in prod (deletion_protection, prevent_destroy in the data modules, and the double-apply ceremony). The module's closing ADR: "TicketFlow's infra is code (44): the state in GCS with per-environment lock, the plan in the PR, secrets in the manager, 41's digest as the only deployment input. The full environment rebuild from scratch must take <30 min — disaster recovery (47) is a terraform apply, not a week of memory".
Self-assessment
- Which three clickops problems does IaC solve, and what does "if clickops created it, it doesn't really exist" mean?
- In §2's.tf: where does
var.image_digestcome from and why NEVER:latest? What doesavailability_typedecide per environment? - The plan: what is a
forces replacement, and what is the difference between a directapplyand applying the savedtfplan? - What is the state, why does it go remote with a lock, and how is clickops drift detected?
- Why don't secrets live in the.tf or the local state, and which ceremony protects prod from destroy?
Continue with the exercises. The solutions only after trying it yourself.