Module 9 · Deployment and operations

Lesson 42 — Cloud

Compute, storage, networking and IAM on AWS/GCP/Azure.

Published
In this lesson
  1. Objectives
  2. 1. The map: TicketFlow in three languages
  3. 2. Compute: where the container runs
  4. 3. The data: managed vs. homemade
  5. 4. Networking: the VPC that separates and the NAT that bills
  6. 5. IAM: the real wall
  7. 6. The bill: TicketFlow's five items
  8. Self-assessment

Stack: AWS/GCP/Azure (overview) · Project: TicketFlow Status: Published — compute, storage, networking and IAM without the syllabus's vendor lock Prerequisite: Lesson 41 — CI/CD


Objectives

  1. Map TicketFlow's concepts (app, Postgres, Redis, storage, DNS) onto the equivalent services of any cloud, without marrying a provider.
  2. Understand IAM as the cloud's real security wall: service roles, least privilege and the eternal Access Key mistake.
  3. Estimate TicketFlow's deployment cost and know what moves the bill (egress, NAT, high-availability overhead).

1. The map: TicketFlow in three languages

The project needs 6 pieces; all three providers sell them under other names:

PieceAWSGCPAzure
Compute (container)ECS/Fargate, EKSCloud Run, GKEContainer Apps, AKS
Managed PostgresRDSCloud SQLAzure Database
Managed RedisElastiCacheMemorystoreAzure Cache
Objects (GDPR export, logos)S3Cloud StorageBlob Storage
DNS/TLSRoute53 + ACMCloud DNS + certsDNS + Front Door
Queues/messagingSQS/SNSPub/SubService Bus

The concept is the same; the API and the bill vary. The project's discipline: dockerize and configure per environment (27/40) so the provider is a reversible decision — the image is the same; only the deployment manifest and the IaC (44) change. The beginner's mistake: writing the business against the provider's API (boto3 everywhere): 23's storage service defines its Protocol and the S3/GCS implementation stays at the boundary (24).

2. Compute: where the container runs

The three general options: container PaaS (Cloud Run/Fargate/Container Apps: you upload the image, the provider scales and patches; pay per use), managed K8s (GKE/EKS/AKS: full control, operation cost and complexity — 43 decides when it's worth it), classic VM (one machine with docker-compose: the minimum-budget option with the operation on you). For TicketFlow today: container PaaS — 40's image deploys unchanged, autoscaling (55) comes free, and the team of 1 operates no nodes. The decision gets reviewed with 30's threshold (the ADR setting the deadline): if volume ×10 or the business demands complex inter-service networking, K8s enters.

The PaaS configurations that DO matter: concurrency per instance (Cloud Run: simultaneous requests per container — with sync gunicorn, 1; with async, more), the request timeout (32/54: the platform's timeout must exceed your endpoints' p99, or the edge kills healthy requests), and the health probe (40: the container's healthcheck IS the cloud deployment's liveness signal).

3. The data: managed vs. homemade

Managed Postgres (RDS/Cloud SQL): the provider does backups, read replicas (55), failover and patching — for ~2-3× the cost of a VM with hand-rolled Postgres. The team-of-1 arithmetic: an hour of a DBA's (your) time is worth more than the month's delta. The decision: managed ALWAYS for prod; the VM with docker compose for the lab. What managed does NOT gift you: expand-contract migration (11) is still yours, 39's pool/arithmetic still applies (managed max_connections is small: 100-500 per tier — 39's arithmetic is the first meeting with the cloud), and backups must be TESTED (47's restore drill: a never-restored backup is a pendrive waiting to fail).

Object storage (S3/GCS/Blob): the GDPR exports (23), static assets, project logos. Three rules: buckets WITHOUT public listing (Block Public Access on), objects with expiring signed URLs (the user's export downloads via a 15-min link, not an open bucket), and versioning ON for what the GDPR requires being able to prove.

4. Networking: the VPC that separates and the NAT that bills

The minimal topology: a VPC with two subnets (private: DB, Redis, workers — NO public IP; public/proxy: the app serving traffic), security groups allowing ONLY what's needed (the app talks to Postgres's 5432; nothing else reaches Postgres — 22 applied to the network). The detail that bills: private resources without direct egress use a NAT Gateway (~€32/month + egress) — the worker that only talks to the DB and the broker does NOT need NAT if its egress goes through private endpoints (VPC endpoints/Private Service Connect). Egress (data LEAVING the cloud to the Internet) is the bill's surprise item: 38/39's CDN doesn't just speed up: cached traffic leaves through the CDN's edge (cheaper than the origin's egress).

5. IAM: the real wall

The root account's Access Key pasted into a config file is the cloud's original sin. The correct design: service identities (IAM roles attached to the container/function — Cloud Run service account, ECS task role) with NO long-lived keys: the app "is" its identity and the token rotates itself. Least privilege per service:

json
{
  "Role": "ticketflow-export-worker",
  "Allows": ["s3:PutObject on bucket exports/*"],
  "Forbidden": ["s3:*", "s3:DeleteBucket", "iam:*", "the rest of the world"]
}

The GDPR export worker (23) can WRITE to the exports bucket and nothing else: if compromised (22), the blast radius is that bucket. The Access Keys policy when unavoidable (CI that can't use OIDC): scheduled rotation (27: you already designed dual rotation), in the secrets manager (Secrets Manager/Secret Manager), never in the repo nor the image (40 proved it). And the audit: CloudTrail/Cloud Audit Logs ALWAYS on — 47's "who deleted the bucket?" has an answer only if it was recorded.

6. The bill: TicketFlow's five items

Monthly estimate (team of 1, ~50k reservations/month): container PaaS 2 vCPU/4GB ~€35-60; managed Postgres small ~€25-50; managed Redis small ~€15-30; storage+CDN ~€5-15; NAT ~€0-35 (if §4's design avoids NAT: 0). Total: ~€80-190/month. What moves it: egress (the CDN mitigates), multi-zone high availability (×2-3: 46's SLO decides whether it's paid), and forgotten environments (the 24/7 staging nobody uses: turn it off by schedule — and 53's ADR says the modular monolith shrinks the bill too). The monthly bill review is a task of 31's: orphan resources (an elastic IP with no instance, a deleted DB's disk) are growth's tax.


Self-assessment

  1. Why does dockerizing per environment (27/40) keep the provider decision reversible, and which layer SUFFERS if not?
  2. Container PaaS vs K8s vs VM: what does each sell, and which PaaS configurations matter for TicketFlow?
  3. What does managed Postgres gift you, what does it NOT, and which arithmetic stays yours?
  4. Draw TicketFlow's minimal VPC: who lives in the private subnet, and why can the worker NOT need NAT?
  5. Why does the service role beat the eternal Access Key, and what blast radius does the export worker have under least privilege?

Continue with the exercises. The solutions only after trying it yourself.