Module 9 · Deployment and operations

Lesson 42 — Cloud

Compute, storage, networking and IAM on AWS/GCP/Azure.

Published
In this lesson
  1. Exercise 1 — The lock-free map
  2. Exercise 2 — The compute
  3. Exercise 3 — The data
  4. Exercise 4 — The network and the bill
  5. Exercise 5 — The IAM
  6. Professor's summary

Exercise 1 — The lock-free map

  1. The GCP table (the course's): Cloud Run (2M requests free tier, then ~€0.40/M), Cloud SQL PostgreSQL (no free tier: ~€25 for the small), Memorystore Redis basic 1GB (~€35), Cloud Storage (5GB free), Cloud DNS (~€0.20/zone), Pub/Sub (10GB free). The lab's start: ~60% free tier — prod's real bill: ~€100-150.
  1. The direct-SDK finding: the GDPR export (23) with import boto3 inside the service (the URL and SDK nailed in). The fix with 24's Protocol:
python
class ObjectStore(Protocol):
    def put(self, key: str, data: bytes) -> None: ...
    def signed_url(self, key: str, ttl: timedelta) -> str: ...

class GCSStore:          # the boundary's implementation; the service doesn't know it
    ...

The service signs the contract; the provider is one line in settings (27).

  1. The 15-line translation: what changed between GCP and AWS: names (Cloud Run→Fargate, Cloud SQL→RDS, Memorystore→ElastiCache) and the manifests. What did NOT change: 40's image (the same digest deploys on both), 27's env vars (same structure), 41's pipeline (build/push to another registry), 35's smoke. The provider-change cost: ~1 day of manifests — that is the price of the avoided lock.

Exercise 2 — The compute

  1. 36's plateau (200 VUs, realistic mix): with 1 concurrency/instance and p95 380 ms → ~52 concurrent instances at the peak (200 VUs × 1 req each ~1 s with think time / 380 ms per req ≈ 76; with the 70/20/9/1 mix the cached browse (38) never reaches the app: ~52 real). Cloud Run pay-per-use cost: ~€45/month at that peak × 1 h/day + minimal baseline. The conclusion: PaaS usage-based autoscaling is the budget's thermostat: the peak only gets paid when it happens.
  1. The edge timeout: the synchronous GDPR export took 8-12 s (23's job moved it to async with 31, but the big export's download endpoint: 14 s). With a 10 s edge timeout: the request dies AT the edge (the client sees a 504, the work continues and the download is lost). The fix: the export is ALWAYS async with a signed URL (§3) — the 60 s edge timeout covers the p99 × 2 of synchronous endpoints, and nothing synchronous takes >5 s by design.
  1. The pod's MTTR: with healthcheck + restart: ~25-30 s (the healthcheck detects in ~15 s with interval 5s×3, restart+bind in ~10). On PaaS: instance replacement ~20-60 s. The "dead process" MTTR is cheap; the "zombie process" one (answers 500 without dying) demands the healthcheck with DEPENDENCY (40's /healthz touching the DB) — without it, the healthy-looking pod serves eternal 500s and 47 wonders why that incident's MTTR was 40 min.

Exercise 3 — The data

  1. The arithmetic against managed: Cloud SQL small = 100 max_connections (the basic tier). 39's table (90 app connections) leaves 10: not enough for the 20% margin. Tuning in order: (a) pool max_size 10→6 per process (total 68: 32% margin); (b) if it doesn't fit: 43's PgBouncer transaction mode (the arithmetic becomes "instances × bouncer_pool" with ITS rules — advisory locks out). The exercise's finding: the managed limit is the deployment's FIRST real constraint, not the CPU.
  1. The restore drill: the runbook (5 steps): (1) identify the target backup (47's incident moment's PITR); (2) restore to a NEW instance (never over the live one: restoring over live deletes the only copy if it fails); (3) point the test app at the restored one and run the smoke (35): the data is as expected (the check is the BUSINESS's: "is the customer's reservation there?"); (4) time it (the small: 8-14 min restore + 5 min smoke); (5) the app's DNS/config switchover. The drill runs 2×/year in 31's calendar — the untested backup is a hypothesis.
  1. The bucket: the signed-link evidence: download OK at 5 min; 403 Forbidden at 16. The retention: 23's policy — expired exports get deleted at 30 days (lifecycle rule) and 56's audit records who downloaded what. The bucket is not a disk: it is a contract with expiry and an owner.

Exercise 4 — The network and the bill

  1. The VPC (ASCII):
VPC 10.0.0.0/16
├── private-subnet 10.0.1.0/24        # NO public IP
│   ├── Cloud SQL 10.0.1.10   ← only: web(5432), worker(5432)
│   └── Memorystore 10.0.1.11 ← only: web, worker
└── app-subnet 10.0.2.0/24
    └── Cloud Run (NAT egress only if it talks to external APIs: gateway)

The rules: Postgres accepts ONLY 5432 from the app-subnet's range; Redis likewise; the app accepts 8080 from the edge/proxy. Nothing private has direct Internet egress: the export worker uses GCS's private endpoint (Private Service Connect): NAT doesn't exist in the architecture and its €35/month doesn't either.

  1. The egress: browse 70% of traffic ≈ 22 GB/month of JSON without CDN → egress ~€1.8 (cheap but grows with 55's ×10 traffic: €18); WITH 38's CDN: 90% leaves from the edge (~€0.6) and the origin only sees the misses. At this scale the CDN doesn't "pay for itself" in euros: it pays in p95 (38 ms vs 210 ms) — the bill is the excuse, latency is the reason.
  1. The orphans and their tax: the 36 k6 experiment's static IP (€3.6/month), 34's test PG disk (€4/month), the restore drill's snapshot (€2/month), last month's provisional NAT (€35/month), 53's K8s cluster spun up "to try" (€70/month). Total: ~€115/month = 70% of the system's REAL bill. The monthly review (31) with the question "does this serve any living environment?" is the calendar's most profitable task.

Exercise 5 — The IAM

  1. The export worker's policy:
json
{
  "Version": "2027-01-01",
  "Statement": [
    {"Effect": "Allow", "Action": ["storage.objects.create"],
     "Resource": "projects/ticketflow/buckets/exports-rgpd/objects/*"},
    {"Effect": "Deny",  "Action": ["storage.objects.delete", "storage.buckets.*"],
     "Resource": "*"}
  ]
}

It can: create objects in the exports bucket. It CANNOT: delete them, touch another bucket, or anything IAM/DB. Blast radius if compromised (22): writing garbage into a bucket (sanitizable with lifecycle) — not reading another client's bucket (the policy doesn't allow it even by mistake), not touching the DB (it doesn't hold the identity).

  1. The anti-pattern found: in the repo's early history (06), a .env with the sandbox gateway's key committed for 4 months. The fix: (1) revoke NOW (27's dual rotation in the secrets manager); (2) the key is in git history: git filter-repo/BFG is the treatment, BUT the operating assumption is "every key in history is compromised" — it gets rotated and that's it, history cleanup is cosmetic; (3) 06's pre-commit with the secret-scan (gitleaks) as the guard preventing the next one.
  1. The audit: Cloud Audit Logs on per project (data access logs for the export bucket). The bucket-deletion event records: actor (IAM identity, not IP), timestamp, source (console/API), and the request's trace. Without the log: the postmortem's (47) "who deleted the bucket?" is a conversation of faith — with the log: one line with a name and a time. The audit log is cheap (~free); the unaudited one is expensive the day the bill comes.

Professor's summary

  • The image + per-environment config make the provider reversible; the business speaks Protocols (24), the SDK lives at the boundary.
  • Managed for the data (the connection arithmetic stays yours), PaaS for compute, the object store with signed expiring URLs.
  • The wall is IAM: service role without eternal keys, least privilege per worker, audit log always — and the bill gets reviewed for orphans, not out of fear.