Exercise 1 — The lock-free map
- The GCP table (the course's): Cloud Run (2M requests free tier, then ~€0.40/M), Cloud SQL PostgreSQL (no free tier: ~€25 for the small), Memorystore Redis basic 1GB (~€35), Cloud Storage (5GB free), Cloud DNS (~€0.20/zone), Pub/Sub (10GB free). The lab's start: ~60% free tier — prod's real bill: ~€100-150.
- The direct-SDK finding: the GDPR export (23) with
import boto3inside the service (the URL and SDK nailed in). The fix with 24's Protocol:
class ObjectStore(Protocol):
def put(self, key: str, data: bytes) -> None: ...
def signed_url(self, key: str, ttl: timedelta) -> str: ...
class GCSStore: # the boundary's implementation; the service doesn't know it
...The service signs the contract; the provider is one line in settings (27).
- The 15-line translation: what changed between GCP and AWS: names (Cloud Run→Fargate, Cloud SQL→RDS, Memorystore→ElastiCache) and the manifests. What did NOT change: 40's image (the same digest deploys on both), 27's env vars (same structure), 41's pipeline (build/push to another registry), 35's smoke. The provider-change cost: ~1 day of manifests — that is the price of the avoided lock.
Exercise 2 — The compute
- 36's plateau (200 VUs, realistic mix): with 1 concurrency/instance and p95 380 ms → ~52 concurrent instances at the peak (200 VUs × 1 req each ~1 s with think time / 380 ms per req ≈ 76; with the 70/20/9/1 mix the cached browse (38) never reaches the app: ~52 real). Cloud Run pay-per-use cost: ~€45/month at that peak × 1 h/day + minimal baseline. The conclusion: PaaS usage-based autoscaling is the budget's thermostat: the peak only gets paid when it happens.
- The edge timeout: the synchronous GDPR export took 8-12 s (23's job moved it to async with 31, but the big export's download endpoint: 14 s). With a 10 s edge timeout: the request dies AT the edge (the client sees a 504, the work continues and the download is lost). The fix: the export is ALWAYS async with a signed URL (§3) — the 60 s edge timeout covers the p99 × 2 of synchronous endpoints, and nothing synchronous takes >5 s by design.
- The pod's MTTR: with healthcheck + restart: ~25-30 s (the healthcheck detects in ~15 s with interval 5s×3, restart+bind in ~10). On PaaS: instance replacement ~20-60 s. The "dead process" MTTR is cheap; the "zombie process" one (answers 500 without dying) demands the healthcheck with DEPENDENCY (40's /healthz touching the DB) — without it, the healthy-looking pod serves eternal 500s and 47 wonders why that incident's MTTR was 40 min.
Exercise 3 — The data
- The arithmetic against managed: Cloud SQL small = 100 max_connections (the basic tier). 39's table (90 app connections) leaves 10: not enough for the 20% margin. Tuning in order: (a) pool max_size 10→6 per process (total 68: 32% margin); (b) if it doesn't fit: 43's PgBouncer transaction mode (the arithmetic becomes "instances × bouncer_pool" with ITS rules — advisory locks out). The exercise's finding: the managed limit is the deployment's FIRST real constraint, not the CPU.
- The restore drill: the runbook (5 steps): (1) identify the target backup (47's incident moment's PITR); (2) restore to a NEW instance (never over the live one: restoring over live deletes the only copy if it fails); (3) point the test app at the restored one and run the smoke (35): the data is as expected (the check is the BUSINESS's: "is the customer's reservation there?"); (4) time it (the small: 8-14 min restore + 5 min smoke); (5) the app's DNS/config switchover. The drill runs 2×/year in 31's calendar — the untested backup is a hypothesis.
- The bucket: the signed-link evidence: download OK at 5 min;
403 Forbiddenat 16. The retention: 23's policy — expired exports get deleted at 30 days (lifecycle rule) and 56's audit records who downloaded what. The bucket is not a disk: it is a contract with expiry and an owner.
Exercise 4 — The network and the bill
- The VPC (ASCII):
VPC 10.0.0.0/16
├── private-subnet 10.0.1.0/24 # NO public IP
│ ├── Cloud SQL 10.0.1.10 ← only: web(5432), worker(5432)
│ └── Memorystore 10.0.1.11 ← only: web, worker
└── app-subnet 10.0.2.0/24
└── Cloud Run (NAT egress only if it talks to external APIs: gateway)The rules: Postgres accepts ONLY 5432 from the app-subnet's range; Redis likewise; the app accepts 8080 from the edge/proxy. Nothing private has direct Internet egress: the export worker uses GCS's private endpoint (Private Service Connect): NAT doesn't exist in the architecture and its €35/month doesn't either.
- The egress: browse 70% of traffic ≈ 22 GB/month of JSON without CDN → egress ~€1.8 (cheap but grows with 55's ×10 traffic: €18); WITH 38's CDN: 90% leaves from the edge (~€0.6) and the origin only sees the misses. At this scale the CDN doesn't "pay for itself" in euros: it pays in p95 (38 ms vs 210 ms) — the bill is the excuse, latency is the reason.
- The orphans and their tax: the 36 k6 experiment's static IP (€3.6/month), 34's test PG disk (€4/month), the restore drill's snapshot (€2/month), last month's provisional NAT (€35/month), 53's K8s cluster spun up "to try" (€70/month). Total: ~€115/month = 70% of the system's REAL bill. The monthly review (31) with the question "does this serve any living environment?" is the calendar's most profitable task.
Exercise 5 — The IAM
- The export worker's policy:
{
"Version": "2027-01-01",
"Statement": [
{"Effect": "Allow", "Action": ["storage.objects.create"],
"Resource": "projects/ticketflow/buckets/exports-rgpd/objects/*"},
{"Effect": "Deny", "Action": ["storage.objects.delete", "storage.buckets.*"],
"Resource": "*"}
]
}It can: create objects in the exports bucket. It CANNOT: delete them, touch another bucket, or anything IAM/DB. Blast radius if compromised (22): writing garbage into a bucket (sanitizable with lifecycle) — not reading another client's bucket (the policy doesn't allow it even by mistake), not touching the DB (it doesn't hold the identity).
- The anti-pattern found: in the repo's early history (06), a
.envwith the sandbox gateway's key committed for 4 months. The fix: (1) revoke NOW (27's dual rotation in the secrets manager); (2) the key is in git history:git filter-repo/BFG is the treatment, BUT the operating assumption is "every key in history is compromised" — it gets rotated and that's it, history cleanup is cosmetic; (3) 06's pre-commit with the secret-scan (gitleaks) as the guard preventing the next one.
- The audit: Cloud Audit Logs on per project (data access logs for the export bucket). The bucket-deletion event records: actor (IAM identity, not IP), timestamp, source (console/API), and the request's trace. Without the log: the postmortem's (47) "who deleted the bucket?" is a conversation of faith — with the log: one line with a name and a time. The audit log is cheap (~free); the unaudited one is expensive the day the bill comes.
Professor's summary
- The image + per-environment config make the provider reversible; the business speaks Protocols (24), the SDK lives at the boundary.
- Managed for the data (the connection arithmetic stays yours), PaaS for compute, the object store with signed expiring URLs.
- The wall is IAM: service role without eternal keys, least privilege per worker, audit log always — and the bill gets reviewed for orphans, not out of fear.