Module 9 · Deployment and operations

Lesson 43 — Kubernetes and managed services

Basic K8s, Cloud Run or ECS: choosing the complexity your product needs.

Published
In this lesson
  1. Exercise 1 — The documented decision
  2. Exercise 2 — The local cluster
  3. Exercise 3 — The HPA and the CronJob
  4. Exercise 4 — The DB outside
  5. Exercise 5 — The full deployment
  6. Professor's summary

Exercise 1 — The documented decision

  1. The ADR (excerpt): "TicketFlow's platform: container PaaS (Cloud Run). Context: team of 1, 1 modular-monolith app (53), ~50k reservations/month. Reasons: 0 nodes to patch, usage-based autoscaling (~€45/month vs ~€90 for a minimal K8s), 41's pipeline deploys digests without knowing the platform. K8s discarded: the operating cost (updates, RBAC, networking) isn't paid by the product today. Review trigger: a second service with complex networking, or 56's on-prem requirement".
  1. The real triggers (not fashion): (1) a second system that must talk to the first over a private network with per-service policies (53's multi-service): simple PaaS falls short on networking; (2) strict multi-tenancy with per-customer network isolation (56): K8s's Network Policies provide it, PaaS doesn't; (3) the institutional customer's on-prem/edge requirement (56): portable K8s is the deployment standard. The roadmap's most likely: (1) — and it would arrive with 32's payment saga splitting off the monolith, not before.
  1. The 6 objects and their translation: Deployment (compose's --scale replicas with probes), Service (compose's ports: as internal balancing), ConfigMap/Secret (27's .env split into non-secret/secret), HPA (the autoscaling PaaS gives free), CronJob (beat/31 per task), Ingress (39's proxy as a declared object).

Exercise 2 — The local cluster

  1. The normal cycle: kind load docker-image + kubectl apply -f deploy/k8s/ → kubectl get pods:
NAME                              READY   STATUS    RESTARTS   AGE
ticketflow-web-7d4f9c-x2k8l       1/1     Running   0          12s
ticketflow-web-7d4f9c-b9m3p       1/1     Running   0          12s
ticketflow-web-7d4f9c-t4w7q       1/1     Running   0          12s

The local nuance: the image lives in your local docker — kind load pushes it into the cluster's containerd (the registry pull is 41's real path).

  1. The resurrection: killing gunicorn inside the pod → liveness fails in ≤10 s → kubelet restarts the container (same pod, RESTARTS 1) → Ready in ~15-20 s total. The pod's MTTR (42) is now executed by kubelet with no human.
  1. The key contrast: frozen DB → readiness fails (the /ready touches the DB) → the pod leaves the Service's Endpoints (no traffic) BUT does NOT restart (local liveness stays OK). When the DB returns: readiness passes and the pod re-enters balancing by itself. Without the separation: liveness touching the DB would restart ALL pods (cascade) exactly when the DB comes back — §2's thundering herd documented by experiment.

Exercise 3 — The HPA and the CronJob

  1. The HPA finding: 400-VU peak → the HPA scales from 2 to 7 pods in ~25-40 s (metrics-server + decision + pull + readiness): the peak's user SEES the scale-out latency (the initial 2 pods absorb 400 VUs for half a minute: the peak's p95 rises before falling). 55's cure: minReplicas covering the foreseeable peak (the announced presale scales BEFORE — the HPA reacts, the operator anticipates).
  1. The arithmetic with HPA: 10 pods × pool 10 = 100 + worker 8 + beat 2 = 110 > managed's 100. The reasonable tuning: pool max_size 6 (total 70) + maxReplicas 8 (56+10=66: 34% margin) — or 42's PgBouncer if the business demands a high maxReplicas. The rule: the HPA's maxReplicas is part of the connection budget — without that line in 39's table, the scale-out is your own DoS.
  1. The CronJob with the lock: two overlapping kubectl create job --from=cronjob/expirar-reservas → the second's log: run-lock expirar ocupado: skip (31) and it exits 0 in 2 s. The stable logic (28) makes the new transport safe — K8s doesn't need to "make sure" of anything: the lock lives in the service.

Exercise 4 — The DB outside

  1. The documented warning: Postgres in a Deployment + emptyDir → kubectl delete pod → the data leave with the pod (emptyDir lives and dies with it): the DB boots EMPTY. The finding as a runbook warning: "a DB Deployment without a PV isn't a DB, it's a boot cache with 500 users inside".
  1. The PV version: data survive the pod's restart (the PV persists). Even so, the 3 NO-in-prod reasons: (1) failover is artisanal (hand-rolled PG replication vs managed's automatic failover with PITR); (2) backups: nobody tests the PV's restore (42: the drill is the provider's); (3) node drain/maintenance can kick the StatefulSet at the worst moment (the node evacuation carrying the DB = 47's incident with extra steps).
  1. The final connections table:
SourceConnections
web (HPA max 8 × pool 6)48
Celery worker8
beat2
operating margin (20%)~15
Total vs managed (100)73

Exercise 5 — The full deployment

  1. The rolling: maxSurge 1, maxUnavailable 0 → one new pod comes up, passes readiness, enters the Service; one old one drains and dies; ×3. Total time: ~45 s. The old ones coexist with the new during the rollout (the coexistence 11's expand-contract allows).
  1. The measured rollback: broken version (healthcheck 500) → the new pods NEVER pass readiness → the rollout stalls (the Service never sends them traffic: zero 5xx during the bad rollout — readiness IS the guard) → kubectl rollout undo → 20 s to the good version. The contrast with 41's blue-green: K8s protects the user BETTER during the broken deployment (traffic never flips), but blue-green's rollback is more atomic post-switch. The team's answer: rolling for day to day, and 41's blue-green when the release demands it.
  1. The YAMLs in the repo (deploy/k8s/): the directory's README with the kubectl apply -f deploy/k8s/ and 44's principle: manifests get PR'd like code — the 3 a.m. kubectl edit is the drift 44's IaC declares illegal (the terraform plan/kubectl diff catches it the next day).

Professor's summary

  • The platform is chosen by the cost the product pays: PaaS today, K8s when networking/isolation demands it — with the trigger written in the ADR.
  • readiness (do I take traffic?) ≠ liveness (am I alive?): a frozen DB must pull the pod from balancing, not restart it in cascade.
  • The HPA's maxReplicas enters 39's connection arithmetic; the DB lives managed outside the cluster; the YAMLs get reviewed in PRs.