Stack: basic K8s / Cloud Run / ECS · Project: TicketFlow Status: Published — choosing the complexity level your product needs Prerequisite: Lesson 42 — Cloud
Objectives
- Read and write TicketFlow's minimal K8s objects: Deployment, Service, ConfigMap/Secret, HPA and CronJob.
- Decide with judgment between K8s, container PaaS and compose: the decision matrix by team and product size.
- Apply the managed Postgres + K8s rules (the correct pair) and know why the DB inside K8s is the classic mistake.
1. The complexity level the product pays for
Kubernetes is not a maturity trophy: it is a permanent operating cost (cluster upgrades, networking, RBAC, infinite YAML). The honest matrix:
| Scenario | Tool | Why |
|---|---|---|
| Personal project / MVP (team of 1, 1 app) | compose / PaaS | K8s's cost > its benefit |
| Growing SaaS (team 1-5, 2-10 services) | container PaaS (Cloud Run/Fargate) | scaling and patching free, 0 nodes |
| Multi-service with complex networking, strict multi-tenant, or a dedicated platform team | K8s | the platform IS the product |
| Edge/on-prem/regulatory | K8s (or VMs) | when the cloud doesn't reach |
TicketFlow today: PaaS (42). This lesson teaches K8s anyway — because K8s-literacy is the backend's literacy (every big deployment speaks it) and because the migration arrives when the product asks: reading a Deployment stops being optional.
2. TicketFlow's minimal K8s objects
apiVersion: apps/v1
kind: Deployment
metadata: {name: ticketflow-web}
spec:
replicas: 3
selector: {matchLabels: {app: ticketflow-web}}
template:
metadata: {labels: {app: ticketflow-web}}
spec:
containers:
- name: web
image: ghcr.io/ticketflow/app:git-a1b2c3d # 41's digest/SHA, never :latest
ports: [{containerPort: 8000}]
envFrom:
- configMapRef: {name: ticketflow-config} # non-secret config (27)
- secretRef: {name: ticketflow-secrets} # the secrets (27/42)
readinessProbe: {httpGet: {path: /healthz/, port: 8000}, initialDelaySeconds: 5}
livenessProbe: {httpGet: {path: /healthz/, port: 8000}, periodSeconds: 10}
resources:
requests: {cpu: 250m, memory: 512Mi}
limits: {cpu: "1", memory: 1Gi}
---
apiVersion: v1
kind: Service
metadata: {name: ticketflow-web}
spec: {selector: {app: ticketflow-web}, ports: [{port: 80, targetPort: 8000}]}40's dictionary translated: the compose service → Deployment (replicas are explicit or HPA-driven), the port → Service, the .env → ConfigMap+Secret, the healthcheck → readiness (can it take traffic?) + liveness (is it alive and should it restart?). The two probes ARE different: a liveness touching the DB can kill pods in cascade (the DB goes down → all pods "not alive" → mass restart → the DB comes back → thundering herd) — liveness's /healthz is cheap and local; the DB's /ready goes in readiness (the pod leaves the load balancing without dying).
3. The HPA and the CronJob: the automatic and the periodic
CPU/memory autoscaling (K8s's answer to 55's scaling):
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: {name: web-hpa}
spec:
scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: ticketflow-web}
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource: {name: cpu, target: {type: Utilization, averageUtilization: 70}}36's nuance: the HPA reacts to CPU, but TicketFlow's bottleneck was the DB (the event's lock) — scaling web replicas doesn't fix the UPDATE's contention: the HPA scales the tier that scales (web), and the real bottleneck gets attacked with 36/55's decisions (partitioning, presale turns). And the CronJob (31's beat in K8s language): the Celery beat stays a singleton (Deployment replicas=1) OR migrates to a K8s CronJob per task — 31's run lock makes both safe: 28's "stable logic, replaceable transport" pattern.
apiVersion: batch/v1
kind: CronJob
metadata: {name: expirar-reservas}
spec:
schedule: "*/1 * * * *"
jobTemplate:
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: expirar
image: ghcr.io/ticketflow/app:git-a1b2c3d
command: ["python", "manage.py", "expirar_reservas"]4. The database: managed outside, never inside
The K8s-enthusiast's classic mistake: a Postgres StatefulSet inside the cluster. The NO reasons: pod storage (PersistentVolume) is the ecosystem's most fragile resource (the volume failing to mount after a zone outage = the DB that doesn't boot), failover is artisanal compared to managed's (42), the provider's backups (with PITR) disappear, and node drain/maintenance can kick the DB in production. The correct pair: K8s (or PaaS) for everything stateless in the app + managed Postgres/Redis outside the cluster, talking over the VPC (42). The corollary: 39's connection arithmetic becomes critical (each pod × pool × HPA replicas: the HPA scaling to 10 pods can knock over managed's max_connections — 39's budget includes the HPA's maxReplicas).
5. Deployment in K8s: 41's objects translated
The rolling update replaces artisanal blue-green (41): the Deployment with maxSurge: 1, maxUnavailable: 0 brings up new pods (with readiness deciding), drains the old ones: 41's blue-green is still better for atomic rollback, but rolling + kubectl rollout undo is the K8s world's default. Secret/config stays 27's (the ConfigMap/Secret gets versioned in the repo via 44's IaC, never hand-edited with kubectl — the 3 a.m. "kubectl edit" is the config nobody will find tomorrow). And migrations (41): a pre-deploy Job in the pipeline (the migrate job runs BEFORE the rollout, from the same image): the release's canonical order is identical on any platform — that is what dockerizing well means (40).
Self-assessment
- Which scenario pays for K8s and which pays for PaaS? Where is TicketFlow and what would the real migration trigger be?
- Deployment/Service/ConfigMap+Secret/readiness+liveness: which compose object does each translate, and why are the two probes different?
- Why can a liveness touching the DB cause a cascade restart, and which probe does what?
- Why doesn't the HPA fix the checkout's bottleneck, and which piece of 39 becomes critical with the HPA?
- Which three reasons kill a Postgres StatefulSet inside the cluster, and how does the app connect to the managed one?
Continue with the exercises. The solutions only after trying it yourself.