Module 9 · Deployment and operations

Lesson 43 — Kubernetes and managed services

Basic K8s, Cloud Run or ECS: choosing the complexity your product needs.

Published
In this lesson
  1. Objectives
  2. 1. The complexity level the product pays for
  3. 2. TicketFlow's minimal K8s objects
  4. 3. The HPA and the CronJob: the automatic and the periodic
  5. 4. The database: managed outside, never inside
  6. 5. Deployment in K8s: 41's objects translated
  7. Self-assessment

Stack: basic K8s / Cloud Run / ECS · Project: TicketFlow Status: Published — choosing the complexity level your product needs Prerequisite: Lesson 42 — Cloud


Objectives

  1. Read and write TicketFlow's minimal K8s objects: Deployment, Service, ConfigMap/Secret, HPA and CronJob.
  2. Decide with judgment between K8s, container PaaS and compose: the decision matrix by team and product size.
  3. Apply the managed Postgres + K8s rules (the correct pair) and know why the DB inside K8s is the classic mistake.

1. The complexity level the product pays for

Kubernetes is not a maturity trophy: it is a permanent operating cost (cluster upgrades, networking, RBAC, infinite YAML). The honest matrix:

ScenarioToolWhy
Personal project / MVP (team of 1, 1 app)compose / PaaSK8s's cost > its benefit
Growing SaaS (team 1-5, 2-10 services)container PaaS (Cloud Run/Fargate)scaling and patching free, 0 nodes
Multi-service with complex networking, strict multi-tenant, or a dedicated platform teamK8sthe platform IS the product
Edge/on-prem/regulatoryK8s (or VMs)when the cloud doesn't reach

TicketFlow today: PaaS (42). This lesson teaches K8s anyway — because K8s-literacy is the backend's literacy (every big deployment speaks it) and because the migration arrives when the product asks: reading a Deployment stops being optional.

2. TicketFlow's minimal K8s objects

yaml
apiVersion: apps/v1
kind: Deployment
metadata: {name: ticketflow-web}
spec:
  replicas: 3
  selector: {matchLabels: {app: ticketflow-web}}
  template:
    metadata: {labels: {app: ticketflow-web}}
    spec:
      containers:
        - name: web
          image: ghcr.io/ticketflow/app:git-a1b2c3d     # 41's digest/SHA, never :latest
          ports: [{containerPort: 8000}]
          envFrom:
            - configMapRef: {name: ticketflow-config}    # non-secret config (27)
            - secretRef: {name: ticketflow-secrets}      # the secrets (27/42)
          readinessProbe: {httpGet: {path: /healthz/, port: 8000}, initialDelaySeconds: 5}
          livenessProbe:  {httpGet: {path: /healthz/, port: 8000}, periodSeconds: 10}
          resources:
            requests: {cpu: 250m, memory: 512Mi}
            limits:   {cpu: "1",  memory: 1Gi}
---
apiVersion: v1
kind: Service
metadata: {name: ticketflow-web}
spec: {selector: {app: ticketflow-web}, ports: [{port: 80, targetPort: 8000}]}

40's dictionary translated: the compose service → Deployment (replicas are explicit or HPA-driven), the port → Service, the .env → ConfigMap+Secret, the healthcheck → readiness (can it take traffic?) + liveness (is it alive and should it restart?). The two probes ARE different: a liveness touching the DB can kill pods in cascade (the DB goes down → all pods "not alive" → mass restart → the DB comes back → thundering herd) — liveness's /healthz is cheap and local; the DB's /ready goes in readiness (the pod leaves the load balancing without dying).

3. The HPA and the CronJob: the automatic and the periodic

CPU/memory autoscaling (K8s's answer to 55's scaling):

yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: {name: web-hpa}
spec:
  scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: ticketflow-web}
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource: {name: cpu, target: {type: Utilization, averageUtilization: 70}}

36's nuance: the HPA reacts to CPU, but TicketFlow's bottleneck was the DB (the event's lock) — scaling web replicas doesn't fix the UPDATE's contention: the HPA scales the tier that scales (web), and the real bottleneck gets attacked with 36/55's decisions (partitioning, presale turns). And the CronJob (31's beat in K8s language): the Celery beat stays a singleton (Deployment replicas=1) OR migrates to a K8s CronJob per task — 31's run lock makes both safe: 28's "stable logic, replaceable transport" pattern.

yaml
apiVersion: batch/v1
kind: CronJob
metadata: {name: expirar-reservas}
spec:
  schedule: "*/1 * * * *"
  jobTemplate:
    spec:
      template:
        spec:
          restartPolicy: OnFailure
          containers:
            - name: expirar
              image: ghcr.io/ticketflow/app:git-a1b2c3d
              command: ["python", "manage.py", "expirar_reservas"]

4. The database: managed outside, never inside

The K8s-enthusiast's classic mistake: a Postgres StatefulSet inside the cluster. The NO reasons: pod storage (PersistentVolume) is the ecosystem's most fragile resource (the volume failing to mount after a zone outage = the DB that doesn't boot), failover is artisanal compared to managed's (42), the provider's backups (with PITR) disappear, and node drain/maintenance can kick the DB in production. The correct pair: K8s (or PaaS) for everything stateless in the app + managed Postgres/Redis outside the cluster, talking over the VPC (42). The corollary: 39's connection arithmetic becomes critical (each pod × pool × HPA replicas: the HPA scaling to 10 pods can knock over managed's max_connections — 39's budget includes the HPA's maxReplicas).

5. Deployment in K8s: 41's objects translated

The rolling update replaces artisanal blue-green (41): the Deployment with maxSurge: 1, maxUnavailable: 0 brings up new pods (with readiness deciding), drains the old ones: 41's blue-green is still better for atomic rollback, but rolling + kubectl rollout undo is the K8s world's default. Secret/config stays 27's (the ConfigMap/Secret gets versioned in the repo via 44's IaC, never hand-edited with kubectl — the 3 a.m. "kubectl edit" is the config nobody will find tomorrow). And migrations (41): a pre-deploy Job in the pipeline (the migrate job runs BEFORE the rollout, from the same image): the release's canonical order is identical on any platform — that is what dockerizing well means (40).


Self-assessment

  1. Which scenario pays for K8s and which pays for PaaS? Where is TicketFlow and what would the real migration trigger be?
  2. Deployment/Service/ConfigMap+Secret/readiness+liveness: which compose object does each translate, and why are the two probes different?
  3. Why can a liveness touching the DB cause a cascade restart, and which probe does what?
  4. Why doesn't the HPA fix the checkout's bottleneck, and which piece of 39 becomes critical with the HPA?
  5. Which three reasons kill a Postgres StatefulSet inside the cluster, and how does the app connect to the managed one?

Continue with the exercises. The solutions only after trying it yourself.