Module 1 · Foundations that hold everything up

Lesson 02 — Networking basics

DNS, TCP, latency, load balancers and Nginx in front of your Django.

Published
In this lesson
  1. Exercise 1 — DNS with
  2. Exercise 2 — The phase breakdown
  3. Exercise 3 — The simulated 502
  4. Exercise 4 — Balancing across workers
  5. Exercise 5 — Proxy headers
  6. Exercise 6 — Health check by design
  7. Exercise 7 — Reservation latency
  8. Exercise 8 — Route tracing
  9. Professor's summary

Do not read this without having tried the exercises. Diagnosis is the skill; the solution is only the confirmation.


Exercise 1 — DNS with dig

  1. It returns A records (IPv4; if the machine has IPv6, also AAAA). +short hides the type; in the full output you see it in column 4 (IN → type).
  2. The TTL goes down with each read from cache: dig asks your resolver, and it answers from its cache with the remaining TTL of its copy. When it reaches 0, the resolver re-asks the authoritative server and the TTL returns to its original value. It is not that "the record changed": it is the countdown of the cached copy.
  3. Typical chain: root servers (.) → TLD (org.) → the domain's authoritative server. Each level delegates: nobody knows everything, each server knows who to ask next.
  4. Both resolvers answer the same if the zone is stable, but they may hold cached copies of different ages: during a migration with a short TTL you will see the new IP on one resolver and the old one on another. Practical implication: don't expect instant global consistency; lower the TTL first, verify afterwards from several resolvers.

Exercise 2 — The phase breakdown

  1. Against 127.0.0.1: DNS ≈ 0 (no real resolution, or the OS resolves it instantly), TCP ≈ 0.1-0.5 ms (loopback: no physical network), TLS does not apply (plain HTTP). All the time is your Django processing.
  2. Against external sites time_appconnect (TLS) and time_connect (TCP) dominate: each is worth ~1 RTT, and that RTT is the physical distance (Madrid→Virginia ~90 ms). Distance "shows up" exactly where physics predicts.
  3. Yes, they drop: the 1st request pays DNS (resolver cache), TCP (full handshake) and TLS (full handshake); the 2nd and 3rd reuse the OS/resolver DNS cache and, with modern curl, each invocation opens a new connection, but the resolver already has the answer cached and TLS can resume (session resumption) if the server allows it. To see real keep-alive you need several URLs in the same curl (or HTTP/2): that is the experiment linking back to Lesson 01.

Exercise 3 — The simulated 502

1-2. Gunicorn serves on 127.0.0.1:8000 and curl receives 200 OK.

  1. With Gunicorn dead, the client gets curl: (7) Failed to connect to 127.0.0.1 port 8000: Connection refused. Honest note: that is exactly what Nginx would see when doing proxy_pass against a dead Gunicorn; and to that, Nginx answers the client 502 Bad Gateway. The mental chain: live proxy + dead backend = 502. (If the port itself refused the connection and there were no proxy, the client sees "connection refused", not an HTTP code.)
  2. In error.log you would look for: connect() failed (111: Connection refused) while connecting to upstream (wrong port or dead process) and upstream prematurely closed connection (the worker died mid-request, typically OOM). Then: systemctl status gunicorn, and check that the proxy_pass socket/port matches the service's.

Exercise 4 — Balancing across workers

2-3. In the logs you will see two different PIDs serving requests. Gunicorn with sync workers uses an accept-based pattern: the listening socket is shared and each worker accepts the connection that "wins" (the kernel wakes one). It is not an HTTP balancer's round robin (which decides before forwarding the request); it is the kernel spreading the accept. Similar effect, different mechanism: with very uneven durations, accept-based can get unlucky (one worker piles up the slow requests).

  1. With one worker blocked for 10 s, its next 3 requests wait in the socket queue while the other worker stays free: an HTTP balancer with least connections would have sent them to the free one. Lesson for TicketFlow: payment POSTs (slow) and GETs (fast) coexist better with least connections or with async worker types (gevent, Lesson 03).

Exercise 5 — Proxy headers

  1. With SECURE_PROXY_SSL_HEADER configured and a request without the X-Forwarded-Proto: https header, Django interprets the scheme from what arrives; with a plain HTTP request and no header, request.scheme is http and nothing odd happens. The real poison happens the other way around: if Nginx does send the header (or a client forges it) and you don't strip client headers, an attacker sending X-Forwarded-Proto: https over plain HTTP fools Django about the connection's security. That is why the setting demands the proxy always sets or strips that header.
  2. Without X-Forwarded-For, Django would see the client IP as 127.0.0.1 (Nginx's own): all your audit logs and rate limits would point at the same "IP". Without X-Forwarded-Proto, request.scheme would be http even though the client speaks HTTPS.
  3. On Nginx: proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; and proxy_set_header X-Forwarded-Proto $scheme; — and in Django trust those headers only if traffic comes from your proxy (SECURE_PROXY_SSL_HEADER plus a firewall blocking direct hits to Gunicorn's port).

Exercise 6 — Health check by design

  1. Levels: liveness (does the process respond? cheap, always) and readiness (can I really serve? DB/Redis). Checking the DB on every check every 5 s is 17,280 queries/day per instance: acceptable for a SELECT 1, expensive if the check queries real tables. Recommendation: cheap liveness always; readiness with SELECT 1 (not business queries).
  2. If the DB is slow but responding: 200 on liveness; readiness depends on policy. Returning 503 makes the balancer drop the instance from the pool: if the slowness is a blip, you amplified the outage (less capacity exactly when things go wrong). If it is structural, dropping it is right. Decide with criteria: latency thresholds and flapping (several consecutive failures before dropping).
  3. Deployment version (which build am I talking to?), boot time (is it restart-looping?) and PID (does it match the process I see in top?): those three hints turn "the API is misbehaving" into "the v1.4.2 deployed at 10:03 is misbehaving". Cost: zero. Value in an incident: enormous.

Exercise 7 — Reservation latency

  1. With the DB at 90 ms RTT, 5 trips = ~450 ms added. In a single transaction the trips don't disappear, but they get planned better: with transaction.atomic() you still have 5 statements but you can reduce round trips (batch INSERTs, RETURNING id) and, above all, the lock is held for 450 ms: more contention. The senior solution is minimize-RTT (DB in the same region as the app) — physics is not negotiable; Lesson 10 handles the rest.
  2. It is eventual consistency between cache and origin: the CDN served a copy up to 60 s stale. Business options: low TTL on availability (5-10 s) + invalidation on confirmation; or accept the contradiction and resolve it at reservation time (Lesson 01's 409) with a clear message. The second is more robust: the cache accelerates; the database decides.

Exercise 8 — Route tracing

  1. The hops are intermediate routers; the * are hops that do not answer ICMP with an expired TTL (or drop it). It does not mean "down": it means "does not answer traceroute", and the route continues beyond.
  2. The first one is your router (always answers: it is yours). The last one may not answer ICMP for security policy (rate-limit or dropping ICMP towards its public IP), while serving your traffic perfectly on TCP 443.

Professor's summary

  • The journey is DNS → TCP → TLS → HTTP: every phase is paid for and curl -w breaks it down for you. Diagnose by phases, not by gut feeling.
  • Latency is physical (RTT) and bandwidth is almost never the problem for a JSON API.
  • The load balancer spreads and watches (health checks); your app must be stateless to allow replicas.
  • Nginx in front of Gunicorn: static files, TLS and the X-Forwarded-* headers — without them, Django lies about the world it sees.

When you submit, we correct and move on to Lesson 03 — Concurrency and asynchrony: the event loop, the threads and the race Lesson 01 promised to demonstrate.