Module 1 · Foundations that hold everything up

Lesson 02 — Networking basics

DNS, TCP, latency, load balancers and Nginx in front of your Django.

Published
In this lesson
  1. Objectives
  2. 1. The journey of a request (in 6 hops)
  3. 2. DNS: the Internet's address book
  4. 3. TCP: reliability before speed
  5. 4. Latency: what it is made of and how to measure it
  6. 5. Load balancers: spreading the work
  7. 6. Nginx as reverse proxy (your concrete case)
  8. 7. Diagnostics: the tool stack
  9. Self-assessment (answer me in the chat)

Stack: Django + DRF · Project: TicketFlow Status: Published — taught when you submit the 00b and 01 exercises Prerequisite: Lesson 01 — HTTP in depth


Objectives

By the end of this lesson you will be able to:

  1. Explain the full journey of a request: DNS → TCP → TLS → HTTP → response.
  2. Reason about latency: what it is made of, how it is measured and why bandwidth is almost never the problem.
  3. Decide where a load balancer goes and which algorithm suits each case.
  4. Configure Nginx as a reverse proxy in front of Gunicorn, with TLS and the right headers.
  5. Diagnose network problems with dig, curl, ping, traceroute and ss.

1. The journey of a request (in 6 hops)

When the browser requests https://ticketflow.app/api/events/, this happens:

1. DNS        "Who is ticketflow.app?"  → 203.0.113.42
2. TCP        3-way handshake (SYN, SYN-ACK, ACK)
3. TLS        handshake (1-RTT in TLS 1.3)
4. HTTP       request → Nginx (reverse proxy) → Gunicorn → Django
5. Render     Django queries PostgreSQL, Redis and builds the JSON
6. Return     response back the other way, with keep-alive the connection stays reusable

Every hop adds latency. A senior does not memorize the list: they know how much each hop costs, and that is why they place things where they place them.

Key: Lesson 01 studied hop 4 in depth (HTTP). This lesson studies hops 1-3 and the physical "where": who receives the connection and how it is spread across machines.

2. DNS: the Internet's address book

DNS translates names to addresses. The full chain:

Browser cache → OS cache → ISP resolver (recursive)
  → Root (.) → TLD (.app) → Authoritative (ns.ticketflow.app) → A/AAAA

Record types you will really use:

RecordWhat it answersUse in TicketFlow
A / AAAAIPv4 / IPv6ticketflow.app → 203.0.113.42
CNAMEalias to another namewww → ticketflow.app
ALIAS/ANAMEalias at the apexticketflow.app → load balancer
MXmail serversconfirmation emails must not bounce
TXTarbitrary textdomain verification and SPF

Two properties that matter in production:

  • TTL: how long a response lives in cache. Lower the TTL (say to 60 s) before an IP migration so you can roll back fast; with an 86400 s TTL you are chained to your mistake for a whole day.
  • Propagation: it is not that "the network updates"; it is that each cache keeps its copy until its TTL expires. That is why DNS changes take "from a few minutes to 48 hours".

Note: DNS is usually UDP 53 (fast, stateless); TCP 53 for large responses or zone transfers. If "the domain resolves on your machine but not in production", the first question is which resolver you asked.

3. TCP: reliability before speed

HTTP runs on top of TCP (or QUIC over UDP in HTTP/3). What TCP gives you:

  • 3-way handshake (SYN → SYN-ACK → ACK): one RTT before a single application byte is sent.
  • Reliable, ordered delivery: retransmits what was lost, reorders what arrived out of order.
  • Congestion control (slow start): starts slow and speeds up as long as no packets are lost. A freshly opened connection is slower than an established one: another reason for keep-alive.

Total cost of opening a new connection to ticketflow.app from Europe to a server in Virginia:

HopTypical cost
DNS (cached / not cached)~0 ms / 20-150 ms
TCP handshake1 RTT (~90 ms)
TLS 1.3 handshake1 RTT (~90 ms)
First HTTP response1 RTT + server time

Total latency of the first request: ~270 ms + processing. With keep-alive and reused TLS, later requests stay at ~RTT + server. This is Lesson 01 (keep-alive) seen from the wire.

Key: bandwidth matters for large responses (files, videos). For small JSON APIs, it is almost all latency and request count. Tuning the compression of a 2 KB JSON pays off less than removing one request from the chain.

4. Latency: what it is made of and how to measure it

Request latency = DNS + connection + TLS + transmission + server processing + return. Tools:

bash
ping ticketflow.app            # RTT to the machine (ICMP; sometimes ignored)
curl -w '...' -o /dev/null URL # breakdown by phases (we will do it in exercises)
traceroute ticketflow.app      # network hops to the destination

The curl breakdown you will use all the time:

text
time_namelookup:  DNS
time_connect:     TCP (cumulative)
time_appconnect:  TLS (cumulative)
time_starttransfer: first byte (TTFB, cumulative)
time_total:       end of the response (cumulative)

Pocket rules:

  • Light in fibre: ~1 ms per 100 km round trip approximates well across regions. Madrid-Virginia ~90 ms is not "your slow code": it is physics.
  • High TTFB with low DNS/TCP/TLS = the server is taking long: look at your app (N+1 queries, full queue, GC).
  • Slow DNS = cache or resolver: do not fix it by optimizing Python.

5. Load balancers: spreading the work

With a single server you have a single point of failure and a capacity ceiling. The load balancer receives traffic and spreads it across several identical instances (your app must be stateless! Lesson 27).

Algorithms you must be able to name:

AlgorithmHow it spreadsWhen to use it
Round robinrotating turnsidentical instances, general cases
Least connectionsto whoever has fewest active connectionsvariable-duration requests (your case: long payments)
IP hash / stickysame IP always to the same backendin-memory sessions (avoid: external sessions)
Weightedby weight (bigger machine carries more)canary deployments or mixed hardware

In the course stack: Nginx in front of 2+ Gunicorn workers is already a local balancer; in production, the balancer comes from the cloud (ALB on AWS, Cloud Load Balancing on GCP) or Nginx/HAProxy on your VM.

Two balancer decisions people forget:

  1. TLS termination: the certificate lives on the balancer, which decrypts; towards the app traffic travels plain over the private network (or mTLS if compliance demands it). You saw this in Lesson 01, Exercise 8.
  2. Health checks: the balancer asks GET /api/health/ on each backend and drops failing ones from the pool. That is why Lesson 00's first endpoint was not a whim.

6. Nginx as reverse proxy (your concrete case)

TicketFlow architecture on one VM:

Internet → Nginx (:80/:443) ─┬→ /api/  → Gunicorn (:8000) → Django
                             ├→ /admin/→ Gunicorn
                             └→ /static/ → files served by Nginx

Minimal configuration (commented in the exercises):

nginx
server {
    listen 443 ssl;
    server_name ticketflow.app;

    ssl_certificate     /etc/letsencrypt/live/ticketflow.app/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/ticketflow.app/privkey.pem;

    # Static: Nginx is fast at serving files; Django never sees them.
    location /static/ {
        alias /srv/ticketflow/static/;
        expires 30d;
    }

    location / {
        proxy_pass http://127.0.0.1:8000;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
        proxy_read_timeout 60s;
    }
}

Why each X-Forwarded-* header exists: your Django, behind the proxy, would see the client IP as Nginx's own and the scheme as http unless you tell it. With USE_X_FORWARDED_HOST and the right middleware, Django reconstructs the truth. Without it: pagination with broken URLs, CSRF screaming and useless audit logs.

Fine note: the default proxy_read_timeout (60 s) kills long responses. A report export that takes 3 minutes will die with a 504 even while Django keeps working. The right answer is to return 202 and process in a queue (Lesson 06), not to raise the timeout to 10 minutes.

7. Diagnostics: the tool stack

SymptomToolWhat to look at
"Doesn't resolve"dig +short domainrecord, TTL, resolver used
"Doesn't connect"curl -vwhich phase it stops at (DNS/TCP/TLS/HTTP)
"It's slow"curl -w by phaseswhich phase accumulates the time
"502 Bad Gateway"Nginx error.logGunicorn down or wrong port
"504 Gateway Timeout"proxy_read_timeoutrequest slower than the timeout
"Hanging connections"ss -tananomalous CLOSE_WAIT/ESTAB states

The rule of the craft: locate the phase before touching anything. Changing Python code because DNS takes 300 ms fixes nothing.


Self-assessment (answer me in the chat)

  1. Why is lowering the DNS TTL the first step before an IP migration, and not after?
  2. A request takes 400 ms; DNS 5 ms, TCP 90 ms, TLS 90 ms, TTFB 210 ms. Where is the problem and what would you look at?
  3. Why is "least connections" better than "round robin" for TicketFlow, where the payment POST is slower than a GET?
  4. What breaks in Django if Nginx does not send X-Forwarded-Proto https? Name two symptoms.
  5. Which endpoint of your API is the balancer's health check and why must it be cheap?

Continue with the exercises. The solutions only after trying it yourself.