Module 3 · API design

Lesson 16 — gRPC, GraphQL, WebSockets and SSE

When REST isn't enough and what each alternative costs you.

Published
In this lesson
  1. Exercise 1 — Matrix
  2. Exercise 2 — SSE
  3. Exercise 3 — From commit to push
  4. Exercise 4 — Load
  5. Exercise 5 — ADR-0007 (skeleton)
  6. Professor's summary

Exercise 1 — Matrix

(a) REST + OpenAPI: a universe of clients, caching, curl. (b) gRPC: strict binary contracts between your own services, streaming, no cache needed. (c) GraphQL could (configurable dashboards) — but if the dashboard has 3 fixed shapes, REST with a light BFF wins; GraphQL pays off with real combinatorics. (d) WebSocket: real bidirectionality. (e) SSE or queue-push (notification): one-way, "ready" happens once — neither WS nor GraphQL.

  1. Today none is urgent: cached polling of availability carries you until real users stare at maps simultaneously. Push gets bought when the business demands the instant — ADR-0007 records it.

Exercise 2 — SSE

  1. curl -N shows the events arriving (event: availability, data: {"free": 41}): the connection stays open.
  2. The PUBLISH from redis-cli appears instantly: the stream is a channel, state lives in Redis/the DB.
  3. Without traffic, intermediaries (LB proxy, Nginx buffers) assume a dead connection and close it (30-60s typical). Keepalive under the timeout keeps NAT/proxies alive; proxy_buffering off stops Nginx from buffering the stream and delivering it in bursts (or never, if the buffer never fills).

Exercise 3 — From commit to push

  1. In code: transaction.on_commit(lambda: redis.publish(...)) — Django runs it only if the commit really happened.
  2. If you publish inside atomic() and then there's a rollback, the client received "seat sold" that never happened (the rollback undid it): the push LIES. Lesson 30's general rule (outbox): publish only after commit; on_commit is the simple version.

Exercise 4 — Load

  1. 200 ESTAB and the process holds (async I/O). Uvicorn's memory per connection: KB, not MB. With keepalive 1s vs 15s: only the ping traffic changes (not the connection count); more pings = trivially more CPU and fewer proxies cutting in.
  2. With 2,000 connections: ulimit -n (open files per process — every socket is an fd) trips at the 1024 default: raise it (systemd LimitNOFILE) and size LB/keepalive. The ops lesson: the limit that matters is not CPU, it's fds and memory per connection — Lessons 05/55.

Exercise 5 — ADR-0007 (skeleton)

Context: buyers stare at maps and seats sell under their feet; cached polling generates duplicate queries and the "instant" reduces checkout frustration. Decision: SSE with Redis pubsub + Last-Event-ID; polling as fallback. Alternatives: polling (not instant), WS (bidirectional overhead with no case). Consequences: persistent connections (size fds/LB), keepalive and buffering off on Nginx, duplicates on reconnect (client-side dedup per event), and a new ops piece (Redis pubsub is already there: Lesson 29's broker uses it).


Professor's summary

  • REST public, gRPC internal if there are microservices, GraphQL only with real combinatorics, push only with a business demanding the instant.
  • One-way SSE + native reconnect covers live availability; WS waits for bidirectionality.
  • Publish after commit (on_commit): a push from a rollback is a lie to the client.
  • The real stream limit: fds and proxies (keepalive/buffering), not CPU.

Next: Lesson 17 — Webhooks and third-party APIs.