Why this matters
Every website you've ever used without thinking about it was quietly load balanced. One server can only handle so many requests before it starts sweating — latency climbs, requests queue up, and eventually the whole thing falls over. A load balancer is how we go from "one big server doing its best" to "many servers sharing the work," and it's the reason your favorite apps survive traffic spikes, celebrity tweets, and the occasional server that just gives up on life.
Video: Load Balancers, Properly Explained — Part 1: The Why — whiteboardinterview
Opens with one server buckling under ten thousand users, then shows why every serious service needs a balancer.
Picture a restaurant host
Meet Rita, the host at the busiest restaurant in town. Rita does three things, and they are exactly what a load balancer does:
- She greets everyone at the door — one entrance, all the incoming guests (requests).
- She seats guests across the dining room — spreading them among waiters and tables (servers) so no single waiter gets 40 tables while another stands around polishing glasses.
- She knows who's actually on shift — if a waiter called in sick (a failed server), Rita simply stops seating people at their section. The dining room barely notices.
That's the whole concept. One smart greeter in front of a pool of workers:
flowchart LR
C[Clients] --> LB[Rita the Load Balancer]
LB --> S1[Server 1]
LB --> S2[Server 2]
LB --> S3[Server 3]
LB -. "health checks: who's on shift?" .-> S1
LB -. "health checks: who's on shift?" .-> S2
LB -. "health checks: who's on shift?" .-> S3
(Rita never takes a bathroom break. We'll come back to why that's a problem.)
Video: What is a Load Balancer — Dargslan
20-minute visual explainer with an explicit "restaurant-host analogy" chapter, plus L4 vs L7, algorithms, and health checks in plain English. Caveat: non-trusted channel — quick audio check advised
Where Rita works: DNS, Layer 4, Layer 7
Requests pass through several layers before reaching your app, and Rita can stand at more than one door:
- DNS load balancing — the restaurant's phone book lists several addresses; callers pick one. Crude, but cheap — it works before a connection is even opened.
- Layer 4 (transport) — Rita glances at the name tag (IP address and port) and points you to a table without reading your reservation. Fast and works for any protocol, but she can't route you to the sushi bar versus the grill.
- Layer 7 (application) — Rita actually reads your reservation: "party of four, sushi bar, and you're a regular." She routes based on URL path, host header, or cookies. Smarter, but it costs more of her attention per guest.
Video: What is Load Balancing? Load Balancing Explained — Learn CS
Walks through hardware, software, and cloud balancer types plus L4 vs L7 with timestamped chapters — like sorting mail by address vs. reading the letter inside. Caveat: channel identity unverified; description is detailed human-written
How Rita picks a table: the algorithms
Seating a Friday-night crowd is a strategy question, and Rita has several playbooks:
| Algorithm | How it picks a server | Good for |
|---|---|---|
| Round robin | Cycles through servers in order | Uniform, stateless workloads |
| Weighted round robin | Round robin, biased by server capacity | Mixed hardware fleets |
| Least connections | Sends to the server with fewest active connections | Long-lived / uneven requests |
| IP hash | Hashes client IP to pick a server | Session affinity without shared state |
| Consistent hashing | Maps both servers and keys onto a ring | Caches, sharded stores — minimizes reshuffling when nodes change |
Round robin is the classic "your turn, my turn" — dead simple, and surprisingly effective when every request costs about the same. Least connections is what Rita reaches for when one table is a family of twelve settling in for the evening and another is a coffee-and-go: raw headcounts lie, so she watches active load instead.
flowchart TD
R{New request arrives} --> A{Is it special?}
A -->|Just a regular| RR[Round robin: next table in line]
A -->|Needs the sushi bar| L7["Layer 7: /sushi/* goes to the sushi pool"]
A -->|A returning regular| IH[IP hash: same table as last time]
Interactive diagram: LoadBalancerSim (loads in the app)
Video: Top 6 Load Balancing Algorithms Every Developer Should Know — ByteByteGo
Breaks down six algorithms — round robin, least connections, hashing — with simple diagrams showing when each fits.
Health checks: who's actually on shift
Rita's superpower is that she never sends a guest to a table whose waiter just quit. The load
balancer periodically probes each server — a TCP connect, or an HTTP GET /health — and pulls
unhealthy nodes out of rotation. This turns a single server crash into a brief capacity dip
instead of an outage. Your users keep eating; they never know a waiter left mid-shift.
Video: Load Balancers Part 2: Algorithms, Health Checks & Sticky Sessions — whiteboardinterview
Frames health checks as a morning standup that finds dead servers and pulls them out of rotation.
Layer 7 example: routing by path
A Layer 7 balancer like Rita-with-a-reservation-list can front an entire fleet of independently scaled services behind one hostname:
flowchart LR
LB[Load Balancer]
LB --> N1["/api/orders/* → order pool"]
LB --> N2["/api/users/* → user pool"]
LB --> N3["/static/* → CDN pool"]
This is what lets one public address — app.example.com — quietly serve orders, users, and
images from completely different server pools, each scaled for its own appetite.
Video: EP-09 AWS Application Load Balancer (ALB) Explained — Our Cloud School
Human walkthrough of listener rules, path-based routing, target groups, and health checks with a step-by-step EC2 demo. Caveat: non-trusted channel
The bathroom-break problem
Rita is indispensable. But if Rita is the only host and she steps away for five minutes, the entire restaurant grinds to a halt. A lone load balancer is a single point of failure. The fix is delightfully simple: hire two hosts. Run at least two load balancer nodes behind one floating (virtual) IP, managed by a protocol like VRRP or a cloud provider's managed load balancer. One is active, the other waits in the wings — if the active host faints, the standby picks up the clipboard without the diners noticing:
flowchart LR
VIP["Virtual IP (VRRP)"] --> LB1[Active LB]
VIP -. "failover" .-> LB2[Standby LB]
LB1 --> POOL[Server Pool]
LB2 --> POOL
So Rita finally gets her bathroom break — as long as there's a second Rita on standby. (They split tips.)
Video: AWS Network Load Balancer: Connection Draining & Deregistration Delay Explained — Network Ninja
Explains deregistration delay so in-flight requests finish before a removed target stops serving them.
Takeaways
- A load balancer is the host of your restaurant: one door, smart seating, and a live sense of who's actually on shift.
- Layer 4 is faster (glances at the name tag); Layer 7 is smarter (reads the reservation) — most production systems use both, in layers.
- The algorithm matters most when backends are stateful or unevenly sized — otherwise round robin and a smile go a long way.
- Even Rita needs a backup: always run redundant load balancers behind a virtual IP.
Check your understanding
In the lesson's analogy, what does Rita the restaurant host represent?
- The load balancer — greeting requests, seating them across servers, and tracking server health
- A database server
- The DNS phone book
- A firewall blocking unwanted guests
What is the key difference between Layer 4 and Layer 7 load balancing?
- There is no functional difference, only naming
- Layer 4 is only for HTTPS; Layer 7 is only for HTTP
- Layer 7 is faster because it skips health checks
- Layer 4 balances on IP/port without inspecting content; Layer 7 can route by path, host, or cookies
Which routing algorithm minimizes key/request reshuffling when servers are added or removed?
- IP hash
- Round robin
- Consistent hashing
- Weighted round robin
Why do production teams typically run at least two load balancer nodes?
- To support both Layer 4 and Layer 7 simultaneously
- Because a single load balancer instance would be a single point of failure
- To double the number of routing algorithms available
- Load balancers are required to run in pairs by TCP/IP standards
Go deeper
Want to keep pulling this thread? These talks and tutorials go further than we did here:
- Load balancing in Layer 4 vs Layer 7 with HAPROXY Examples — Hussein Nasser, The Backend Engineering Show (~38 min). The L4/L7 split with hands-on HAProxy examples for both.
- AWS re:Invent 2025 – Deep dive: The evolution of AWS load balancing (NET334) — Matt Lehwess, Milind Kulkarni (AWS), AWS re:Invent 2025 (~1 hr). ALB/NLB internals, weighted target groups, DNS front-ends.
- The System Design Concept That Fixes Traffic Overload (Load Balancing) — The Logic Blueprint, YouTube. Round-robin, least-connections, health checks, sticky sessions, redundant balancers.
- System Design Course – APIs, Databases, Caching, CDNs, Load Balancing & Production Infra — freeCodeCamp.org (~2h 05m). The load-balancing and health-check chapters — how traffic distribution eliminates the single point of failure.
- System Design Concepts Course and Interview Prep — freeCodeCamp.org (~54m). Proxies and load balancers in one concise pass — a quick second angle.
Sources & further reading
- Betsy Beyer et al., Site Reliability Engineering (Google / O'Reilly, 2016), Ch. 19–20 — load balancing at the frontend and within the datacenter.
- David Karger et al., "Consistent Hashing and Random Trees," STOC 1997 — the consistent-hashing technique.
- Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017) — scaling and request routing.
- NGINX and HAProxy documentation — production Layer 4 / Layer 7 configuration and balancing algorithms.