Load Balancers

How traffic gets distributed across a fleet of servers, the layers at which it happens, and the algorithms used to pick a target.

Beginner · 15 min read

Why this matters

Every website you've ever used without thinking about it was quietly load balanced. One server can only handle so many requests before it starts sweating — latency climbs, requests queue up, and eventually the whole thing falls over. A load balancer is how we go from "one big server doing its best" to "many servers sharing the work," and it's the reason your favorite apps survive traffic spikes, celebrity tweets, and the occasional server that just gives up on life.

Video: Load Balancers, Properly Explained — Part 1: The Why — whiteboardinterview
Opens with one server buckling under ten thousand users, then shows why every serious service needs a balancer.

Picture a restaurant host

Meet Rita, the host at the busiest restaurant in town. Rita does three things, and they are exactly what a load balancer does:

  1. She greets everyone at the door — one entrance, all the incoming guests (requests).
  2. She seats guests across the dining room — spreading them among waiters and tables (servers) so no single waiter gets 40 tables while another stands around polishing glasses.
  3. She knows who's actually on shift — if a waiter called in sick (a failed server), Rita simply stops seating people at their section. The dining room barely notices.

That's the whole concept. One smart greeter in front of a pool of workers:

flowchart LR
    C[Clients] --> LB[Rita the Load Balancer]
    LB --> S1[Server 1]
    LB --> S2[Server 2]
    LB --> S3[Server 3]
    LB -. "health checks: who's on shift?" .-> S1
    LB -. "health checks: who's on shift?" .-> S2
    LB -. "health checks: who's on shift?" .-> S3

(Rita never takes a bathroom break. We'll come back to why that's a problem.)

Video: What is a Load Balancer — Dargslan
20-minute visual explainer with an explicit "restaurant-host analogy" chapter, plus L4 vs L7, algorithms, and health checks in plain English. Caveat: non-trusted channel — quick audio check advised

Where Rita works: DNS, Layer 4, Layer 7

Requests pass through several layers before reaching your app, and Rita can stand at more than one door:

Video: What is Load Balancing? Load Balancing Explained — Learn CS
Walks through hardware, software, and cloud balancer types plus L4 vs L7 with timestamped chapters — like sorting mail by address vs. reading the letter inside. Caveat: channel identity unverified; description is detailed human-written

How Rita picks a table: the algorithms

Seating a Friday-night crowd is a strategy question, and Rita has several playbooks:

AlgorithmHow it picks a serverGood for
Round robinCycles through servers in orderUniform, stateless workloads
Weighted round robinRound robin, biased by server capacityMixed hardware fleets
Least connectionsSends to the server with fewest active connectionsLong-lived / uneven requests
IP hashHashes client IP to pick a serverSession affinity without shared state
Consistent hashingMaps both servers and keys onto a ringCaches, sharded stores — minimizes reshuffling when nodes change

Round robin is the classic "your turn, my turn" — dead simple, and surprisingly effective when every request costs about the same. Least connections is what Rita reaches for when one table is a family of twelve settling in for the evening and another is a coffee-and-go: raw headcounts lie, so she watches active load instead.

flowchart TD
    R{New request arrives} --> A{Is it special?}
    A -->|Just a regular| RR[Round robin: next table in line]
    A -->|Needs the sushi bar| L7["Layer 7: /sushi/* goes to the sushi pool"]
    A -->|A returning regular| IH[IP hash: same table as last time]

Interactive diagram: LoadBalancerSim (loads in the app)

Video: Top 6 Load Balancing Algorithms Every Developer Should Know — ByteByteGo
Breaks down six algorithms — round robin, least connections, hashing — with simple diagrams showing when each fits.

Health checks: who's actually on shift

Rita's superpower is that she never sends a guest to a table whose waiter just quit. The load balancer periodically probes each server — a TCP connect, or an HTTP GET /health — and pulls unhealthy nodes out of rotation. This turns a single server crash into a brief capacity dip instead of an outage. Your users keep eating; they never know a waiter left mid-shift.

Video: Load Balancers Part 2: Algorithms, Health Checks & Sticky Sessions — whiteboardinterview
Frames health checks as a morning standup that finds dead servers and pulls them out of rotation.

Layer 7 example: routing by path

A Layer 7 balancer like Rita-with-a-reservation-list can front an entire fleet of independently scaled services behind one hostname:

flowchart LR
    LB[Load Balancer]
    LB --> N1["/api/orders/* → order pool"]
    LB --> N2["/api/users/* → user pool"]
    LB --> N3["/static/* → CDN pool"]

This is what lets one public address — app.example.com — quietly serve orders, users, and images from completely different server pools, each scaled for its own appetite.

Video: EP-09 AWS Application Load Balancer (ALB) Explained — Our Cloud School
Human walkthrough of listener rules, path-based routing, target groups, and health checks with a step-by-step EC2 demo. Caveat: non-trusted channel

The bathroom-break problem

Rita is indispensable. But if Rita is the only host and she steps away for five minutes, the entire restaurant grinds to a halt. A lone load balancer is a single point of failure. The fix is delightfully simple: hire two hosts. Run at least two load balancer nodes behind one floating (virtual) IP, managed by a protocol like VRRP or a cloud provider's managed load balancer. One is active, the other waits in the wings — if the active host faints, the standby picks up the clipboard without the diners noticing:

flowchart LR
    VIP["Virtual IP (VRRP)"] --> LB1[Active LB]
    VIP -. "failover" .-> LB2[Standby LB]
    LB1 --> POOL[Server Pool]
    LB2 --> POOL

So Rita finally gets her bathroom break — as long as there's a second Rita on standby. (They split tips.)

Video: AWS Network Load Balancer: Connection Draining & Deregistration Delay Explained — Network Ninja
Explains deregistration delay so in-flight requests finish before a removed target stops serving them.

Takeaways

  1. A load balancer is the host of your restaurant: one door, smart seating, and a live sense of who's actually on shift.
  2. Layer 4 is faster (glances at the name tag); Layer 7 is smarter (reads the reservation) — most production systems use both, in layers.
  3. The algorithm matters most when backends are stateful or unevenly sized — otherwise round robin and a smile go a long way.
  4. Even Rita needs a backup: always run redundant load balancers behind a virtual IP.

Check your understanding

  1. In the lesson's analogy, what does Rita the restaurant host represent?

    • The load balancer — greeting requests, seating them across servers, and tracking server health
    • A database server
    • The DNS phone book
    • A firewall blocking unwanted guests
  2. What is the key difference between Layer 4 and Layer 7 load balancing?

    • There is no functional difference, only naming
    • Layer 4 is only for HTTPS; Layer 7 is only for HTTP
    • Layer 7 is faster because it skips health checks
    • Layer 4 balances on IP/port without inspecting content; Layer 7 can route by path, host, or cookies
  3. Which routing algorithm minimizes key/request reshuffling when servers are added or removed?

    • IP hash
    • Round robin
    • Consistent hashing
    • Weighted round robin
  4. Why do production teams typically run at least two load balancer nodes?

    • To support both Layer 4 and Layer 7 simultaneously
    • Because a single load balancer instance would be a single point of failure
    • To double the number of routing algorithms available
    • Load balancers are required to run in pairs by TCP/IP standards

Go deeper

Want to keep pulling this thread? These talks and tutorials go further than we did here:

Sources & further reading