Why this matters
Almost every performance win in backend engineering is secretly a caching win. When a homepage loads in 40 milliseconds instead of 400, when a Black Friday sale doesn't melt the database, when an API serves a million identical requests without breaking a sweat — there is almost always a cache doing the heavy lifting. Caching is the best bang-for-buck optimization in backend engineering — and also one of the easiest to get subtly wrong.
Think of a cache like a chef's mise en place — the little bowls of chopped ingredients sitting within arm's reach. The walk-in freezer (your database) holds everything, but walking back and forth to it for every single order would bring the kitchen to its knees. The ingredients you use constantly live right next to the stove. Caching is mise en place for data: keep the hot stuff close.
Video: 7. Caching Explained — Cache Hits, TTL, Eviction & Redis (System Design Basics) — ByteScaler
Like keeping a sticky-note copy of the DB on your desk instead of walking to the filing cabinet, with a TTL alarm so the note never goes stale.
The core trade-off
A cache buys you speed and reduced load in exchange for staleness and complexity. There are only two hard problems in computer science, as the old joke goes — cache invalidation, naming things, and off-by-one errors. (The joke itself contains an off-by-one error. You're welcome.) Every strategy below is a different answer to one question: when this data changes, who is responsible for making the cache agree?
Video: How One Cache Turns 200ms Into 1ms — Caching Strategies Explained — prahack
The 200ms-to-1ms win, then the catch: a second copy drifts out of date.
Cache-aside (lazy loading)
The most common pattern. Your application checks the cache first; on a miss, it loads from the database and stores the result in the cache for next time.
sequenceDiagram
participant App
participant Cache as Cache (Redis)
participant DB as Database
App->>Cache: GET user:42
Cache-->>App: MISS
App->>DB: SELECT * FROM users WHERE id=42
DB-->>App: row
App->>Cache: SET user:42 = row (TTL 300s)
App-->>App: serve row
The beauty: the cache only ever holds data something actually asked for — no wasted memory. The danger: the first request after a deploy (or a cache flush) hits the database hard, and your application code now owns all the invalidation logic. Most teams using Memcached or Redis start here.
Interactive diagram: VsToggle (loads in the app)
Video: A Step-by-Step Guide for the Cache-Aside Pattern + Stampede Protection — Milan Jovanović
An established educator walks cache-aside through implementation, pros and cons included.
Read-through and write-through
In read-through, the cache itself knows how to load from the database on a miss — your application only ever talks to the cache. In write-through, every write goes to the cache and the database synchronously, so the cache stays fresh on the write path. The price: every write waits on both. And even so, concurrent reads racing a write — or a partial failure between the two writes — can still surface briefly stale data.
flowchart LR
subgraph "Write-through"
A1[App] --> C1[Cache] --> D1[Database]
end
subgraph "Write-behind"
A2[App] --> C2[Cache] -. "async flush" .-> D2[Database]
end
Interactive diagram: CacheFlow (loads in the app)
Write-behind (write-back) flips the trade: the write lands in the cache and returns immediately; the database is updated asynchronously. Writes feel instant, but if the cache node dies before the flush, that write is gone forever. Databases like Cassandra look like write-behind at a glance, but they aren't: every write is appended to a durable commit log synchronously, and the write is only acknowledged once it's safe on disk. Cassandra's write path is not write-behind — durability is synchronous; the async part is only compaction, the background merging of SSTables. Application caches rarely have that safety net, so they should avoid write-behind unless you can afford to lose recent writes.
Video: System Design Part 3: Caching Layer — Sketch Mind
Dedicated chapters contrast read-through vs write-through, and the freshness each one buys.
Eviction: when the mise en place overflows
Memory is finite, so caches evict. The classics:
- LRU (least recently used) — evict whatever hasn't been touched in the longest. Great default; Redis and Memcached both offer it.
- LFU (least frequently used) — evict the least popular. Better when a few hot keys dominate, worse when access patterns shift suddenly.
- TTL (time to live) — don't even wait for pressure; expire entries after N seconds. Simple, predictable, and the reason your "new profile picture" takes five minutes to show up.
TTL deserves special respect: it is the dumbest invalidation strategy and therefore the most reliable. Short TTLs mean fresher data and more database load; long TTLs mean the reverse. Picking a TTL is picking where you live on that line — there is no universal right answer, only the one your users will tolerate.
Video: Ep. 21: Cache Eviction Explained in Detail — Systemic Stack English
Compares FIFO, LRU, and LFU with workload trade-offs — exactly this subtopic.
The thundering herd
Here's the failure mode that pages people at 3 AM. A popular key expires (or the cache restarts). A thousand requests arrive at once, all miss, and all stampede the database with the same expensive query. The database, which was happily idle a moment ago, falls over.
sequenceDiagram
participant R as 1000 Requests
participant Cache as Cache
participant DB as Database
R->>Cache: GET hot:key (expired!)
Cache-->>R: MISS × 1000
R->>DB: Same expensive query × 1000
Note over DB: Database falls over
Defenses, from simplest to fanciest:
- Request coalescing / single-flight — only the first request hits the database; the rest wait for its result. Many cache libraries do this automatically.
- Probabilistic early expiration — refresh keys before they expire, with a little randomness so they don't all refresh at once.
- Stale-while-revalidate — serve the slightly-old value immediately and refresh in the background. Users get speed and freshness, eventually.
Video: Why Expired Cache Keys CRASH Databases | Interview Question #42 — Binary Dose
One expired TTL and 10,000 requests rush the DB at once — like a concert crowd squeezing through a single door.
Where caches live
Caching isn't one layer — it's a stack. Browser caches, CDN edge caches (next lesson!), application caches like Redis, database query caches, even CPU caches. Each layer trades a different amount of staleness for a different amount of speed. When debugging "why is this slow," walk the stack top-down: the fix is usually a missing layer, not a faster database.
Video: Caching at Different Levels: Client, CDN, Redis & DB | System Design — InfraWithDipankar
Walks the four layers — client, CDN, Redis, database — and when each one goes wrong.
Takeaways
- Cache-aside is the default starting pattern: lazy, simple, and your app owns invalidation.
- Write-through keeps the cache fresh at the cost of write latency; write-behind is fast but can lose data.
- TTLs are the most reliable invalidation strategy — dumb beats clever.
- Plan for the thundering herd before your hottest key expires: coalesce, pre-refresh, or serve stale.
Check your understanding
In the cache-aside pattern, what happens on a cache miss?
- The request is dropped and the client retries with a cache-bypass header
- The cache serves a placeholder while a background job loads the row
- The application loads the row from the database, SETs it into the cache, then serves it
- The cache itself fetches the row from the database and returns it
What is the main risk of the write-behind strategy?
- Unflushed writes are lost if the cache fails before the database is updated
- Reads slow down because every read must verify the cache against the database
- The database receives every write twice, once from the cache and once from the app
- Hit ratios collapse because write-behind entries are never allowed to expire
A 'thundering herd' (cache stampede) happens when...
- Too many cache nodes join the cluster in the same minute
- Eviction deletes keys faster than the application can rewrite them
- Two regional caches disagree and clients see conflicting values
- A popular key expires and hundreds of requests hit the database with the same query at once
Which of these is a defense against the thundering herd, as described in the lesson?
- Raising every TTL to the maximum the cache allows
- Coalescing requests: only the first request queries the database while the others wait
- Turning the cache off entirely during traffic spikes
- Switching the eviction policy from LRU to LFU
Go deeper
Want to keep pulling this thread? These talks and tutorials go further than we did here:
- Caching in System Design: Redis, Strategies, Clustering & Real-World Patterns — Bagherani, YouTube tutorial (~27 min). Cache-aside, write-through, write-behind, and stampede control in Redis.
- Caching Explained — Cache Hits, TTL, Eviction & Redis — ByteScaler (formerly System Design From Zero), YouTube tutorial (~8 min). Hits, misses, TTL, and LRU — the essentials in eight minutes.
- Caching Explained: Client-Side, CDN, and Server-Side — System Design HLD, YouTube tutorial. Multi-layer caching and the freshness-versus-invalidation trade-off.
- Redis Course – In-Memory Database Tutorial — freeCodeCamp.org (~1.5h). Data structures, transactions, Pub/Sub, Lua scripting, security, benchmarking — the toolkit behind the strategies.
- System Design Course – APIs, Databases, Caching, CDNs, Load Balancing & Production Infra — freeCodeCamp.org (~2h 05m). The caching chapter covers where caches sit, eviction, and invalidation trade-offs.
Sources & further reading
- Martin Kleppmann, Designing Data-Intensive Applications (O'Reilly, 2017), Ch. 12 — "A cache often contains an aggregation of data" and caches as derived data.
- Redis documentation, "Eviction policies" — LRU/LFU variants in practice.
- Marc Brooker (AWS), "DynamoDB and caching" / AWS Architecture Blog posts on request coalescing and DAX.
- Rajesh Nishtala et al., "Scaling Memcache at Facebook" (NSDI 2013) — cache consistency and thundering-herd mitigation at massive scale.