Design Instagram: A Photo-Sharing System

The capstone case study: upload pipelines, sharded metadata, CDNs, and the fan-out problem behind a feed that serves a billion photos.

Advanced · 20 min read

Why this matters

This is the capstone — the lesson where everything in the course finally shares one roof. It's also the single most-asked system design interview question in the industry: design Instagram. Interviewers love it because there is no single right answer, only a series of good decisions under pressure. The trick isn't knowing every technology; it's showing a clean thought process: scope the problem, do the math, design the write path, design the read path, and name the hard part before they do. That's exactly how this lesson is structured — think of it as a rehearsal with the answers included.

Video: Designing INSTAGRAM: System Design of News Feed — Gaurav Sen
Builds Instagram from ER diagram to celebrity fan-out so every design choice feels earned — like watching the machine get assembled before it's switched on.

Step one: clarify the scope

Every strong interview answer starts the same way: by narrowing the problem before touching a whiteboard. Say this out loud:

Video: Design Requirements: The First Step of Every System Design Interview — AlgoJS
Covers functional vs non-functional requirements — the scoping questions to ask before designing.

Back-of-the-envelope math

Interviewers watch your arithmetic the way driving examiners watch your mirrors. State your assumptions, keep the math simple, and round aggressively:

The math has already told you the shape of the system: cheap durable byte storage for writes, a CDN absorbing nearly all reads, and a feed path engineered for sub-100ms responses. Summarized as the capacity plan you'd sketch on the whiteboard:

ResourceBack-of-envelopeDrives the decision
Photo bytes written~50TB/day (100M × 500KB), more with variantsObject storage, never the DB
Photo bytes storedTens of petabytes/yearMulti-region replication, lifecycle tiers
Photo bytes read~5PB/day at 100:1 read:writeCDN absorbs nearly all of it
Feed readsOne cache lookup per app openTimeline cache, not computation

Now let's build it.

Video: System Design Interview Course • Back of the Envelop Calculation aka Capacity Estimates — Anubhav Sethi
Drills the numbers worth memorizing and the four estimation formulas, then works them through URL-shortener and Google-Drive examples — times tables for scale.

The upload pipeline: the photo lab

This lesson's running analogy is an old-school photo printing lab. You drop a roll of film at the counter, the clerk hands you a claim ticket immediately, and you leave. Behind the counter, the real work happens out of sight: the negatives go into the archive vault, the lab prints your photos in standard sizes, and the order gets logged in the ledger. The counter never makes you wait for the developing.

Map it onto the system and the design almost draws itself:

Why async? Resizing images burns CPU. Doing it synchronously makes every upload slow, couples your API to your image pipeline's worst day, and turns one failed thumbnail into a failed upload. Async means the user gets their claim ticket in milliseconds, and a failed thumbnail job just retries invisibly.

Interactive diagram: PacketFlow (loads in the app)

Watch one photo's journey: the phone hands it to the API, the API drops the bytes into object storage and returns immediately — then the photo fans out to the thumbnail workers, the metadata writer, and CDN prefetch, all without the user waiting.

flowchart LR
    Phone["Phone app"]:::client --> API["Upload API<br/>the counter"]:::service
    API --> Store[("Object store<br/>the vault")]:::data
    Store --> Lab["Thumbnail workers<br/>the lab"]:::service
    Lab --> Store
    API --> Ledger[("Metadata DB<br/>the ledger, sharded")]:::data

Video: Design a Video Streaming Platform (Netflix/YouTube) — System Design Interview — Phani Thaticharla
Details the upload and transcoding pipeline: resumable uploads, workers, and async processing.

Metadata: the ledger, sharded

Here's a decision the math forces on you: photo bytes and photo metadata get different storage, because they have different shapes and different query patterns. The bytes are large, immutable, and fetched by ID — object storage. The metadata is tiny, structured, and queried by owner and time ("show me Maya's latest 20 photos") — a database.

But one database won't hold the metadata for billions of photos, so you shard it — the Sharding Strategies lesson, applied. And here you face a real fork in the road, because the shard key is a trade-off with no painless answer:

Shard by user ID, and every photo a person uploads lands on the same slice of the ledger. "Show me Maya's latest 20 photos" is a single-shard lookup — the system's most common read stays cheap. The cost is the celebrity: all of their uploads land on one shard, and every fan's profile view hammers it until that shard runs like a furnace. Real systems survive this by splitting the hot user — celebrity sharding, where one lopsided account is spread across several shards — or by keeping a secondary index.

Shard by photo ID instead, and writes spread beautifully: a celebrity's uploads scatter across every shard, and no single shard sweats. The cost is on the read side — Maya's photos now live on all shards, so "show me her latest 20" fans out across the fleet in a scatter-gather. Unless you build a secondary user→photos index so the router knows where her photos live — for the sharp reader who just asked: yes, that index exists in real systems, and yes, keeping it consistent is a cost of its own.

This isn't hypothetical: Instagram's engineering team has written about running sharded PostgreSQL in production, with photo IDs generated to encode the shard they live on. The pattern is the point, not the vendor: the ledger must be partitionable by the key you query by, or your most common read becomes a scatter-gather across every shard.

Video: What is Database Sharding? — Anton Putra
A senior engineer explains sharding like splitting a phone book across offices — same data, smarter routing — and how to pick a shard key so one office doesn't drown.

Delivery: the prints go to the corner store

When a user scrolls their feed, the actual photo bytes don't come from your servers — they come from a CDN, the way a photo lab distributes finished prints to corner stores instead of making every customer drive to the central lab. The phone app gets signed CDN URLs from the feed response and fetches bytes from the nearest edge PoP. Popular photos hit 95%+ edge cache ratios (you did this math in the CDNs lesson); the origin object store only serves the first fetch per region. That 5PB/day of read bandwidth? The CDN absorbs nearly all of it, and your origin bill stays sane.

Video: Amazon S3 Explained Buckets, Objects & Permissions — Hired in IT (unverified — confirm channel on YouTube)
Teaches object storage like renting storage units — buckets, keys, and who gets a key — including the permission mistakes that bite, plus S3-plus-CloudFront serving.

The feed: the newspaper routes

Now the hard part — and the part interviewers are really testing. Extend the analogy: the feed is a newspaper delivery route. There are two ways to get the morning paper to subscribers.

Fan-out on write (push) is the paperboy: when Maya posts a photo, the feed service stuffs her photo ID into the pre-built home timeline of every follower. Reading a feed is then trivially cheap — one lookup of your own timeline. But the write cost scales with follower count: a post to 1,000 followers is 1,000 timeline writes.

Fan-in on read (pull) is the newsstand: timelines are assembled at request time by gathering recent posts from everyone you follow and merge-sorting them. Writing a post costs almost nothing; reading costs scale with how many people you follow.

And then there's the celebrity problem — the interviewer's favorite trap, sometimes called the Justin Bieber problem. A user with 100 million followers posts a photo. Under pure push, that's 100 million timeline writes for a single post — write amplification so extreme it can melt the feed service. Under pure pull, every one of those 100 million followers pays merge-sort cost on every feed load instead.

The production answer is a hybrid: push for ordinary users (where fan-out is cheap), and treat high-follower accounts specially — their posts are fetched at read time and merged into the pre-built timeline. You pay the pull cost only for the tiny fraction of accounts that would break push, and everyone else gets the cheap pre-built read. Twitter's timeline architecture famously works this way, and the reasoning transfers directly.

Interactive diagram: StepThrough (loads in the app)

flowchart LR
    Phone["Phone app"]:::client --> Feed["Feed service"]:::service
    Feed --> Cache[("Timeline cache<br/>pre-built photo IDs")]:::data
    Feed --> Star[("Celebrity posts<br/>fetched live")]:::data
    Feed --> Phone
    Phone --> CDN["CDN<br/>serves the bytes"]:::cloud

Notice the separation the whole design keeps returning to: the feed service deals in IDs (tiny, cacheable, mergeable), while the bytes flow through the CDN. That split — metadata path vs byte path — is the single most important architectural line in the system. Every time the design gets confusing, ask "is this about IDs or about bytes?" and it gets clear again.

Video: System Design 101: News Feed Architecture — Ptolémé
Compares fan-out on write vs read for feed generation, caching, and ranking tradeoffs.

The follow graph: the other database

Photos aren't the only thing that needs sharding — the follow graph does too. Follows are edges (follower → followee), and at Instagram scale there are hundreds of billions of them. They live in a sharded store of their own, because the query patterns differ from photo metadata: "who does Maya follow?" (fan-out reads walk out from a user) and "does Alex follow Maya?" (a single indexed lookup on a profile view) shard naturally by follower, while "who follows the celebrity?" shards by followee and gets paginated — nobody renders a 100M-row follower list in one response.

The follow graph is also where the celebrity problem quietly lives a second life: the feed service needs the follower list to do fan-out on write, so that read must be fast and paginated, not a full table scan wearing a nice query. In the interview, mentioning the follow graph as a separate sharded concern — distinct from both photo bytes and photo metadata — is what separates a good answer from a great one.

Video: Neo4j (Graphdb) vs SQL (RDBMS) | Graph Database Explained for SQL Developers — Learn IT Tech
Maps SQL tables and JOINs onto graph nodes and relationships so the "why graphs" moment clicks — like swapping a filing cabinet for a whiteboard of connections.

Serving the right size: variants and signed URLs

One more detail the lab analogy covers nicely: the lab prints your photo in several standard sizes, and the counter hands you the right one for the frame you bought. The feed response doesn't contain photo bytes — it contains a set of CDN URLs, one per variant (thumbnail, feed-size, full-resolution). The phone app picks the variant for the screen it's rendering: a 100-pixel avatar never downloads the 4K original.

For private accounts, those URLs are signed — they carry an expiry timestamp and a signature, so a leaked URL stops working after an hour instead of living on the internet forever. The CDN verifies the signature at the edge without ever calling your API. Access control, enforced where the bytes actually flow.

Video: System Design Interview: Design YouTube w/ a Ex-Meta Staff Engineer — Hello Interview
An ex-Meta staff engineer whiteboards YouTube end to end — like watching a senior think out loud through every trade-off from upload to playback.

Edge cases: what breaks at a billion photos

The design above is the happy path. The interview points come from naming what breaks:

Video: Design the News Feed (Twitter / X) — System Design Interview — The Hot Path
Starts where the real difficulty lives — the celebrity hot-key that breaks naive fan-out — then shows the hybrid and failover fixes, like a post-mortem before the outage.

What you'd say in the interview

Walk it back in one minute: scope to upload, feed, follow; 100M uploads/day means tens of petabytes a year in object storage, so bytes never touch the database; metadata goes to a sharded DB keyed by the query pattern; the CDN absorbs the 100:1 read amplification; thumbnails are async so uploads feel instant; and the feed is a hybrid — fan-out on write for normal users, fan-in on read for celebrities, timelines cached as ID lists. Then name the trade-off you didn't take: pure push melts on celebrities, pure pull makes every read expensive, and your hybrid is the deliberate middle. That's a passing answer.

Video: Design FB News Feed System Design Interview w/ ex: Meta Senior Manager — Hello Interview (unverified — confirm channel on YouTube)
Full interview-framework walkthrough, requirements to deep dives, with an ex-Meta senior manager — a dress rehearsal for the real round.

Takeaways

  1. Scope first, then math, then design. The back-of-the-envelope (100M uploads/day, ~50TB/day of photo bytes, 100:1 read ratio) dictates the architecture before a single component is chosen.
  2. Bytes and metadata get different storage. Immutable blobs go to object storage; small, queryable records go to a sharded database partitioned by your query key.
  3. Uploads return fast because the slow work is async. The API stores bytes and hands back an ID; thumbnail workers process the queue behind the counter.
  4. The feed is a fan-out problem. Push (fan-out on write) makes reads cheap but breaks on celebrities; pull (fan-in on read) makes writes cheap but reads expensive. The hybrid — push for most, pull for celebrities, timelines cached as ID lists with bytes on the CDN — is the production answer.

Check your understanding

  1. Why does the design store photo metadata in a sharded database instead of alongside the photo bytes in object storage?

    • Object storage caps out at a few terabytes, so the metadata needs a database of its own simply to fit the billions of rows
    • Databases are always faster than object storage for every kind of read, so moving metadata into a database speeds up the whole system
    • Metadata is small, structured, and queried by owner and time; photo bytes are big blobs fetched by ID — each gets storage shaped for its own pattern
    • A sharded database automatically generates the thumbnail variants from the photo bytes, a job object storage was never designed to handle
  2. Using the lesson's back-of-the-envelope math — 100M uploads/day at ~500KB per photo — roughly how much new photo storage is needed per day, before thumbnails?

    • About 5TB per day
    • About 50TB per day
    • About 500TB per day
    • About 5PB per day
  3. Under pure fan-out on write, a celebrity with 100 million followers posts one photo. What happens?

    • Roughly 100M timeline writes for a single post — write amplification that can overwhelm the feed service
    • Each follower's phone downloads the photo 100M times, so the CDN edge caches melt under the repeat traffic
    • The object store must replicate the photo across 100M regions before the post can go live, stalling every upload
    • Nothing special — the CDN edge cache absorbs the extra writes, so the feed service never feels them
  4. Why are thumbnails generated asynchronously, after the upload API has already responded?

    • Thumbnail workers only run during off-peak hours, so they process the queue while the user is already gone
    • The object store only keeps one size per photo, so variants are generated just-in-time instead of at upload
    • The metadata database has to commit the ledger row before any image processing is allowed to begin
    • So uploads feel instant instead of blocking on CPU-heavy image processing — and a failed thumbnail job just retries invisibly

Go deeper

Want to keep pulling this thread? These talks and tutorials go further than we did here:

Sources & further reading