M2 Design Instagram Feed
Loading learning experience...
Lecture transcript
Read the narration for M2: Design Instagram Feed
Design Instagram Feed
Welcome everyone, today we will design an Instagram style home feed, exploring hybrid fan out, caching, ranking, and how to keep P ninety nine latency low at celebrity scale.
From estimation to architecture: why the feed is a numbers game
Dr. Wei: Today we are transitioning from estimation to architecture. The main idea is that a feed only feels simple until you put real scale and latency constraints on it.
Sam: So we are going to treat the feed like a throughput and latency problem, not just a feature question, right?
Dr. Wei: In the last lesson, we used quick estimation to turn user activity into load. That kind of back of the envelope math is what keeps our design grounded in reality.
Dr. Wei: Our driving question for this lecture is: given Instagram scale, how do we design a home feed that stays fast and reliable, even when data is changing and not everything is perfectly up to date everywhere?
The core tension: fast reads vs explosive writes
Dr. Wei: In an Instagram-style feed, the job sounds simple: show people the newest posts from accounts they follow. But at scale there is a central tension you can’t ignore: you want reads to feel instant, while writes can suddenly trigger massive work across the system.
Sam: My default instinct is to precompute on write so reads are fast, but I guess the celebrity case makes that blow up quickly.
Dr. Wei: Most of the time, the product is read-heavy: people open the app and scroll far more often than they create content. So if we optimize for the common case, we’d love to make feed reads cheap, predictable, and cacheable.
Dr. Wei: But the painful case is the write path: when a very popular account posts, the system may need to fan that post out to an enormous number of followers. That single write can behave like millions of operations, and it can overwhelm your storage, queues, or caches if you do it the naive way.
Back-of-envelope scale: justify caching and async fan-out
Dr. Wei: Before we pick a feed architecture, we need a quick back-of-the-envelope sense of scale: how many reads per second and how many writes per second we should expect, because that ratio is what drives choices like caching and asynchronous fan-out.
Dr. Wei: Sam, based on your gut check, what order of magnitude do you expect for feed read requests per second: ten thousand, one hundred thousand, or one million? And for post writes per second: one thousand, ten thousand, or one hundred thousand?
Sam: I would guess reads are around one hundred thousand per second, maybe higher at peaks, and writes feel like a few thousand per second, not tens of thousands.
Dr. Wei: Let’s reveal the math. If we assume about two billion users and roughly ten feed opens per user per day, that is about twenty billion feed opens per day. Divide by eighty six thousand four hundred seconds in a day, and we get about two hundred thirty thousand read requests per second. For writes, if there are about five hundred million posts per day, dividing by eighty six thousand four hundred gives about five thousand eight hundred writes per second.
Dr. Wei: Now compare that to your prediction: the reads land closer to one hundred thousand per second than one million, and the writes are only in the few-thousand-per-second range. That big read-heavy gap is exactly why we lean on caching for reads and use asynchronous fan-out so writes do not have to synchronously update everyone’s feed at request time.
API surface: what clients actually need
Dr. Wei: When we design an Instagram style feed, one of the quickest ways to clarify the system is to name the small set of client actions we must support. Think from the phone app’s point of view: what does it need to do, and what does it need back, to keep the experience smooth?
Dr. Wei: At a minimum, clients need a way to create a post, a way to fetch the next page of the feed for infinite scrolling, and a way to send lightweight engagement events like likes. Keeping the surface area small makes it easier to evolve the backend without breaking clients.
Sam: For the feed read, do we return full post objects, or do we return mostly identifiers plus a cursor and let the client fetch details separately?
Dr. Wei: As you read these endpoints, focus on what each call returns to the app: for posting, an identifier and status; for reading, a page of items plus a cursor for the next request; and for engagement, an acknowledgment that can be recorded asynchronously. This is the contract the rest of our design must satisfy.
Data model: separate truth from derived views
Dr. Wei: Before we talk about ranking or performance, we need a clean data model for the Instagram feed. The key idea is to separate what is always true in the system from what we compute as a convenience. That separation keeps the design simpler and makes it easier to scale and to fix bugs without corrupting the source of truth.
Sam: So posts and the follow graph are the durable truth, and the feed itself is more like an index we can rebuild if it gets out of sync?
Dr. Wei: We will start by listing the core entities that represent truth: users, posts, and the follow relationships. These are the records we would store durably and update carefully, because other features depend on them. If we get these right, the rest of the system can be built as derived views on top.
Dr. Wei: Then we will define a derived view for the feed: a cached, ordered list of post identifiers per user. This view is allowed to be rebuilt, refreshed, or partially stale, because it is not the truth, it is an optimization for read speed. Thinking this way helps us decide what must be strongly consistent versus what can be eventually consistent.
Fan-out choices: push, pull, or hybrid
Dr. Wei: When we design an Instagram-style home feed, one of the biggest architectural choices is where we pay the cost: at write time, at read time, or split it across both. This choice affects latency, storage, and how the system behaves during spikes.
Sam: If we push on write, reads are simple but the write amplification seems scary. If we pull on read, we risk high tail latency. So hybrid feels like the obvious compromise.
Dr. Wei: Before we pick a hybrid cutoff, Sam, predict a plausible follower threshold where we stop pushing and start pulling: closer to 1 thousand, 10 thousand, or 1 million? Now justify it with a quick capacity check: if our fan-out workers can handle about 100 thousand feed inserts per second total, then a single post from an account with F followers costs about F inserts; so at 10 thousand followers, that is about one tenth of the fleet for a second, but at 1 million followers it is about ten seconds worth of full-fleet work from one post.
Dr. Wei: In practice, teams often use a hybrid rule: push for typical accounts, but switch to pull for very large accounts so a single celebrity post does not overwhelm the system. A cutoff around 10 thousand followers is one reasonable knob, but the real point is that it should come from your capacity math and the spike behavior you can tolerate.
High-level architecture: write path, read path, and media
Dr. Wei: This slide is one high-level flow of the system end to end: write path, read path, and how media bytes get delivered.
Sam: Just to check my understanding: the feed service mostly deals with lightweight metadata, and the heavy media bytes are handled by object storage plus the CDN, right?
Dr. Wei: On the write path, a post goes to the Post Service, gets queued, fan-out workers figure out which followers should see it, and they update each user’s feed cache so reads are fast later.
Dr. Wei: On the read path, the Feed Service mostly serves from the feed cache, but it can also pull in extra candidates for very large authors, like celebrities, where pushing to every follower would be too expensive.
Dr. Wei: Finally, treat media separately from feed metadata: images and videos land in object storage and are served through a CDN, so the heavy bytes come from the edge while the feed services return lightweight IDs and ordering.
Ranking pipeline: from $\sim 10^3$ candidates to a page of $20$
Dr. Wei: Chronological is easy; engagement ranking is the product—and the trick is doing it fast enough to hit a P99 under 500 milliseconds.
Sam: So the real question is how we turn a big pile of possible posts into a small, high-quality page without blowing our latency budget.
Dr. Wei: Right—and that tail-latency constraint is exactly why we split the work into stages and give each stage a budget, instead of doing one giant expensive ranking step.
Dr. Wei: We can think of a simple budget like: about 50 milliseconds to fetch candidates, 150 milliseconds to compute features, 100 milliseconds to score with the model, 50 milliseconds for re-ranking and post-processing, and around 100 milliseconds for network and overhead—adding up to under 500 milliseconds at P99.
Fan-out workers: make the write path safe and boring
Dr. Wei: Now let’s talk about the write path for an Instagram feed. The goal is not to make it clever; it’s to make it safe, predictable, and easy to operate under spikes. When someone posts, we want a design that can absorb huge bursts, recover cleanly after failures, and ensure that duplicate deliveries do not cause duplicate feed entries.
Sam: When you say duplicate deliveries are expected, do you mean we need an idempotency or dedup key per user and post so retries do not create multiple feed entries?
Dr. Wei: A common approach is to separate the user-facing request from the heavy lifting: accept the post quickly, then hand off fan-out work to background workers. That way, if the work takes longer than usual, the user still gets a fast response and the system can retry work safely in the background.
Dr. Wei: In interviews, the probing questions are usually about failure modes: what happens when workers crash mid-job, messages are delivered twice, or a celebrity posts to millions of followers. The key is to design the pipeline so duplicate deliveries are normal but duplicate effects are prevented, for example by using a dedup key like the pair u id and p id on the feed write, and so the biggest accounts don’t overload the write path.
Freshness, invalidation, and pagination under churn
Dr. Wei: In an Instagram-style feed, the hard part is that content changes while the user is in the middle of scrolling. People publish and delete posts, follow and unfollow, and the system is constantly updating ranks. On this slide, we focus on keeping cached feed results fresh under churn, especially the tradeoff between event-driven invalidation and time-based expiration.
Sam: So instead of only waiting for TTL to expire, we can invalidate when a follow or a new post event happens, right?
Dr. Wei: Exactly. There are two main approaches. With event-driven invalidation, new posts, deletes, or follow changes trigger updates quickly, but it adds complexity and can create a high fanout of cache updates. With TTL-based expiration, the system is simpler and bounded in work, but it can serve stale results until the timer runs out, especially during bursts of posting.
Dr. Wei: A concrete failure to watch for with TTL-only is freshness lag: a creator posts something, but followers do not see it for minutes because their feed page is still cached. Product-wise, that feels broken, even though the system is behaving as designed. The usual answer is a hybrid: keep a TTL as a safety net, but also invalidate or patch results for high-impact events.
Failures and product lens: connect knobs to metrics
Dr. Wei: Let’s end by connecting failure handling to product outcomes: what knobs we can turn, and what metrics those knobs protect.
Sam: When we pick a knob like the celebrity threshold, are we basically trading write amplification against read latency and freshness for followers of big accounts?
Dr. Wei: First, reliability: if parts of fan-out fall behind, we can allow a slightly stale feed within a clear service level agreement, like under thirty seconds, so we preserve availability and a predictable user experience.
Dr. Wei: Second, product iteration: we run ranking changes through an experiment framework and evaluate them with both offline metrics and online A B tests, ideally without touching the write path, so we can ship safely and learn quickly. Finally, when a metric drops, like Stories engagement down five percent, we debug by tracing the pipeline from candidate generation to features to the model to find the most likely cause.
Exit ticket: quantify the celebrity problem, then choose a knob
Dr. Wei: Exit ticket time: you’ll do one quick back-of-the-envelope estimate, then one design reflection about which knob you would turn.
Dr. Wei: For the estimate, imagine a celebrity post that would be pushed to F equals six times ten to the eighth followers, and assume each per-follower write is about one hundred twenty eight bytes of payload, like a post id and timestamp plus a little structure.
Sam: So payload-only is six times ten to the eighth times one hundred twenty eight bytes, which is on the order of tens of gigabytes, and then the overhead factor makes it much worse.
Dr. Wei: Compute F times b to get the payload-only bytes written, then multiply by an overhead factor of ten to reflect realistic storage encoding and metadata, and notice how quickly this becomes a massive write spike.
Dr. Wei: Then for the reflection: pick one knob, like a celebrity threshold, k, or tau, and say which metric you expect to improve or worsen, for example write amplification, tail latency, cache hit rate, or feed freshness.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!