M1 Estimation for Meta
Loading learning experience...
Lecture transcript
Read the narration for M1: Estimation for Meta
Why estimation is the $\mathrm{SEDE}$ hinge
Dr. Wei: In real interviews and real systems, estimation is where you stop hand waving and start making decisions. Today we are going to build the habit of converting a vague statement like Instagram is huge into queries per second, storage per day, and bandwidth, and then using those numbers to justify the architecture. The promise is simple: we will only use a few reference numbers and unit conversions, and we will follow the same Scope, Estimate, Design, Evaluate loop you saw last time.
Sam: So the point is not to be exact, but to be specific enough that the design basically picks itself?
Dr. Wei: Exactly. When you say a number out loud, you are making a claim the interviewer can test. If your numbers imply two hundred thousand reads per second, a cache is not optional, it is mandatory.
Sam: And if my estimate implies only a few thousand writes per second, that pushes me toward sharding for storage growth rather than sharding purely for write throughput, right?
What interviewers want from your numbers
Dr. Wei: At Meta style interviews, you are not graded on the exact digit. You are graded on whether your assumptions are explicit, your units line up, and your conclusion changes the design.
Dr. Wei: When your assumptions are reasonable, they give you permission to add real components: caches, queues, sharding, precompute. Without numbers, those are just buzzwords.
Sam: I sometimes get stuck on units. Is there a quick way to catch mistakes early?
Dr. Wei: Yes: say the units out loud. Users times actions per day divided by seconds per day must become actions per second. Then sanity check against known limits, like a Memcache box doing about one hundred thousand ops per second or a database shard doing only a few thousand writes per second.
The cheat sheet of Meta-scale constants
Dr. Wei: To estimate quickly, you want a tiny mental table you can reuse across problems: user scale, time scale, and a simple size anchor for text.
Dr. Wei: On users: you can anchor around a couple billion daily active users for a major Meta app. That turns everyone uses it into a number you can multiply by actions per user per day.
Sam: So if I remember only one time constant, it is eighty six thousand four hundred seconds per day, right?
Dr. Wei: Exactly. And for a quick size anchor, one character is about one byte, which helps you sanity-check text payloads before you get into heavier media and bandwidth later.
From $\mathrm\{DAU\}$ to average and peak $\mathrm\{QPS\}$
Dr. Wei: This is the one equation you should reach for first: average queries per second is daily active users times actions per user per day, divided by seconds per day. It is just unit conversion, but it is the conversion that unlocks architecture.
Dr. Wei: Notice what this formula does: it takes a product question, like how often do users do this, and turns it into load on your backend per second.
Sam: And then we multiply by, say, three for peak because users are not evenly spread across the day?
Dr. Wei: Yes. Before we calculate, Avery, make a prediction: for two billion daily active users and ten likes per user per day, what do you think the average likes per second is, and what is a reasonable peak number if we use a factor of about three?
Sam: Roughly two hundred thousand likes per second on average, and maybe around six hundred thousand likes per second at peak with a factor of three.
Dr. Wei: Let us compute it. Two billion times ten is twenty billion likes per day. Divide by eighty six thousand four hundred seconds in a day and you get about two hundred thirty one thousand likes per second on average. With a peak factor of three, that is about six hundred ninety four thousand likes per second. If you were off, the most common slips are forgetting to divide by eighty six thousand four hundred, or misreading two billion as two million.
Example: Instagram feed is read-heavy
Dr. Wei: Let’s do a quick pressure-test estimation. Suppose Instagram has about two billion daily active users, and each person opens the feed about ten times per day. Before we compute it, what do you think the average feed opens per second is, order of magnitude?
Dr. Wei: Now do the same for writes. If the system sees about five hundred million posts per day, what does that feel like in writes per second, roughly?
Sam: Feed opens seem like hundreds of thousands per second, and writes seem like only a few thousand per second. That seems tiny compared to opens.
Dr. Wei: Exactly. If you run the conversion, two billion users times ten feed opens per day comes out to about two hundred thirty thousand feed opens per second on average, and five hundred million posts per day is about five point eight thousand writes per second. Now, one feed open is not one backend read: it can fan out into r backend reads for things like timeline fetch, ranking features, media, and ads. So backend read queries per second is roughly feed opens per second times r, often landing in the millions. The key is not the exact numbers; the key is the ratio and the fanout.
Dr. Wei: When read work outweighs write work by tens or hundreds of times, you precompute and cache the read path aggressively.
Storage estimation: count $\times$ size $\times$ retention
Dr. Wei: For storage, you want a simple multiplicative model: items per day times bytes per item times how long you keep it, then multiply by an overhead factor for replicas and indexes.
Dr. Wei: The overhead factor is where many candidates undercount. Real systems store multiple copies, plus secondary indexes, plus extra metadata for routing and encryption. Two to three times is a reasonable first pass.
Sam: So if I say just raw bytes, I should immediately say I will multiply by an overhead factor for the real footprint?
Dr. Wei: Yes, and you should also say whether you are doing short term sizing or multi year capacity planning. Designing a system that stores data for five years usually forces sharding and tiered storage, even if day one fits on a single cluster.
Example: Messenger storage blows up fast
Dr. Wei: Start with the classic Messenger estimate: one hundred billion messages per day, about one hundred bytes of text. That is about ten terabytes per day, even before we talk about attachments or metadata.
Dr. Wei: Now the real world correction: messages are not just text. You have sender and receiver ids, timestamps, delivery state, spam signals, maybe encryption envelopes. If you add around one kilobyte of metadata per message, you are quickly near one hundred terabytes per day.
Sam: And then five years means we are more like on the order of one hundred to two hundred petabytes, even before replicas. So a single database is not happening.
Dr. Wei: Exactly. This is how estimation drives architecture: you now need sharding, and you probably want tiered storage, where hot recent messages live on fast storage and older messages move to cheaper tiers. Also, you think about compaction and how indexes scale with time.
Read:write ratio is a design compass
Dr. Wei: Once you have rough reads and writes, compress it into one headline number: the read to write ratio. That ratio immediately tells you where to spend complexity.
Dr. Wei: For feeds, reads overwhelm writes, so you pay in write amplification to make reads fast. That is where caching, precompute, and even hybrid fanout come in.
Sam: Messaging is more balanced, so over optimizing reads could backfire because it slows down the write path too much?
Dr. Wei: Exactly. And for logging, writes dominate, so you design for sustained ingest: batching, sequential writes, and durable queues. Reads are often offline and can tolerate higher latency.
Sam: So in an interview, after I state the read to write ratio, I should immediately justify one tradeoff, like paying extra work on writes to make reads fast, or the opposite for logging?
Work backwards from $p_{99}$ latency to cache hit rate
Dr. Wei: A useful way to think about performance is to work backwards: start with a product latency goal, then infer what reliability you need from the fast path versus the slow path.
Dr. Wei: Now we plug in a typical target: you want under two hundred milliseconds at the ninety ninth percentile. Suppose cache hits are around two milliseconds, and a database miss is around fifty milliseconds. The key tail latency intuition is about the probability of taking the slow path: if misses are slow, then to keep p99 near the cache path you need misses to be rare enough that they almost never show up in the slowest one percent of requests.
Sam: So even a one percent miss rate could dominate the tail latency. That is why people say you need ninety nine percent plus hit rate?
Dr. Wei: Exactly. For p99, a clean rule of thumb is: keep the miss probability at or below one percent so a database miss is unlikely to land in the p99 bucket. That is a tail risk requirement about how often you route to the slow path, not something you can prove from an arithmetic mean. Separately, if you want a quick sanity check on overall performance, you can use a simple average latency approximation with hit rate weighting.
Dr. Wei: And now the architecture follows: warm caches, consistent hashing, careful time to live settings, and fast invalidation, because a one percent miss rate is not a small detail, it is your latency budget.
Exit ticket: do the conversion, then state the design consequence
Dr. Wei: For practice, do the same conversion we used all lecture: one billion daily active users times thirty actions per day is thirty billion actions per day. Divide by eighty six thousand four hundred seconds to get about three hundred forty seven thousand actions per second on average, then multiply by three for peak to get about one million actions per second.
Sam: So if I say peak is about one million queries per second, I should immediately talk about caching, sharding, and load shedding, not just leave it as math.
Dr. Wei: Exactly. For the reflection, pick one consequence: for example, one million per second means you cannot hit a single database on every request, so you add a cache layer and set a concrete hit rate goal. That is the habit interviewers are looking for: a number that forces a design move.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!