M5 Tradeoff Deep Dive
Loading learning experience...
Lecture transcript
Read the narration for M5: Tradeoff Deep Dive
From $IC6$ scoping to $IC6$ tradeoffs
Dr. Wei: Today we are turning scoping skills into tradeoff skills: the kind of reasoning expected at the IC6 level. By scoping, I mean you can clearly state the goal and constraints, list two or three plausible approaches, and attach order-of-magnitude numbers so the discussion stays grounded. We will use that same structure to compare options and make a decision you can defend.
Sam: When you say IC6 tradeoffs, what exactly are we expected to produce in a short discussion?
Dr. Wei: A fast, defensible answer: name the key metric like tail latency or cost, state the main risks to correctness, and do quick back-of-the-envelope math to justify your choice. The goal is to make the reasoning legible, not to be perfectly precise. By the end, you should be able to walk through a tradeoff in about two minutes and explain why your decision is reasonable.
The tradeoff articulation framework
Dr. Wei: Before we argue about a design, we need a shared way to talk about tradeoffs that is concrete, comparable, and falsifiable.
Dr. Wei: Sam, quick checkpoint: for a familiar case like adding a cache in front of a database, give me one sentence using the template with at least two numbers. For example, a p ninety-nine latency change and a cost or risk number.
Sam: We choose a cache because p99 read latency drops from 120 milliseconds to 20 milliseconds at 5,000 queries per second, and we accept about 2 gigabytes of extra memory; we mitigate stale reads by setting a 30 second time to live and monitoring a one percent staleness error budget.
Dr. Wei: Good. You picked one primary metric, p ninety-nine latency, and you named real costs and risks. Now we can critique it: are those numbers measured for this workload, is 2 gigabytes per node acceptable, and is the one percent staleness budget actually enforced with an alert and a rollback plan?
Consistency vs availability: CAP in practice
Dr. Wei: When you build a distributed product, you are constantly trading off between consistency and availability. CAP is specifically about what you choose to preserve when the network partitions. Instead of treating the CAP idea as a slogan, we will focus on what it means for user experience and operational cost.
Sam: My default answer is usually, "pick availability for user-facing stuff." But that still feels hand-wavy. How do I make it defensible?
Dr. Wei: Start with the practical framing: during partitions, many products favor availability, and then accept eventual consistency with a staleness SLO, like a maximum of five seconds, when the product can tolerate it. Then quantify the bound and the payoff. On the Instagram likes example, ask: if we allow about five seconds of staleness, how much coordination or invalidation traffic can we avoid, and what is the user-visible downside?
Sam: So to quantify it, I could say something like: we cap staleness at five seconds, and in return we cut invalidation or coordination traffic by about an order of magnitude, while the user downside is only a briefly stale like count.
Latency vs throughput: batch, real-time, and micro-batch
Dr. Wei: Today we are unpacking a classic systems tradeoff: how fast each individual result arrives versus how much total work the system can push through over time.
Sam: Prediction check: if we process one million items in ten minutes, what throughput is that? I think it is about one thousand items per second.
Dr. Wei: Let’s compute it using the throughput line: items over time. Ten minutes is six hundred seconds, so one million divided by six hundred is about one thousand six hundred sixty seven items per second. Your estimate was in the right ballpark, but in interviews I would say, "about one point seven thousand per second," and call out the seconds conversion.
Storage vs computation: precompute vs compute on demand
Dr. Wei: When we choose between precomputing and computing on demand, we are trading storage and write work for faster reads later.
Sam: Prediction check: for the feed numbers, reads are about 230 thousand QPS and writes are about 5.8 thousand QPS. R over W is maybe around 400, so precompute is definitely worth it?
Dr. Wei: Good instinct to compute the ratio, but the arithmetic is off. Look at the feed line: two hundred thirty thousand divided by five point eight thousand is about forty, not four hundred. Forty is still comfortably above the "ten to twenty" guideline, so caching or precomputing often wins, but you must also name the costs on the next bullet: staleness, terabytes of cache, and invalidation complexity.
Accuracy vs latency: bigger models, slower inference
Dr. Wei: When we talk about model quality, it is tempting to chase every last improvement in accuracy or engagement. But in production systems, those gains only matter if users still get a fast, responsive experience. This slide is about that core tension: higher accuracy often comes with higher inference latency.
Sam: So we are not just optimizing accuracy, we are also optimizing how quickly the system responds.
Dr. Wei: A small relative lift, like a plus two percent engagement gain, can be extremely valuable at scale. The catch is that moving to a heavier model can add milliseconds, and that can push you over a latency budget for ranking, search, or recommendations.
Sam: Is this the kind of trade where five milliseconds for a better model might still be too slow compared with a sub-millisecond baseline?
Dr. Wei: Also remember that user experience is often driven by tail latency, not the average. A quick rule of thumb: if one request fans out to about ten downstream calls, and each service has a p ninety-nine of five milliseconds, the end-to-end p ninety-nine is often much higher than five milliseconds because you are effectively waiting for the slowest one. That is why teams may need to budget something like two to three milliseconds per hop, or reduce fanout, to keep the overall tail under control.
Sam: So even if every component looks fine in isolation, combining them can make the tail latency jump because the slowest hop wins.
Dr. Wei: One practical pattern is a cascade: start with a fast model that filters or prunes, then spend the expensive model only where it matters, like reranking a small top k set. That way you keep most of the accuracy benefit while protecting end to end latency.
Sam: That makes sense: pay the expensive latency only on a small shortlist, instead of on every single item.
Cost vs performance: when two times cost is worth ten times speed
Dr. Wei: Performance work is only a good bet when it creates clear business value, not just faster graphs. On this slide, we will translate speed improvements into dollars saved or dollars earned, so we can judge whether a higher cost option is actually the better decision.
Sam: Prediction check: if the cache costs two million dollars per year and it avoids five million dollars per year of database capacity, the net is plus two million per year?
Dr. Wei: Close, but check the subtraction. On the net line, it is negative two million plus positive five million, which equals positive three million per year. The key is you are comparing deltas in the same units, then sanity-checking that the speedup is actually worth pursuing because returns diminish as you get very fast.
Push vs pull fan-out: avoid write or read amplification
Dr. Wei: When you build a social feed, you quickly run into a fan-out decision: do you do the work when someone posts, or when someone reads? This is a classic scale tradeoff where you are choosing which side of the system absorbs the cost.
Sam: I would normally just say, "push is better because reads are faster." But I guess that ignores the cost side. Is that what you are looking for?
Dr. Wei: Yes. Anchor to the first two bullets: push gives fast reads but can create huge write amplification for celebrities; pull keeps writes cheap but creates read amplification at feed load time. Then the third bullet is your mitigation: a hybrid rule that caps worst cases, like pushing only for users below a follower threshold and pulling for the tiny top fraction.
Dr. Wei: At large scale, teams often land on a hybrid: push for most users where fan-out is manageable, and pull for the tiny fraction of accounts with massive follower counts. The goal is not to eliminate amplification, but to cap the worst-case behavior so neither writes nor reads explode for your hottest users.
Exit ticket: articulate one tradeoff with numbers
Dr. Wei: For this exit ticket, I want you to state a real tradeoff in one sentence, and include at least one number so your reasoning is concrete.
Sam: Okay: "Choose availability for read receipts because it is faster, accept some inconsistency, and mitigate with retries." That still feels vague though. What number would you want?
Dr. Wei: Make sure the number shows up in at least one part of the sentence: a latency target, an error rate, a cost, a time window, or a freshness bound like one second. For example: "Choose availability with bounded staleness of one second because it keeps reads responsive; accept that some users may briefly see the wrong state; mitigate with idempotent events and conflict resolution." Keep it short, but specific enough that someone could disagree with the assumption.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!