M4 Design Real-Time Ranking System
Loading learning experience...
Lecture transcript
Read the narration for M4: Design Real-Time Ranking System
Design Real-Time Ranking System
Welcome everyone, today we will explore how to design a real time ranking system that narrows 100 thousand candidates into a great list in 200 milliseconds.
From rate limiting to ranking: protect the pipeline
Dr. Wei: Before we talk about model quality, picture the real time ranking path: a request comes in, we generate candidates, fetch features, score them, and then rerank to produce the final list.
Sam: So even if the model is great, a traffic surge can still make the experience bad because the system cannot answer in time. Are we mainly trying to avoid timeouts and cascading failures?
Dr. Wei: Rate limiting and load shedding are how we bound that load to protect downstream dependencies like candidate generation calls, feature store lookups, and the ranking service itself, because if we cannot cap input pressure, none of the downstream latency budgets matter.
What real-time ranking is really doing
Dr. Wei: Ranking is where product quality and systems scale collide. In a real-time product, you are making a decision for every user, in every moment: what to show first, what to hide, and what to defer. That decision shapes engagement, revenue, and trust, all at once.
Sam: I usually picture ranking as sorting by a score, but it sounds like the core is making a decision under constraints. Is the main challenge that the decision has to be both fast and safe?
Dr. Wei: Over this lesson, we will treat ranking as an end-to-end system: data arrives continuously, features need to be fresh, models need to be served safely, and the output must be explainable enough to iterate on. The goal is to make the ranking decision feel instantaneous to the user, while still being correct, robust, and measurable.
Requirements and scale that drive the architecture
Dr. Wei: Before we pick any databases, queues, or services, we need to agree on what the system must do and what “real time” actually means for ranking. The architecture will be driven by concrete requirements and by the scale we expect to handle.
Sam: For calibration, if the end to end budget is around two hundred milliseconds, is it fair to say the ninety ninth percentile matters more than the average, since a few slow requests are what users feel?
Dr. Wei: Yes, and to keep that two hundred millisecond goal realistic, we need a worked scale scenario. For example, suppose peak traffic is ten thousand requests per second and we score one thousand candidates per request. That is ten million candidate scores per second. If scoring each candidate needs thirty features from a feature store, that becomes about three hundred million feature reads per second unless we cache, batch, or precompute. Now the budget becomes concrete: we might reserve about fifty milliseconds for feature fetch, one hundred milliseconds for scoring, and fifty milliseconds for overhead and network at the ninety ninth percentile. Those numbers are what justify components like caching layers, batching, and earlier candidate narrowing, and they keep our design grounded.
Multi-stage ranking pipeline: spend compute wisely
Dr. Wei: In real-time ranking, we cannot run the most expensive model on everything, so we use a multi-stage pipeline that steadily narrows the set while staying within a roughly 200 millisecond budget.
Sam: Predicting the funnel: I would start with maybe ten thousand candidates, then cut to one thousand, then one hundred, because one hundred thousand sounds too big to even touch in time. Am I underestimating how cheap the early stages can be?
Dr. Wei: Here is a running numeric funnel we will keep using: start with about 100,000 lightweight candidates in candidate generation, trim to 5,000 in first-pass ranking, then to 500 in full ranking, and finally return the top 50 after applying re-rank constraints.
Feature stores: offline, near-line, and online in one join
Dr. Wei: Features come from different time scales and different cost profiles. To keep ranking fast and reliable, we usually split features into three freshness tiers: offline, near-line, and online.
Dr. Wei: Offline features are computed in large batch jobs and change slowly, like a user’s long-term purchase rate over the last 30 days. They are cheap per feature, but the data is not the freshest.
Dr. Wei: Near-line features update from streams with a small delay, like how many times a user viewed this category in the last hour. Online features are computed right at request time, like whether the candidate item is in the user’s current session context. A common pattern is to hide all three behind one feature join interface so the ranking service asks once and gets the best available values.
Feature categories that actually matter in ranking
Dr. Wei: When we design a real time ranking system, the question is not just which model to use, but what kinds of signals the model should pay attention to.
Sam: So I should be able to name feature families and also explain who maintains them and how fresh they are, not just list random inputs to the model.
Dr. Wei: A good design names feature families and their ownership, so teams can reason about latency, debugging, and data drift without guessing where a signal comes from.
Model serving under a hard latency budget
Dr. Wei: When you serve a model in real time, latency is not a vague goal; it is a hard budget you must meet, request by request.
Sam: Prediction checkpoint: if we have two hundred milliseconds end to end, I would allocate maybe fifty milliseconds to candidate generation, one hundred milliseconds to ranking, and fifty milliseconds to reranking and business rules. Is that in the right ballpark?
Dr. Wei: A more complete breakdown is: end to end latency equals candidate generation plus feature fetch and joins plus model inference plus reranking, plus network and queueing overhead. If one stage grows, the others must shrink, or you miss the budget.
Real-time signals: make the feed react within seconds
Dr. Wei: In a real time ranking system, the point is not just to predict what a person will like in general, but to react to what is happening right now. That means the feed can shift within seconds as new information arrives.
Sam: So the key is that the next request should reflect the latest intent. If our streaming pipeline is delayed, the model is effectively scoring with stale context, even if inference is fast.
Dr. Wei: A useful way to think about it is: each new engagement event updates the system’s view of the user’s current intent, and then the very next ranking request should reflect that updated intent. If we cannot apply those updates quickly, the feed will feel stale and out of touch.
Re-ranking: optimize the whole list, not just scores
Dr. Wei: So far we have talked as if we can just score each item and sort. In real ranking systems, the final list has to behave well as a list, not only as a set of independent predictions.
Sam: Prediction checkpoint: if we only sort by score, the list might end up with near duplicates or too many items from the same creator or topic. Is the first thing that breaks usually diversity, or is it more about hard safety and policy rules?
Dr. Wei: A common first fix is a simple creator diversity constraint, like a max per creator in the top N results. Re-ranking takes the scored candidates and builds the final list while enforcing that cap: if the next highest score would exceed the max for that creator, we skip it for now and take the best remaining item from a different creator, continuing until the list is filled. Once you understand that pattern, other list-level needs are extensions of the same idea, like freshness, exploration, business rules, and safety filters.
$A/B$ testing and rollout: prove you improved the feed
Dr. Wei: Before we ship any ranking change, we need a reliable way to prove it actually makes the feed better, not just different. That means planning measurement up front, deciding what “better” means for users, and making sure we can detect improvements without causing hidden damage.
Sam: Can’t we just look at overall engagement after launch and see if it goes up?
Dr. Wei: Overall engagement can move for lots of reasons, and a small harmful effect can be masked by noise or seasonality. A controlled A/B test gives you a fair comparison between users who see the change and users who do not, and rollout guardrails help you stop quickly if key metrics regress, like retention, complaint rates, latency, or content quality signals.
Tradeoffs and level-up: what makes this $IC6$
Dr. Wei: To level up this real-time ranking system design, we need to go beyond listing components and show how we reason about hard choices.
Sam: So the senior signal is not naming a feature store and a model server, it is defending the budgets and the fallbacks. I should be ready to say what I cut first when latency spikes.
Dr. Wei: It also proposes mitigations and validation: what can go wrong, how we reduce risk with fallbacks and guardrails, and how we measure success with clear service level targets and experiments.
Exit ticket: allocate budget and pick a fallback
Dr. Wei: For this exit ticket, you will make one concrete design decision for a real time ranking system: how to spend a small latency budget, and what you will do when the system cannot meet that budget.
Sam: What kind of budget are we deciding, and what counts as a fallback?
Dr. Wei: Assume you have a strict end to end latency budget for ranking, like one hundred milliseconds. Decide how much you allocate to feature fetching, model scoring, and any re ranking or business rules. Then pick one fallback behavior, such as serving a cached ranking, using a simpler model, skipping expensive features, or returning a safe default ordering.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!