M4 Design Experimentation Platform
Loading learning experience...
Lecture transcript
Read the narration for M4: Design Experimentation Platform
Design Experimentation Platform
Welcome everyone, today we will explore how to build an experimentation platform that scales at Meta, from deterministic assignment to statistically sound decisions with robust metrics and analysis.
From ranking to experiments: proving impact, not guessing
Dr. Wei: A great ranking system is only half the story. After you ship a change, you still have to measure whether it truly helped users, safely, and at scale.
Sam: So the point is not just “did metrics go up,” but “can we trust the measurement enough to bet the product on it,” right?
Dr. Wei: In this lecture on designing an experimentation platform, we move from making smart guesses to making trustworthy decisions based on controlled evidence.
Dr. Wei: The goal is simple: when we say a change improved the product, we can explain how we measured it, why we believe it, and what risks we checked along the way.
Sam: And I guess at Meta scale, even a small measurement bug can turn into a huge product decision, because so many teams and users depend on the same numbers.
Why Meta invests in experimentation infrastructure
Dr. Wei: At Meta, experimentation is not a special event reserved for big launches; it is the default way teams decide whether a product change actually helps people.
Sam: Is the main reason speed, like making decisions faster, or is it more about preventing teams from arguing based on anecdotes?
Dr. Wei: When thousands of changes are proposed across many products, you need a consistent way to run tests, measure outcomes, and make decisions that are fast and trustworthy.
Dr. Wei: That is why Meta invests in experimentation infrastructure: it reduces the cost of running a high quality test, improves measurement and safety, and helps teams learn quickly without guessing or debating based on opinions.
Requirements that drive the design
Dr. Wei: In experiment platform design interviews, requirements are not just a checklist; they are the reasons your system looks the way it does. If you cannot tie each component back to a specific need, the design will feel like a pile of random features.
Sam: Ten thousand concurrent experiments sounds extreme. If I had to simplify, could I just cap it at, say, a few hundred and tell teams to coordinate, or is that a hard requirement?
Dr. Wei: So we will make the requirements measurable. We need to run ten thousand or more concurrent experiments at once, assign millions of users per day, and keep assignment deterministic so the same person sees the same variant across sessions, devices, and services.
Dr. Wei: We also need strong isolation so experiments do not contaminate each other, which leads directly to layering and mutual exclusivity rules. And we need data to be useful quickly: exposures should be logged near real time, and core metrics should be available within twenty four hours.
Sam: For determinism across devices, do we need a global identifier? What if we only have a device identifier sometimes, or people are logged out?
Dr. Wei: Finally, we must protect decision quality and compliance. That means statistical controls like power planning, multiple testing guardrails, and unbiased reporting, plus governance like approvals, audit logs, and access control. When we later introduce hashing, layers, pipelines, and statistical checks, each one will map back to these requirements.
Experiment configuration: what must be stored and served
Dr. Wei: When we talk about an experimentation platform, the configuration is more than a form to fill out. It is the shared contract that defines how an experiment is described, understood, and executed across teams and systems.
Sam: If config is the contract, what is the minimum we must store to be safe? I worry we end up with a giant schema that is hard to evolve.
Dr. Wei: A practical minimum contract usually includes: experiment_id, owner, layer_id, and unit_id_type; the variants with allocation weights and a ramp schedule; an assignment salt and assignment_version so assignment is stable and debuggable; eligibility filters; start_time and end_time or stop rules; and the primary metrics and guardrails to evaluate safety and success.
Dr. Wei: So in this section, keep one question in mind: if a service is running in production and needs to know exactly how an experiment should behave, what information must be persisted and retrievable, and what information can be derived later from defaults or conventions?
Traffic splitting: consistency without per-user storage
Dr. Wei: When we say “traffic splitting without per-user storage,” we mean we can compute a user’s assignment on the fly and still get the same answer every time they return. The platform needs this to be stable, fair, and cheap to serve at high volume.
Sam: Why not just store user-to-variant assignments in a database? Then we would not have to worry about hash changes or bucket math.
Dr. Wei: Predict before we reveal the details: what breaks if we hash only on user id and not on experiment id? And what happens if we change the bucket count from one hundred to one thousand, or if we change the variant weights while the experiment is running?
Sam: My guess: if we hash only on user id, the same people might always land in treatment across all experiments, so experiments could interfere. If we change the bucket count, it probably just changes the precision but keeps most users in the same variant. And if we change weights, maybe it only affects new users and existing users stay put.
Dr. Wei: Here’s the robust pattern: hash a full key that includes a namespace or layer, the experiment id, the user id, and a stable salt, then take a large intermediate bucket space, like mod one million. After that, map bucket ranges to variants using configured weights, so you can express fifty–fifty splits, ramp-ups, and holdouts without changing the hash itself.
Sam: Okay, but if we do a ramp from one percent to fifty percent, does that move a bunch of users between variants and mess up the analysis?
Dr. Wei: Let’s close the loop on your predictions. First, you were right that hashing only on user id can couple experiments: the same users can systematically land on the same side across unrelated experiments, creating unwanted correlation across features. Second, changing the bucket count does not just increase precision: it re-randomizes a lot of users because the modulo changes the mapping, so many people jump buckets and you lose consistency. Third, changing weights mid-run is not limited to new users: any weight change means some existing users will switch variants because the bucket ranges move. That can bias results unless you treat it as an explicit ramp with safeguards: record allocation and ramp changes over time, define clear exposure windows, and analyze by intent-to-treat or by time-based exposure, or freeze allocation after the experiment starts unless you are deliberately ramping.
Layers: preventing experiments from stepping on each other
Dr. Wei: Meta runs many experiments on the same surfaces; without isolation you get interference and unreadable results.
Sam: If interference is the problem, why not just forbid overlapping experiments entirely? Would layers be overkill?
Dr. Wei: Layers are how we extend the deterministic assignment idea from the previous slide so that multiple experiments can coexist safely. Think of a layer as its own bucket universe, like a million numbered buckets, and each experiment inside the layer is allocated a disjoint reserved range of those bucket numbers. A user hashes to one bucket number in that layer; if that number falls inside experiment A's reserved range, the user is eligible for A, and if it falls outside, the user is not in A. Because those reserved ranges do not overlap, two experiments in the same layer cannot accidentally assign the same user at the same time.
Sam: So is the layer basically the first gate, and then experiments compete inside that gate? How do you decide which experiment wins if two want the same layer?
Dr. Wei: Here is the simple flow: first, a user is deterministically hashed into a layer bucket. Then we check which experiment range, if any, contains that bucket number; that picks at most one experiment because the ranges are disjoint. Finally, inside the chosen experiment, the user is assigned to a variant like control or treatment.
Serving-time architecture: fast assignment on every request
Dr. Wei: When we talk about serving-time architecture, we mean the system that decides a user’s experiment treatment in the critical path of a request. That decision has to be fast, correct, and consistent, because it happens at the exact moment we are trying to serve the product.
Sam: If assignment is on the critical path, do we call a central service every request, or should each product service compute it locally? I am worried about a dependency outage.
Dr. Wei: Assignment sits on the hot path, so we optimize for ninety-ninth percentile latency and safe rollout. In practice, that means we budget milliseconds, avoid fragile dependencies, and design for predictable behavior under load.
Dr. Wei: We also need operational safety: gradual ramp-up, monitoring, and quick rollback if something goes wrong. The goal is that experiments can be enabled without risking the core serving experience, even during traffic spikes.
Metric pipeline: from event logs to dashboards in 24 hours or less
Dr. Wei: When we build an experimentation platform, we have two promises to keep: decisions must be fast, and they must be based on reality. The assignment service makes experiments run, but the metric pipeline is what makes the results trustworthy.
Sam: If some logs are missing, can we just ignore those users in analysis, or do we have to fail the experiment? I am trying to understand the safe default.
Dr. Wei: If assignment is the hot path, metrics are the truth path: bad logs means bad decisions. A missing event, a duplicated event, or a late event can quietly flip a win into a loss, or the other way around.
Dr. Wei: So the goal of this pipeline is simple: capture events reliably, process them consistently, and land clean metrics in a store that dashboards can query. And we want that full loop, from user action to a plotted metric, to finish in under a day so teams can iterate safely.
Statistical rigor: turning metrics into decisions
Dr. Wei: A platform that computes a mean is not enough; you need controls against false discoveries.
Sam: Could we just run a standard t test per metric and call it a day? If p is less than point zero five, we ship.
Dr. Wei: On this slide, the key idea is that metrics only matter if they lead to reliable decisions: do we ship, roll back, or keep learning? Statistical rigor is what turns noisy measurements into an honest signal you can act on.
Dr. Wei: That means we define guardrails up front: clear hypotheses, a decision rule, and protections for common failure modes like peeking, repeated testing, many metrics, and segment slicing. Without those controls, you can get a plausible looking lift that is actually just randomness.
Power and sample size: why tiny effects need huge n
Dr. Wei: One of the most important intuitions in experimentation is that tiny effects are hard to see through noise. When you ask for a very small detectable lift, you usually have to compensate with a much larger sample size.
Sam: If we have huge traffic, can we skip power planning and just run it for a day or two? It feels like the data will tell us quickly.
Dr. Wei: In many settings, sample size grows roughly with the variance of the metric, and it shrinks with the square of the effect size you want to detect. So cutting the detectable effect by ten can push the required sample up by about a factor of one hundred.
Dr. Wei: For example, if you are targeting a point one percent lift, the required n can easily reach millions per variant, especially for high variance metrics or rare outcomes. The exact number depends on the baseline rate, the metric variance, and your chosen significance level and power, but the direction is the key takeaway: smaller Delta means dramatically larger n.
Interaction effects: when A and B do not add up
Dr. Wei: Even when we run experiments carefully and analyze them correctly, results can still feel surprising once two changes are in play at the same time. The key idea on this slide is interaction effects: the combined impact of two features can be different from what you would expect by just adding their separate impacts.
Sam: If layers prevent overlap, do we still need to worry about interactions? Or is this mainly a problem when we allow multiple changes to touch the same users?
Dr. Wei: Imagine feature A gives a small lift on its own, and feature B also gives a small lift on its own. If the product experience changes in a way that makes them reinforce each other, the lift together could be much larger than either one suggests. But if they compete for the same user attention, the combined effect could shrink, or even flip direction.
Dr. Wei: This matters for an experimentation platform because independence is not guaranteed in the real world. If we ignore interactions, we can ship the wrong combination, misread results from overlapping tests, or over generalize from a single experiment. So here we are building intuition for why two good ideas do not always add up to an even better outcome.
Tradeoffs and what separates I C 4 vs I C 6 answers
Dr. Wei: Let us make the tradeoffs explicit, because interviewers score decisions, not diagrams. In an experimentation platform, nearly every design choice is a trade between speed, accuracy, reliability, and cost.
Sam: In an interview, should I prioritize the assignment side or the metrics and stats side first? I am never sure what the interviewer expects at the start.
Dr. Wei: At one level, a solid answer names a reasonable default and shows you understand what can go wrong. At the next level, you proactively surface the biggest tensions, explain how you would measure impact, and justify why you picked one side of the tradeoff for this product.
Dr. Wei: As you listen to yourself, check for this: are you just listing components, or are you making decisions like coverage versus latency, strict correctness versus availability, and flexibility versus simplicity? That decision making, with clear assumptions and consequences, is what separates an I C 4 style answer from an I C 6 style answer.
Exit ticket: deterministic assignment plus layers plus safe decision
Dr. Wei: For this exit ticket, you will pull together three ideas from the experimentation platform: deterministic assignment, layered configuration, and making a safe decision when the data is not perfectly clean.
Sam: Should my answer include what we do when the user identifier changes, like someone logs in mid-session? I feel like that is where deterministic assignment gets tricky.
Dr. Wei: First, do one concrete assignment: pick a user identifier and describe, in plain words, how you would assign that user to a variant in a way that is repeatable across sessions and services.
Dr. Wei: Then, explain how layers affect that outcome: for example, how an experiment layer interacts with global defaults, app specific overrides, and a kill switch, and which layer should win when they disagree.
Dr. Wei: Finally, reflect on a hard edge case and make a safe decision: what should happen if the identifier is missing, if the hashing input changes, or if traffic is too small to trust the metric yet. State the conservative behavior you would ship to protect users and data quality.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!