M5 Mock Full Meta System Design Interview
Loading learning experience...
Lecture transcript
Read the narration for M5: Mock Full Meta System Design Interview
Mock Full Meta System Design Interview
Welcome, everyone, glad you are here. Today we will walk through designing a Meta scale real time notifications system, with coaching on interview signals and I C 6 calibration.
Interview simulation: design real-time notifications
Dr. Wei: Today we are switching from discussing tradeoffs in isolation to running a full mock system design interview. The goal is to practice thinking clearly under time pressure while keeping your answers structured and easy to follow.
Sam: What should I expect in this simulation, and how should I pace myself during the interview?
Dr. Wei: We will follow a realistic interview loop: clarify the problem, define success metrics, propose a high-level design, then drill into key components and tradeoffs. We will pause periodically for coaching so you can adjust your approach, and by the end you should be able to tell a coherent end-to-end story for real-time notifications.
How we spend the 45 minutes ($S-E-D-E$)
Dr. Wei: In a forty-five minute system design interview, your biggest advantage is controlling the pacing. When you timebox out loud, you show structure, you keep the conversation moving, and you make it easy for the interviewer to follow your decisions.
Sam: Before we start, can I try a quick prediction? In most interviews I either over-scope or get stuck too long on estimating. My guess is I will underestimate how much time design and evaluation actually need.
Dr. Wei: We will use a simple rhythm called S E D E. It stands for Scope, Estimate, Design, and Evaluate, and it is a reliable way to cover the right depth without getting stuck too early on details.
Phase $1$: scope the problem like an owner
Dr. Wei: Before we design anything, we have to scope the problem like an owner, not just a candidate trying to guess what the interviewer wants. In this phase, the goal is to turn a vague request into a crisp problem statement with clear boundaries, success criteria, and constraints.
Sam: Let me predict what you will care about in scope. I think the first decision is whether this is just user facing notifications, or also internal notifications and experimentation. I would start by narrowing to user notifications across push and in-app, then ask what success means: delivery latency, send accuracy, and user fatigue.
Dr. Wei: By the end of Phase 1, we should be able to say, in one or two sentences, exactly what system we are building, what we are not building, and how we will judge success. If we skip this, we risk designing the wrong thing extremely well.
Phase 1: estimate the scale (numbers drive design)
Dr. Wei: Before we pick any databases or queues, we need to estimate the scale. The goal of this phase is simple: write down a small set of concrete assumptions, then turn them into back of the envelope numbers for requests per second and bandwidth so the rest of the design is grounded.
Sam: Prediction checkpoint: if we are aiming for around a million events per second on average, I think the first bottleneck will be network and fanout, not compute. The raw ingress is already big, and multiplying by internal consumers and regions will really amplify bandwidth.
Dr. Wei: Let’s do a quick warm-up first so the arithmetic is clear: 10 million daily active users times 10 events per user per day is 100 million events per day, which is about 1,160 events per second on average. With a 10 times peak factor, that is about 11,600 events per second peak. At 1 kilobyte each, that is about 11.6 megabytes per second, or about 93 megabits per second of peak ingress.
Dr. Wei: Now scale that up to the Meta-scale target that matches our interview: say 1 billion daily active users and 100 events per user per day. That is 100 billion events per day, or about 1.16 million events per second on average, and about 11.6 million events per second at 10 times peak. At 1 kilobyte each, peak ingress is on the order of 11.6 gigabytes per second, around 93 gigabits per second, before any multipliers. Then include the often-missed multipliers: three internal consumers can mean about three times downstream egress, and active active across two regions can mean roughly another two times for replication and cross-region traffic. Your prediction about fanout and regions was on target: at this scale, those multipliers are often what force architectural choices.
Phase $2$: A-P-I surface (internal + user-facing)
Dr. Wei: In this phase, we define the A P I surface: the set of calls and contracts that connect clients, internal services, and storage. Getting this right early keeps the rest of the design consistent, because it forces us to be explicit about inputs, outputs, and guarantees.
Sam: Prediction checkpoint: I think we should expose a single event ingestion endpoint, and make it idempotent with an explicit idempotency key so clients can retry safely. For reads, we need a paginated list endpoint for a user’s in-app inbox, plus maybe an acknowledge or mark-read call later.
Dr. Wei: A solid contract here is usually just two endpoints: an internal post to ingest an event and return an intent identifier, and a user-facing get that pages through a user’s notifications. This matches your prediction: the key interview signal is that you named idempotency and pagination as first-class contract details, not implementation trivia.
Phase 2: data model (partitioning and $T-T-L$)
Dr. Wei: In phase 2, we lock down the data model: what we store, how we partition it, and how long it lives. This is where we translate product requirements into concrete records and keys that keep the system fast at scale.
Sam: Prediction checkpoint: I expect two core stores. One is a per-user notification inbox keyed by user id with a time-ordered sort key, so listing recent items is easy. The other is a dedup store keyed by notification id with a time to live, so retries do not create double sends.
Dr. Wei: Now point your eyes at the table. It shows exactly that split: NotificationStore partitioned by user id with created at milliseconds as the sort key, plus DedupStore partitioned by notification id with retention on the order of days. Your prediction is the right mental model, and the table makes the read patterns and retention explicit so we can reason about hot shards and costs.
Phase $3$: high-level architecture (async, prioritized, multi-channel)
Dr. Wei: At a high level, we want one clear flow from an event to a delivered notification, with clean boundaries between stages so each part can evolve without forcing a redesign of the whole system.
Sam: Prediction checkpoint: my high-level architecture would be an async pipeline. Producers call an ingest service, then we buffer with a queue, then a decisioning stage fans out into two delivery paths: in-app writes plus realtime sockets, and push dispatch to external providers.
Dr. Wei: Now compare that to the diagram: Event Producers flow into Notification Ingest, then into a Priority Queue, then Decisioning, which splits into Push Dispatcher and In-app Dispatcher. In-app writes to the Notification Store and can also go through a WebSocket Gateway, while Push goes to APNs and FCM. Your prediction matches, and the extra design choice we make visible here is the priority queue between ingest and decisioning.
Deep dive $1$: priority, batching, and fatigue control
Dr. Wei: Not all notifications deserve the same latency or delivery cost.
Sam: Prediction checkpoint: I would define a small number of priority tiers, and I suspect most load can be smoothed with batching for lower tiers. I would also expect fatigue controls to be a bigger product win than micro-optimizing the dispatch path.
Dr. Wei: We will look at three levers: priority tiers to protect critical messages, batching windows to reduce load and improve experience, and fatigue control so the system respects user attention and avoids repeated interruptions across channels. This lines up with your prediction: tiers and batching handle system pressure, while fatigue control is where product quality and trust show up.
Deep dive $2$: deduplication (retries without double-sends)
Dr. Wei: In an at-least-once delivery system, retries are normal, so we need a clear definition of “the same notification” and a way to ensure we never send it twice even if we process it multiple times.
Sam: Prediction checkpoint: I think the key is to define a canonical notification id based on intent, not attempt. For example, derive it from user id, event type or template, and source event id, then use it as the idempotency key at every stage.
Dr. Wei: Then we store that notification id in a fast dedup store like Redis or Cassandra with a time to live, say 7 to 30 days depending on retry and backfill needs. Each stage becomes idempotent: enqueue checks and sets the key, render can cache by notification id, and provider send uses the same id as an idempotency key so repeated sends become no ops. That is exactly the intent-versus-attempt distinction you predicted, and it is what makes at-least-once workable.
Deep dive add-on: preferences filtering $+$ Gatekeeper
Dr. Wei: In a large-scale notification system, there is a step after we assemble a set of possible notifications and before we decide what to send, where we apply user preferences and safety rules. This step protects the experience while keeping latency under control, and it has to work reliably across many experiments and product surfaces.
Sam: Prediction checkpoint: for notifications, I would separate user preferences from safety or eligibility checks. Preferences can fail open in some cases, but safety rules should fail closed or at least degrade in a clearly conservative way.
Dr. Wei: Think of it as two related layers: preferences filtering, which removes notifications the user explicitly does not want, and a Gatekeeper layer, which is a policy and eligibility check service that blocks notifications that fail rules like integrity, policy, recipient eligibility, or required constraints. Together they must be fast, safe, and easy to iterate on without destabilizing delivery.
Delivery guarantees: match correctness to priority
Dr. Wei: A single guarantee for everything is either too expensive or too weak.
Sam: Prediction checkpoint: I would map guarantees to tiers. Low priority can be best effort with occasional drops, medium priority should be at least once with dedup, and top priority might need stronger end-to-end accounting even if we still cannot truly promise exactly once across external providers.
Dr. Wei: Treat delivery guarantees like a budget. For low priority events, you can accept duplicates or occasional drops. For high priority actions, you may need stronger protection against loss, duplication, and reordering, even if it adds latency and cost. Your tiered prediction is the right framing, and in interviews it shows you match correctness to user impact instead of over-engineering everything.
Phase $4$: failure modes (design for the incident)
Dr. Wei: Phase 4 is where we stop assuming the happy path and ask what happens when the system is stressed, partially broken, or behaving in unexpected ways. If you cannot describe degradation behavior, the design is incomplete.
Sam: Prediction checkpoint: the first incident I expect is queue lag during a spike, followed by provider failures for push. My default mitigation would be to shed or batch low priority notifications, keep high priority flowing, and make degradation visible in metrics like queue age and send error rate.
Dr. Wei: That is the right way to think in this phase: pick concrete failure modes, name the user impact, decide what signals will reveal the problem early, and choose mitigations that degrade safely. Your prediction lines up with common incidents, and this is the interviewer signal: you did not just name failures, you described user impact, how you detect it, and how the system degrades safely.
Scaling $+$ multi-region routing (where the user is)
Dr. Wei: Now let’s zoom out and talk about what happens when our system has users and services spread across multiple geographic regions. The core idea is simple: we want requests and real time events to feel fast and reliable no matter where the user is located, even as load grows and failures happen.
Sam: Prediction checkpoint: I would route users to a home region for reads and websocket connections, and replicate write-ahead events across regions asynchronously. My worry is cross-region fanout cost, so I would try to keep decisioning close to where the user is, and only replicate the minimum state needed for correctness.
Dr. Wei: But we also need to handle the messy parts: users travel, traffic spikes, and a region can degrade or go down. So the routing strategy has to include sensible failover and a plan for what data must be local versus what can be replicated asynchronously. The goal is predictable user experience even when the underlying topology changes.
Tradeoffs $+$ what made this $I-C-6$
Dr. Wei: To close out, let’s zoom out and connect the design back to the interview goals. We’ll summarize the biggest tradeoffs we made, why we chose them, and how we would roll this out safely in phases.
Sam: Prediction checkpoint: the tradeoff you will highlight is probably at least once delivery with dedup, instead of trying to promise exactly once. And another one is prioritizing the pipeline so we can shed load for low priority, rather than scaling everything linearly.
Dr. Wei: Exactly. In a system design interview, it’s not enough to propose a solution. You want to show you can reason about what you gave up, what risks you accepted, and what you’d do differently if the constraints changed. Then you tie that reasoning to a rollout plan that reduces blast radius and proves the design with real data.
Exit ticket: capacity math $+$ design reflection
Dr. Wei: We are going to close with an exit ticket that checks two things: capacity math and design judgment. You will do a quick back of the envelope estimate, then you will reflect on a tradeoff you made and explain why it was reasonable.
Sam: Prediction checkpoint: I expect my biggest mistake will be forgetting multipliers, like fanout to multiple internal consumers and cross-region replication. I will try to state those explicitly and then call out which resource I think will saturate first.
Dr. Wei: Second, the reflection part: name one principled tradeoff you chose, for example latency versus cost, consistency versus availability, or simplicity versus flexibility. State what you optimized for, what you gave up, and what signal would make you revisit that decision later.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!