M2 Design Notification System
Loading learning experience...
Lecture transcript
Read the narration for M2: Design Notification System
Bridge: from encrypted messaging to user-facing alerts
Dr. Wei: Today we bridge from encrypted messaging to user-facing alerts. The core idea is the same: you are trying to get the right information to the right person at the right time, and you need a system that stays reliable even when networks and devices are flaky.
Sam: So we are basically reusing messaging ideas, but the output is an alert instead of a chat message. Should I think of it as fan-out plus retries plus a state machine, just tuned for user attention?
Dr. Wei: We will reuse three messaging concepts as our mental model. First, device fanout: one user often has multiple devices, so a single event may need to reach several endpoints, like a phone, a watch, and the web.
Sam: If I were designing it, I would start by listing all endpoints per user, then deciding whether we send to every endpoint or pick a primary one. Is that where device registration and last-active tracking come in?
Dr. Wei: Second, delivery receipts and retries: messaging systems watch for acknowledgments and retry when they do not arrive. For notifications, we will do the same kind of tracking, using provider responses and retry with backoff to improve delivery without spamming.
Sam: For retries, would you expect a per-channel strategy? Like a few fast retries for in-app delivery, but more conservative backoff for push or email so we do not annoy users or get rate-limited by providers?
Dr. Wei: Third, inbox and unread state: messages have a lifecycle, like sent, delivered, and read. Notifications also have state, like queued, delivered, shown, and dismissed, and that state is what lets the product be consistent across devices and across time.
Sam: And that state is also what drives badges and cleanup, right? Like if a notification is dismissed on the phone, the web client should not still show it as unread.
Problem framing: what are we building (and not building)?
Dr. Wei: Before we design anything, we need to frame the problem clearly: what we are building, what we are not building, and how we will decide whether it works. This keeps the design focused and prevents us from solving the wrong problem.
Sam: To scope it: are we optimizing more for transactional alerts and product updates, not marketing campaigns? And do we assume events are already legitimate, so we are not building heavy abuse detection here?
Dr. Wei: For this notification system, the core goal is simple: deliver the right message to the right user at the right time, reliably. That means we care about channels like email, push, or text message, and we care about prioritization so urgent notifications do not get stuck behind low-value ones.
Sam: For success metrics, should I treat opt-out rate as a hard guardrail, like we will sacrifice some engagement if it keeps opt-outs low? Or do you want a balanced scorecard where latency and success rate are leading indicators and engagement is lagging?
Dr. Wei: We also need to set boundaries and metrics. In scope are things like user preferences, templates, retries, and failure handling. Out of scope are adjacent products like a full chat app, deep spam detection, or a marketing analytics platform. And our success metrics should include delivery latency, send success rate, opt-out rate, and downstream engagement.
Estimation: events, deliveries, storage
Dr. Wei: Before we pick queues, caches, or sharding, we need one shared story about scale: how many active users we serve, how many notification events each user generates per day, how big one stored event is, and how long we retain it.
Dr. Wei: Quick prediction, Sam: if we had about fifty million daily active users and around one hundred notification related events per user per day, do you expect average traffic to be in the hundreds, the tens of thousands, or the millions of events per second?
Sam: My gut says tens of thousands of events per second on average, not millions, because five billion a day divided by a day is still only on the order of ten to the five per second.
Dr. Wei: Now we sanity check: fifty million times one hundred is five billion events per day, and dividing by seconds per day gives roughly fifty eight thousand events per second on average. If your guess was far off, that is the point: this number pushes us toward partitioning and buffering, and we still need headroom for peak traffic above the average.
Sam: For peak, would you assume something like five to ten times the average because of diurnal patterns and big launches? And is it fair to say the queue and worker autoscaling are what absorb that burstiness?
Dr. Wei: Storage uses the same assumptions, so it stays consistent: events per day times bytes per record times replication. With five hundred bytes per record and three copies, that is about seven and a half terabytes per day using decimal terabytes, and with seven days of retention you are around fifty two and a half terabytes before indexes and compression. Those back of the envelope numbers tell us what kind of database, sharding strategy, and retention policy are realistic.
APIs: internal event in, user inbox out
Dr. Wei: When we design a notification system, we want a clear contract between what the rest of the product asks for and what the user ultimately sees in their inbox. That contract is what lets teams move fast without breaking each other as we scale.
Sam: On the input side, is the contract basically an event like notification requested with an idempotency key, and then everything else is our system’s responsibility? And on the output side, the inbox is the stable interface for clients?
Dr. Wei: On the input side, we define an internal event boundary: producers send a single, stable event like notification requested, including who the user is, what template to use, and the data needed to render it. We also include an idempotency key so retries do not create duplicates.
Dr. Wei: On the output side, we define user-facing inbox APIs: list notifications with pagination and filters, and an endpoint to mark items as read or acknowledged. By separating these two surfaces, we simplify ownership, testing, and versioning, while keeping the user experience consistent.
Data model: events vs. notifications vs. preferences
Dr. Wei: In a notification system, the data model has to separate what happened from what we show to users, and from what users want. If we blur those ideas together, it becomes hard to scale, hard to debug, and easy to send the wrong message to the wrong person.
Sam: So events are immutable facts, notifications are derived user-facing records with delivery state, and preferences are separate rules. That separation also means we can reprocess past events to regenerate notifications, right?
Dr. Wei: An event is the source of truth: an immutable fact like an order shipped or a comment created. A notification is a derived, user-facing record that we can create, format, retry, and track delivery for. Preferences are the rules that decide whether and how a user should be notified, such as channel choices, quiet hours, and opt-outs.
Dr. Wei: Keeping these distinct gives us clean behavior: events stay append-only and auditable, notifications can evolve as templates and delivery mechanisms change, and preferences can be updated without rewriting history. This separation also supports reprocessing, like rebuilding notifications from past events after a bug fix, while still respecting current user preferences when we deliver.
High-level architecture: the notification decision pipeline
Dr. Wei: At a high level, a notification system is a decision pipeline: a stream of events comes in, we decide what to notify, and we learn from what happens after we send.
Sam: When you say decision pipeline, am I right that it is not just delivery? It includes steps like deduping, aggregating similar events, ranking, and then choosing a channel based on preferences and constraints.
Dr. Wei: We make the pipeline explicit so every stage is measurable, independently scalable, and easy to reason about when something goes wrong.
Dr. Wei: As we reveal it step by step, keep the verbs in mind: ingest and queue events, dedupe, aggregate, rank, apply a budget and final decision, choose a channel, send, then log outcomes and feed that back into future decisions.
Deep dive: aggregation to prevent notification storms
Dr. Wei: Let’s tackle a classic notification problem: a post suddenly gets ten thousand likes, and we need to avoid spamming the author with ten thousand separate push notifications.
Sam: So we need some kind of grouping rule. Is the usual approach a short time window plus a key like user and event type, so likes on the same post roll up into one message?
Dr. Wei: The core idea is aggregation: instead of emitting one notification event per like, we group many similar events over a short time window and send a single, meaningful message.
Sam: Would you store raw like events separately and have an aggregator generate the user-facing notification record? That way the rollup can be recomputed if the window size or template changes.
Dr. Wei: So the user sees something like, “Your post got ten thousand likes,” possibly with a few representative names, while our system stays stable under bursts and avoids a notification storm.
Deep dive: priority tiers $+$ per-user budgets
Dr. Wei: Now we are going to make notification delivery feel respectful instead of noisy by thinking in terms of priority tiers and per-user budgets.
Sam: How do you usually separate tiers from budgets in interviews? Like, tiers decide importance of an individual notification, while budgets cap the overall interruption rate per user across notifications and channels?
Dr. Wei: Priority tiers answer, “How urgent is this?” while budgets answer, “How often can we interrupt this particular person?” These are separate knobs that work best together.
Sam: So in implementation terms, I might rank within a tier, then apply a per-user token bucket for that tier, and spill lower tiers into an in-app inbox only. Is that a reasonable mental model?
Dr. Wei: The key idea is to reserve the highest tier for truly time-sensitive events, and then use per-user rate limits to prevent repeated pings, even when many events happen at once.
Delivery: push, in-app realtime, offline store-and-forward
Dr. Wei: Now let’s focus on delivery: getting a notification from our system to the user in a way that is timely, reliable, and consistent across all of their devices.
Sam: Should I describe delivery as a router plus channel-specific adapters? Like push for offline reach, in-app realtime when the user is online, and store-and-forward so the inbox stays correct even if delivery fails.
Dr. Wei: In practice, delivery is not one channel. The same user might be online in the app, temporarily offline on mobile, and also have a device that can receive a push alert even when the app is closed.
Sam: If we care about consistency, do we write to the inbox first and treat push as a best-effort side effect? That way the user can always fetch the authoritative list even if the push provider drops something.
Dr. Wei: Our goal is to route each message to the right devices, handle retries without spamming duplicates, and keep read and delivered state consistent so the user sees a coherent experience everywhere.
Correctness: dedupe, idempotency, and read-state races
Dr. Wei: Before we talk about dedupe and races, let’s agree on what “correct” looks like for a notification. A simple state machine is: created, then sent, then delivered, then seen, then read. Most product expectations assume these states move forward over time, not backward.
Sam: If we pick at least once processing, duplicates are expected internally. So the real question is where we place idempotency keys and dedupe: at the inbox write, at the channel send, or both?
Dr. Wei: We also need to be explicit about ownership. Typically the server is authoritative for created and sent, while the device is the source of truth for delivered, seen, and read events, and the server reconciles and persists those updates. If we do not define that contract, correctness bugs become impossible to reason about.
Sam: For read-state races, would you rather model state as a numeric progress level, so updates can only increase it? That seems simpler than trusting timestamps from multiple devices.
Dr. Wei: Now the delivery pipeline goal: we often choose at-least-once processing so we do not drop notifications when workers retry. But at-least-once means duplicates can happen, so every step that writes or fan-outs must be idempotent, and we must dedupe so a user never sees the same notification twice.
Dr. Wei: Read-state races are the subtler version of the same theme. If delivered, seen, and read updates arrive late, out of order, or are retried, the stored state can appear to go backward, like read becoming seen again. Prevent that by enforcing monotonic updates, for example with sequence numbers, last-write rules that respect ordering, or merge logic that always chooses the furthest state.
Reliability and tradeoffs and scaling scenarios
Dr. Wei: Before we wrap up, let's connect three big themes in notification systems: reliability, the tradeoffs we accept, and how those choices change as we scale from a small launch to a large platform.
Sam: When you say tradeoffs, do you want me to call out specific ones like at least once versus exactly once, realtime versus batch, and how hard we push rate limits to protect users even if engagement dips?
Dr. Wei: When we say reliability here, we mean things like delivering the right message to the right user, at the right time, without losing it, duplicating it, or sending it after it is no longer relevant.
Sam: And for scaling scenarios, should I describe what breaks first at ten times traffic? For example, ranking features might get expensive, inbox storage might become the bottleneck, and third-party push providers might impose throughput limits.
Dr. Wei: And the key recap point is that reliability is never free: improving it usually costs latency, complexity, and money, so mature systems often roll out in phases, tightening guarantees as traffic grows and as the business impact justifies the added engineering.
Exit ticket: size the system and justify one tradeoff
Dr. Wei: For the exit ticket, you will practice two skills you need in system design: sizing the system with quick, reasonable estimates, and explaining a product tradeoff clearly using numbers and constraints.
Sam: If I pick one tradeoff, I might choose immediate fanout versus queued delivery, and argue with a latency target and a retry policy. Is that the kind of justification you want, with a metric and a clear failure mode?
Dr. Wei: Assume a notification system where you need to estimate daily and peak traffic, the rough storage needed for message records, and the throughput you expect from the delivery pipeline. State your assumptions out loud so someone else can follow and critique them.
Sam: For assumptions, I will explicitly pick a daily active user count, events per user per day, average payload size, and a peak multiplier. Then I will compute average and peak events per second and translate that into partitions, worker capacity, and storage per day.
Dr. Wei: Then choose one tradeoff and justify it with a metric. For example: prioritize lower latency versus higher reliability, immediate fanout versus queued delivery, or sending to all devices versus sending only to the most recent device. Explain what you gain, what you give up, and how you would measure success.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!