M2 Design WhatsApp at Scale
Loading learning experience...
Lecture transcript
Read the narration for M2: Design WhatsApp at Scale
Design WhatsApp at Scale
Welcome everyone, today we will explore how WhatsApp is designed at scale, building encrypted messaging, group and multi device sync, and efficient status, all for billions of users.
From Messenger to WhatsApp: same scale, different constraints
Dr. Wei: Before we design anything, let’s name what changes when we move from a feature-rich messenger to WhatsApp. The scale can be similar, but the rules you have to play by are very different.
Dr. Wei: In a typical plaintext or feature-heavy messenger, the server can store messages, index them, scan content, and power features like search, smart replies, and moderation based on what it can read.
Dr. Wei: For WhatsApp, the core requirement is end-to-end encryption by default, plus minimal server state and retention. That means the server should keep only what it needs to route and retry delivery, support multi-device and offline use, and still provide some path for safety and abuse reporting without full content visibility.
WhatsApp vs Messenger: what changes when content is unreadable
Dr. Wei: Same parent company, but the user promise is not the same. In Messenger, the service can, in principle, read the message content to power features like search, safety tooling, and smart suggestions. In WhatsApp, the core idea is different: the content is designed to be unreadable to the service itself, so the product has to be built around that constraint.
Sam: So do we basically assume the server only sees metadata like who is talking to whom and when, but not the content itself?
Dr. Wei: That single decision changes what data you can collect, what you can debug, and what kinds of reliability and abuse prevention you can do. It also changes the threat model: who could learn what, and from which signals. So when we say, "design WhatsApp at scale," we are really saying, "design a global messaging system when the server cannot rely on message text."
Numbers that justify the architecture
Dr. Wei: Before we talk databases and queues, we need a quick back-of-the-envelope traffic estimate. The goal is not a perfect number; it is the right order of magnitude so we know what kind of architecture pressure we are dealing with. One important caveat: we should split traffic into at least two buckets, because “send” requests and “metadata” requests have very different costs.
Sam: Okay, I will guess: maybe 2 billion sends per day, and peak traffic is like 5 times the average? And I guess there are also a bunch of acknowledgments and fetches on top of that.
Dr. Wei: Exactly. Let us make the definitions concrete. Bucket one is SendMessage actions: client sends that hit the send hot path. Bucket two is metadata traffic: delivery acknowledgments, read receipts, and message fetch or sync calls. For a simple mapping, think: one text send is about one SendMessage call, then N delivery attempts to recipients, plus roughly N acknowledgments back, and later fetches when devices reconnect. Media sends can dominate bandwidth because the upload and download paths are large and sometimes use separate media servers, so we do not treat media bytes as one tiny request. Worked example: suppose we have 10 billion SendMessage calls per day, and ks equals 8 for a busy hour peak. Average send requests per second is 10 billion over 86,400, about 116 thousand per second on average.
Dr. Wei: Peak SendMessage requests per second is ks times that, so about 8 times 116 thousand, roughly 930 thousand per second. Now add the second bucket: if acknowledgments plus fetches average, say, three times the send rate, then metadata average is about 350 thousand per second and peak can be multiple millions per second. These buckets drive different constraints: SendMessage is CPU and tail latency on the hot path, metadata is storage and cache pressure plus a lot of small network round trips, and media transfer is mainly bandwidth and egress cost. Splitting them keeps the architecture implications grounded instead of treating everything as one uniform QPS number.
API and data model: the server stores envelopes, not meaning
Dr. Wei: In this part, we’ll pin down the simplest contract between client and server that still gives us reliable messaging at scale.
Sam: When you say contract, are we talking about a small set of endpoints like send, fetch pending, and acknowledge, plus a minimal message envelope schema?
Dr. Wei: Yes: think of three endpoints, send, fetch pending, and acknowledge, and a compact envelope with fields like message id, sender id, recipient id and device id, timestamp, content type, ciphertext, time to live, and attempt metadata like attempt count and last attempt time.
Encrypted message flow: the server is a delivery pipeline
Dr. Wei: In an end-to-end encrypted messenger, the server is not a place that reads messages. It is a delivery pipeline: it accepts a sealed packet from one device and tries to get that same sealed packet to the other device as reliably and quickly as possible.
Sam: If the server cannot read it, does it still validate anything beyond basic size limits, like spam signals or malformed payload checks?
Dr. Wei: Think of the flow as a chain of handoffs. The sender connects to the closest entry point, the system forwards the ciphertext through internal relays, and delivery happens if the recipient is online at that moment.
Media message flow: send a pointer, not the bytes
Dr. Wei: When you design chat at WhatsApp scale, media is what threatens to overwhelm the system. Photos and videos are huge compared to text, and they arrive in bursts during peak hours.
Sam: Can we put numbers on that? Say peak traffic is one hundred thousand messages per second, and five percent of messages include a two megabyte video. What egress does that create if we send the bytes through the message path?
Dr. Wei: Great instinct. Five percent of one hundred thousand is five thousand videos per second. At two megabytes each, that is about ten gigabytes per second of payload, before overhead and before any fanout or retries. That is why the key idea is: do not push the raw bytes through the same path that delivers messages. Instead, treat media like a separate data plane with storage and caching that is built for large files, and send a small pointer plus decryption info in the message.
Group messaging: avoiding O of n encryption per send
Dr. Wei: In group chat, scale changes the rules: a design that feels clean for one-to-one can become expensive when every send touches many people.
Dr. Wei: Sam, suppose you did the most straightforward thing: for each message, you encrypt it separately for each group member’s device. What does that do to the sender’s work as the group grows?
Sam: It grows linearly with the group size, because the sender is doing one encryption per recipient.
Dr. Wei: Right. So what optimization might keep per-message sender work roughly constant, while still letting every member decrypt?
Sam: Have the sender use one shared sender key to encrypt the message once, and make sure group members have that sender key.
Dr. Wei: Exactly. With a sender key, the sender’s per-message compute can be constant time. But there’s still a linear cost when membership changes: you need to distribute or update keys across the group, and recipients collectively still do work to process messages. So the win is about per-message sender cost, not making all group overhead disappear.
Multi-device: from phone-mirror to independent devices
Dr. Wei: Multi-device support is mostly a consistency and key-state problem, not a throughput problem. Our goal from the start is: multi-device plus offline delivery, while keeping the same end-to-end encryption guarantees. So we have to decide whether devices are just mirrors of the phone, or truly independent endpoints that can receive and decrypt while the phone is offline.
Sam: If we want independent devices, does that mean each device needs its own set of session keys, and the sender might have to encrypt separately per recipient device rather than per recipient user?
Dr. Wei: Exactly. A common starting point is the phone-mirroring model: the phone is the anchor for identity, session keys, and conversation state, and other clients are extensions that depend on the phone staying connected or at least staying authoritative. To reach the independent-device goal, we move to per-device sessions and server fanout to per-device queues: the sender encrypts once per recipient device, the relay stores only opaque ciphertext and minimal routing metadata until each device pulls it, and that preserves offline delivery without the relay learning message contents.
$Status/Stories$: twenty four hour TTL with contact fan-out
Dr. Wei: Now let’s talk about Status or Stories in a WhatsApp-like system. This feature looks a bit like a feed, but it has two defining constraints: it expires automatically after twenty four hours, and it’s mostly seen by your existing contacts rather than the whole internet.
Sam: So the main difference from chat is that it’s time-limited and one-to-many instead of a direct conversation?
Dr. Wei: Exactly. A status update is posted once and then fanned out to many viewers, usually a few hundred contacts at most. Because it expires quickly, we can treat storage and caching differently, but we still have to handle bursts when someone popular posts and many people open the app to view it.
Efficiency primitives: how the backend stays small
Dr. Wei: At WhatsApp scale, reliability and efficiency usually do not come from a huge number of clever features. They come from a small set of boring, well-instrumented primitives that you can operate confidently day after day.
Sam: When you say primitives, are you thinking things like idempotent send, a durable pending store with TTL, and a simple retry and backoff policy, rather than lots of server-side features?
Dr. Wei: When we say the backend stays small, we mean the number of core moving parts stays limited, and each part has a clear responsibility, a stable interface, and predictable failure modes.
Failure modes and tradeoffs (what interviewers probe)
Dr. Wei: To wrap up, we will name the costs that come with our design choices in a WhatsApp style system. Interviewers are not looking for perfection; they want to hear that you can predict where things break, what you are optimizing for, and what you are willing to give up.
Sam: So it is less about memorizing one correct architecture, and more about explaining the tradeoffs clearly?
Dr. Wei: Exactly. For example, if you choose stronger consistency for message delivery state, you may add latency or reduce availability during partitions. If you optimize for availability and low latency, you may accept brief inconsistencies like duplicated sends, out of order receipts, or delayed sync on a second device, and you must explain how you detect and recover.
Sam: What are the failure modes interviewers usually probe for in a chat system like this?
Dr. Wei: They often probe overloaded services, hot partitions, data loss risk, retry storms, and cascading failures. They also ask about client behavior under bad networks, like reconnect loops, duplicate sends, and long offline periods. The key is to pair each risk with a mitigation, such as idempotency keys, backoff, rate limits, queues, graceful degradation, and clear operational signals for detection and rollback.
Exit ticket: quantify and decide
Dr. Wei: Exit ticket time: we will do one quick back of the envelope estimate, then make a design choice and defend it with a clear trade off.
Sam: For a quick estimate, would you rather I size peak sends per second, or the pending store footprint if we keep undelivered messages for, say, a day?
Dr. Wei: Then decide: based on your estimate, what would you prioritize first, write availability, low latency reads, or cost efficiency? State the assumption you relied on, what could break it, and one mitigation you would add if the assumption is wrong.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!