M2 Design Live Video Streaming
Loading learning experience...
Lecture transcript
Read the narration for M2: Design Live Video Streaming
From notifications to live video: the same fan-out problem, harder physics
Dr. Wei: Before we dive into live video streaming, let’s connect it to something you already know: notifications. In both cases, one event has to reach many people quickly, and that fan-out shape drives the design.
Sam: So it’s still one publisher and lots of receivers, but with a constant stream instead of a single message. Should I expect the same kinds of queues and fan-out services, just pushed a lot harder by timing?
Dr. Wei: The difference is the physics of time. A notification can wait a little and still feel fine. Live video cannot, because viewers are watching a stream of moments that are being created right now.
Requirements: three latency tiers, one brutal thundering herd
Dr. Wei: Before we design anything, we need crisp requirements, because live video systems are mostly a set of latency and scale tradeoffs. Today we’ll anchor on three concrete latency tiers, and we’ll also name the big scaling enemy we have to survive: the thundering herd.
Sam: When you say three tiers, are we choosing one for the whole product, or can different viewers or features land in different tiers?
Dr. Wei: Tier one is ultra-low latency: under one second end to end. This is the feel of real-time interaction, like live auctions or watch parties with voice chat, and it usually means keeping buffers tiny and treating delivery more like a real-time session than a queued download.
Dr. Wei: Tier two is low latency: roughly two to five seconds. That’s often the sweet spot for big broadcasts where you still want near live, but you also want the internet-friendly benefits of chunked, cacheable delivery.
Dr. Wei: Tier three is standard latency: about ten to thirty seconds. This is where you get the most buffer room and the most caching leverage, which usually makes it the easiest tier to scale. But no matter the tier, we have to plan for the thundering herd: huge numbers of viewers requesting the next segment at once after a boundary, a reconnect storm, or a peak moment.
Back-of-envelope: bandwidth is the boss
Dr. Wei: Before we pick CDNs, transcoding ladders, or caching strategies, let’s do a quick back-of-the-envelope bandwidth check. The goal is not a perfect number; it’s to learn which links and components will dominate the design.
Sam: Are we estimating this mostly to size network links, or to justify why we need a CDN in the first place?
Dr. Wei: Both. Quick prediction first: which do you think is bigger, ingest or viewer egress, and by about what factor, like ten times or a hundred times? Now for ingest: suppose we have about one hundred thousand live streams coming in, and each arrives at about four megabits per second. That’s roughly four hundred gigabits per second of ingest into the system.
Dr. Wei: Now do the same for a hot stream on the way out. Prediction: are we closer to a few terabits per second, or tens of terabits per second? If a single event has about five million viewers and each is watching around two megabits per second, that’s about ten terabits per second of egress. That’s not ten times ingest, it’s more like twenty-five times bigger than our four hundred gigabits per second ingest estimate, so egress dominates. Design consequence: you cannot serve that from one region or one provider network; you need a CDN or multi-CDN strategy, good peering, and aggressive caching so most traffic is delivered from edge caches rather than your origin.
APIs: separate control-plane from media-plane
Dr. Wei: When we design APIs for live video streaming, a key idea is to separate the control-plane from the media-plane. The control-plane is where clients create and manage a live session, while the media-plane is where clients fetch the actual playback timeline and related real-time streams.
Sam: So control-plane is like session setup and permissions, and media-plane is the heavy traffic. Should we keep them on different domains or even different infra to isolate failures?
Dr. Wei: Looking at the two groups here, the control plane endpoints start and stop a live session. In the media and realtime plane, the manifest endpoint is the entry point for playback, because it tells the player what to fetch next. Then interactions are separate too: one endpoint sends a comment, and another lets you consume a stream of reactions in near real time.
Data model: metadata small, segments huge
Dr. Wei: To design live video streaming, we need a simple data model that separates small, durable metadata from the huge volume of media created every second.
Sam: For segments, do we treat the object store as the source of truth, and the database just indexes sequence numbers to URIs? I’m trying to picture what must be strongly consistent.
Dr. Wei: Alongside media, we attach time-based data like comments and viewer estimates, because those need to align to the stream timeline for chat, analytics, and later replay finalization.
Ingest $+$ transcoding: turn one upload into an ABR ladder
Dr. Wei: When you design live streaming, a key goal is to accept one clean upstream and still serve many different viewers smoothly. This is where ingest plus transcoding comes in: we take a single contribution stream and turn it into multiple quality levels that playback can switch between.
Sam: If the transcoder falls behind, do players just buffer more, or do we drop quality aggressively to keep the live edge from drifting too far?
Dr. Wei: After transcoding, the stream is segmented and packaged into formats like HLS and DASH, then stored and distributed so edges can fetch and cache what viewers request. This whole chain is why a single upload can reliably reach a diverse audience with different devices and network conditions, while keeping latency and buffering under control.
CDN distribution: win the thundering herd at the edge
Dr. Wei: Imagine a big live moment: a goal is scored, and suddenly millions of people ask for the same few video segments at nearly the same time. If every request went back to your origin or object store, you would create a thundering herd and the backend would collapse under load.
Sam: Is the main trick here request coalescing at the edge, so one miss does one upstream fetch instead of a million?
Dr. Wei: To protect the origin even further, many deployments add a regional shield cache behind the edges. When an edge has a cache miss, it can coalesce requests so only one fetch goes upstream, the shield fills once, and then the edge can fan out the segment to everyone, while origin control still handles signing and authorization decisions.
Adaptive bitrate (ABR): switch quality only at segment boundaries
Dr. Wei: Adaptive bitrate, or ABR, is the idea that the player can change video quality on the fly to keep playback smooth as network conditions change.
Sam: If switching only happens at segment boundaries, does that mean shorter segments give you faster adaptation but worse overhead and caching behavior?
Dr. Wei: This design makes switching reliable, but it also ties responsiveness and latency to segment duration and how much content the player buffers. So when we pick segment length and buffering targets, we are also deciding how quickly ABR can react and how much end to end delay viewers experience.
Live comments and reactions: never fan-out raw events to every viewer
Dr. Wei: Now let’s talk about live comments and reactions on huge streams. The key idea is that interactivity is a different workload than video: it creates lots of tiny events, and the system has to deliver them quickly enough to feel real-time to viewers.
Sam: My first instinct is a publish subscribe bus and push every message to every viewer, but that feels like it explodes. Is that the trap you’re pointing at?
Dr. Wei: So the rule for this part of the design is simple: never fan-out raw events to everyone. Instead, we design a way to filter, aggregate, or sample the interaction stream so each viewer gets something timely and relevant, without the platform doing impossible amounts of delivery work.
Viewer count: approximate, fast, and hard to game
Dr. Wei: When we build live video at scale, we need a way to show how many people are watching right now. That number sounds simple, but in practice it has to be approximate, fast to compute, and difficult to manipulate.
Sam: Why can’t we just count every connected viewer and display the exact number?
Dr. Wei: Because “connected” is fuzzy in real life: viewers refresh, pause, lose network, switch devices, or sit behind shared addresses. If we try to keep an exact global count, we pay a big cost in coordination and latency, and the number can jump around. We also have to assume some clients will try to inflate the count, so our approach needs to be robust, not just precise.
Recording and replay: live segments become VOD almost for free
Dr. Wei: When we design live video streaming, we usually care about what happens right now. But the moment the broadcast ends, viewers still want to watch, share, and rewatch the same content. This slide is about how we can turn a live stream into something stable and reusable, without building a totally separate pipeline.
Sam: So we’re basically keeping the same segments and just switching from a sliding window to a finalized playlist. Is the hard part making sure the last few segments and metadata are consistent?
Dr. Wei: So instead of thinking about recording as capturing one giant file, think about it as finalizing what you already generated during the live event. You preserve the segments, you stop sliding the window, and you publish a final index that players can seek through reliably. That is how live segments become an on-demand asset almost for free.
Failure modes: degrade gracefully, never take down the whole live plane
Dr. Wei: Now let’s talk about what “failure modes” really means for live video streaming. In the real world, something is always a little bit wrong somewhere: a region gets flaky, a dependency slows down, or traffic spikes beyond what you planned for.
Sam: What’s the first thing you intentionally drop when the system is stressed: quality, latency targets, or interactive features like comments?
Dr. Wei: The key principle is: never take down the whole live plane. A problem in one service or one region should not cascade into a platform wide failure. We design isolation boundaries, safe defaults, and backpressure so the system can shed load and keep the core stream alive for as many viewers as possible.
Tradeoffs and rollout: hybrid latency, hybrid distribution
Dr. Wei: As we wrap up, think about shipping live video as a series of explicit tradeoffs, not a single perfect setting. Every choice affects the viewer experience, the infrastructure bill, and operational risk. This is where we decide what we optimize first and what we accept as a cost.
Sam: In an interview, is it better to pick one target tier and defend it, or to propose a hybrid and explain how you decide who gets which path?
Dr. Wei: Hybrid distribution means you can mix delivery methods to balance reach, cost, and control. You might use one approach to handle massive scale efficiently, and another approach to keep critical regions reliable or to reduce startup time. When you can, quantify the impact with metrics like end to end latency, rebuffer rate, video start time, and delivery cost per hour.
Dr. Wei: Finally, roll out in phases. Start with a small percentage of traffic, compare against a stable baseline, and have a clear rollback plan. The goal is to learn safely, so you can improve latency and distribution without surprising viewers or destabilizing the live event.
Exit ticket: reason about latency and interaction scaling
Dr. Wei: For the exit ticket, you will do two things: a quick latency calculation and a design judgment about how interaction scales as the audience grows.
Sam: For the latency calculation, do you want a number for standard HTTP streaming, or should I pick a tier and justify the assumptions I use?
Dr. Wei: Second, make a scaling call: if viewers can react, chat, or join a live Q and A, what happens to fan out, moderation load, and server cost as the audience moves from hundreds to tens of thousands? Propose one concrete change that keeps interaction usable, such as rate limits, aggregation, or smaller rooms, and explain the tradeoff you accept.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!