M4 Design Ad Click Analytics
Loading learning experience...
Lecture transcript
Read the narration for M4: Design Ad Click Analytics
From trending counts to money counts
Dr. Wei: Today we are shifting from analytics that are mainly about insight to analytics that drive real dollars. When a number is used for billing, small errors stop being close enough and become real money problems. This is the bridge into ad click analytics.
Sam: So the bar changes from good enough to spot a trend to good enough to send an invoice. Are we mainly optimizing for correctness first, or freshness first, or do we need both?
Dr. Wei: In the last lecture, we focused on trending style questions where approximation is often acceptable: you want the big movers, the top items, and you can trade a bit of accuracy for lower cost and faster results. That mindset works well for exploration and monitoring.
Sam: Right, and with money counts, even a tiny percentage error can be huge. I am guessing we still want minute-level dashboards, but we cannot let that leak into billing logic.
Dr. Wei: Now we will apply similar streaming ideas to a stricter setting: counting clicks and impressions with exactly once behavior, plus freshness on the order of minutes. The theme is how system design changes when the output is tied directly to revenue and contractual reporting.
Scope and success metrics (what are we building?)
Dr. Wei: Before we draw boxes and arrows, we need a crisp scope for what this system is and is not. For ad click analytics, that means agreeing on the user-facing outcomes, the promises we must keep, and the kinds of questions the system should answer reliably.
Sam: Let me try to state it: we ingest impressions and clicks, we compute billing-grade counts and near-real-time metrics like click-through rate, and we serve dashboards and APIs. What exact service-level objectives should we assume for freshness and correctness?
Dr. Wei: We are building a pipeline that captures raw ad events, turns them into trustworthy metrics, and then delivers those metrics quickly to people and downstream systems. The key is to define success in a way that balances speed, accuracy, and cost, instead of optimizing only one of them.
Sam: So success is not one number. We want minute-level freshness for the dashboard experience, but we also need an auditable trail so final invoicing can be corrected if late events or fraud decisions change counts.
Dr. Wei: So for this design, we will be explicit about the main capabilities end to end: ingest at global scale, compute billing-safe counts alongside timely click-through rate, and serve results for dashboards, application interfaces, exports, and fraud detection. As we go, we will use these scope decisions as our yardstick for tradeoffs and for what we test.
Estimate the firehose (so the design is justified)
Dr. Wei: Before we choose tools, we need a gut-check on scale: roughly how many events per second are we talking about, and how much data per day will that create?
Sam: If it is billions per day, I expect we are at least in the tens of thousands per second on average, and much higher at peak. I want to sanity-check the per-second numbers before we talk partitioning.
Dr. Wei: Sam, quick estimate: if we have about ten billion impressions per day and five hundred million clicks per day, what do you think the average impressions per second and clicks per second are? Just divide by the number of seconds in a day.
Sam: Seconds per day is about eighty six thousand four hundred. Ten billion divided by that is a bit over one hundred thousand per second for impressions, and five hundred million divided by that is a few thousand per second for clicks.
Dr. Wei: Now check the math: ten billion over eighty six thousand four hundred is about one hundred sixteen thousand impressions per second, and five hundred million over eighty six thousand four hundred is about five point eight thousand clicks per second. If your estimate was off, notice whether it was a factor of ten mistake or forgetting how many seconds are in a day.
Sam: Given that, daily volume at five hundred bytes per event is multiple terabytes per day, and peak could be five times that rate. That suggests we need enough bus capacity and consumer parallelism to absorb bursts without dropping or duplicating events.
Dr. Wei: One more checkpoint: assume each event is about five hundred bytes. Multiply by events per day to estimate daily volume, which is about five and a quarter terabytes per day of payload for ten and a half billion events, before overhead like headers and envelopes. Then remember the peak factor: during launches or daily spikes, three to five times the average can be the real design target for ingestion and buffering.
APIs and event contract (make the data usable)
Dr. Wei: In ad click analytics, the main goal is to make raw activity usable later for reliable reporting and fast decisions.
Sam: So the contract is almost as important as the pipeline. If the event fields are inconsistent, we cannot group by country or placement correctly later, and debugging billing will be painful.
Dr. Wei: That means we keep the online ad serving path fast, and we log events that carry enough context for downstream analytics.
Sam: I would expect separate log calls for impressions and clicks, and each should include a stable event id, the campaign and ad identifiers, and a timestamp that lets us bucket by minute reliably.
Dr. Wei: Concretely, our ingestion APIs like log impression and log click should emit an event contract with fields like event id, campaign id, ad id, a user id or device id, a timestamp, country, placement, and the cost model.
Sam: And for queries, we want a simple analytics endpoint that can filter by time range and group by a dimension, plus a realtime endpoint that focuses on the last few minutes. That gives the product team both historical reporting and a live dashboard feel.
Dr. Wei: When those fields are present and consistent, the analytics endpoints can group and filter cleanly, and the realtime endpoint can summarize the last few minutes without extra lookups.
Data model (raw facts vs rollups)
Dr. Wei: Before we talk about tables, think about the tension in ad analytics: you want numbers that are fast to query, but you also need a trustworthy record you can audit when something looks off.
Sam: So we need both raw facts and rollups. If a dashboard number looks wrong, we need to trace it back to individual events, but we still cannot query raw events for every page load.
Dr. Wei: A good data model separates those goals. You keep a durable trail of what actually happened, and you also keep precomputed summaries that answer the common questions in seconds.
Sam: And then we have changing context like campaign settings and fraud decisions. That feels like it should be modeled separately so we can apply adjustments without rewriting history.
Dr. Wei: On top of that, you need supporting information that changes over time, like campaign settings, fraud verdicts, and corrections. Treating these as side tables lets you update business logic without rewriting the historical event stream.
High-level architecture (hot path $+$ audit path)
Dr. Wei: In ad click analytics we usually design for two goals that naturally pull in different directions: very fresh numbers for dashboards and alerts, and provably correct numbers for billing and audits. That is why we separate the system into a hot path and an audit path, and make them agree over time.
Sam: So the hot path is what drives the near-real-time dashboard, and the audit path is what we trust for invoices. The interesting part is how we keep them consistent when events arrive late or get reprocessed.
Dr. Wei: The hot path is optimized for speed: events arrive continuously, we deduplicate and aggregate quickly, and we write rollups to an analytic store plus a small realtime cache so queries feel instantaneous.
Sam: And the cache is really about smoothing query load for the most recent window. Even if the OLAP store can answer it, caching the last minute or two can protect us from dashboard refresh storms.
Dr. Wei: The audit path is optimized for correctness and replay: we append raw logs to a durable lake, run batch reconciliation to catch late arrivals and fixes, and then backfill corrections into the same rollups. The key idea is that the hot path gives you fast answers now, and the audit path ensures the final answers are right.
Exactly-once counting (the billing core)
Dr. Wei: In billing, “exactly-once counting” means each logical billable event contributes at most once to the final rollups, even when the system is under failure and retry pressure.
Sam: When people say exactly once, I always worry it is marketing. Are we guaranteeing it end to end, or only within a single component like the stream processor?
Dr. Wei: The threat model is practical: producers can retry and emit the same logical click again, the broker can redeliver, consumers can replay after a crash, and sinks can see duplicate writes or partial commits.
Sam: So we need a stable identity for each logical event, and then we need idempotent processing and writes. If the event id is not truly stable, dedup becomes guesswork and we will either undercount or overcount.
Dr. Wei: So the identity for dedup must be immutable: dedup on a true event id from the producer, like request id, impression id, or click id; if that is unavailable, use a stronger composite like server request id plus ad render id plus a sequence number, and use time windows only for retention, not as the identity. Also, hashing that id is just an indexing choice; the real guarantee needs three conditions: a stable unique event id, a dedup retention window that covers your maximum replay and late-event horizon, and sink writes that are idempotent or transactional and aligned with the stream processor’s checkpointing.
Multi-dimensional aggregation (avoid the explosion)
Dr. Wei: In ad click analytics, every extra dimension multiplies the number of possible group-bys, and that can explode faster than your budget and latency targets.
Sam: So why not just precompute everything, so every dashboard query is instant?
Dr. Wei: Because most combinations are rarely queried. The practical approach is to materialize the common, high-value combinations, and let the long tail be answered by scanning in the OLAP store when needed.
Serving dashboards (fast, interactive, and cheap)
Dr. Wei: Now let’s focus on the serving layer for ad click analytics dashboards: users expect the page to feel instant, they want to filter and pivot freely, and we still need to keep infrastructure costs under control.
Sam: What kind of latency should we target for the common dashboard interactions? And is caching mainly about speed, or mainly about shielding the OLAP cluster from repeated refreshes?
Dr. Wei: A good mental model is that most dashboard traffic is repetitive: the same charts, the same filters, and usually a recent time window like the last hour or last day, refreshed over and over by many viewers.
Sam: So we should cache at the query result level for popular keys, like campaign and time range and group by, with a short time to live. And for huge exports, we should push them to async jobs so one user does not take down the interactive cluster.
Dr. Wei: So we serve the interactive experience with fast, slice-and-dice queries, and we add caching and background exports so that heavy requests do not slow down everyone else using the dashboard.
Fraud detection $+$ tradeoffs (what can go wrong?)
Dr. Wei: When we add fraud detection to ad click analytics, the hardest part is not just catching bad activity. The hard part is keeping our numbers trustworthy when the system is allowed to block, delay, or revise events.
Sam: If fraud scoring can revise decisions later, does that mean the dashboard and the invoice can disagree for a while? How do we explain that to an advertiser without losing trust?
Dr. Wei: The main takeaway is that dual accounting is the safety net: one set of counts optimized for fast product feedback, and another set optimized for correct billing after reconciliation. That way, fraud controls can be aggressive without silently corrupting invoices.
Sam: So we set expectations: realtime numbers are provisional, and reconciled numbers are authoritative. And we need a clear adjustment story, like what changed and when, so support and finance can audit it.
Dr. Wei: With that framing, the other tradeoffs fit underneath it. Fraud filtering means some events get labeled or filtered before rollups, so we need clear rules for what is excluded and how we recover if we made a mistake. And during spikes, we may temporarily drop non-critical dimensions or rebalance partitions, but we should preserve raw logs so reconciliation can still produce a correct final count.
Exit ticket: compute CTR and reason about dedup
Dr. Wei: For this exit ticket, we are connecting two things: a simple performance metric and the reliability of the data system that produces it.
Sam: So I should compute click-through rate and also explain what happens if clicks are duplicated or arrive late. The point is to connect the math to the failure modes.
Dr. Wei: First, think about click through rate as a ratio: how many clicks happened relative to how many impressions were shown, and what that number means for an ad campaign.
Sam: If impressions are two thousand and clicks are fifty, then click-through rate is fifty divided by two thousand, so two point five percent. If two click records arrive within five seconds with the same event id, we treat one as a duplicate, so the corrected click count is forty nine and the click-through rate drops slightly. If two click records arrive within five seconds with different event ids, they are two distinct clicks, so the corrected click count stays fifty and the click-through rate stays the same.
Dr. Wei: Then, shift to systems thinking: if the same user action can be recorded more than once, we need a clear deduplication rule, plus a clear source of truth so we can explain and correct discrepancies.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!