F1 Redis — Crash Course to Advanced Patterns
Loading learning experience...
Lecture transcript
Read the narration for F1: Redis — Crash Course to Advanced Patterns
From consistent hashing to Redis Cluster
Dr. Wei: If you already know the idea of sharding, here’s how Redis Cluster applies it in practice.
Sam: So if I already know consistent hashing, what should I watch for as the main difference in Redis Cluster?
Dr. Wei: You may have seen consistent hashing described as a ring: both keys and nodes land on the ring, and when a node is added or removed, only a small fraction of keys need to move.
Sam: So instead of the ring moving keys around implicitly, Redis makes the movement explicit by moving slot ranges?
Dr. Wei: Redis Cluster does not use the classic consistent hashing ring. Instead, it uses 16,384 fixed hash slots: each key maps to a slot using CRC 16, and then each slot is assigned to a particular node.
Sam: Does that mean the number of slots stays constant, and scaling is just changing which node owns which slots?
Dr. Wei: The practical motivation is operational simplicity: rebalancing is done by migrating slot ranges between nodes, which makes movement explicit and predictable while still keeping resharding manageable.
Why Redis shows up in so many system designs
Dr. Wei: Before we talk features, think about the kind of problems that make a system feel slow or unreliable: waiting on repeated reads, coordinating many workers, or handling sudden traffic spikes. A common pattern is to keep the hottest, most time sensitive data closer to the application so responses feel instant.
Sam: In an interview, should I pitch Redis as just a cache, or as a general-purpose database?
Dr. Wei: Pitch it as a fast, in memory data store that you can use for caching, shared state, and coordination. For example, if you store login sessions in Redis, you can validate requests quickly without hitting your primary database on every page view.
Sam: And the moment I say sessions, I should also mention time to live and what happens on eviction, right?
Dr. Wei: The tradeoff is that speed comes with constraints: keeping lots of sessions in memory can get expensive, and you must be deliberate about what happens when keys expire or are evicted under memory pressure. In this course we will learn when Redis is the right fit, and how to use it safely in real system designs.
Mental model: an event loop executing commands
Dr. Wei: Let’s build a concrete mental model for how Redis runs your requests, because it’s the key to why some operations feel “instant” and others suddenly slow down.
Sam: When people say Redis is single-threaded, is that the reason it is fast, or the reason it gets slow sometimes?
Dr. Wei: Imagine a simple cache flow: a client sends GET for user:42, Redis misses, then the app computes the value and sends SET user:42 value with a time to live of 60 seconds. Those are separate commands arriving one after another.
Sam: So the danger is when one of those commands is expensive, everyone else queues behind it in the same line?
Dr. Wei: Now here’s the event loop idea: Redis reads one command, executes it to completion, sends the reply, then moves on to the next command. That single file line of execution is why each individual command is atomic, but it also explains performance cliffs when one command takes a long time and everyone else waits in line.
Core structures that interviewers love
Dr. Wei: Redis is most impressive when you treat it like a data-structure service, not just a simple key value cache.
Sam: Which few structures are the highest signal to mention in an interview without turning it into a laundry list?
Dr. Wei: In interviews, that framing matters because it shows you can model a problem using the right structure, and you understand the tradeoffs in speed, memory, and correctness.
Sam: So I should pick a few and tie each one to a concrete feature, like counters for quotas or sorted sets for leaderboards?
Dr. Wei: On this slide, we will focus on the core Redis structures that come up again and again, and the kinds of product features they naturally unlock, like leaderboards, counters, queues, and real time feeds.
Advanced structures: big leverage, small memory
Dr. Wei: These are the patterns that separate "cache user profile" from "design a feature."
Sam: Is the main idea here that some structures compress better, or that they let you answer different questions faster?
Dr. Wei: In Redis, the real leverage comes from choosing the right structure for the question you are asking: not just storing data, but shaping how fast you can answer, how safely you can update, and how much memory you pay to do it.
Sam: If I only remember one rule for interviews, is it basically: state the query you need, then pick the Redis structure that makes that query cheap?
Dr. Wei: As we move into the advanced options, listen for the tradeoffs: what you gain in speed or simplicity, what you give up in precision or flexibility, and what assumptions your design is making about scale and access patterns.
TTL and eviction: the cache is a policy engine
Dr. Wei: When you use Redis as a cache, it is not just storing data for you. It is enforcing rules about how long data is allowed to live and what should be thrown away when memory gets tight. That means your cache behaves like a policy engine, not a simple bucket of key value pairs.
Sam: So the cache is making decisions for me, not just holding onto things until I delete them?
Dr. Wei: Definition first: a time to live, or TTL, is a per-key expiration timer, and eviction is what Redis does when memory hits a configured limit. Expiration is about freshness; eviction is about survival under pressure.
Sam: Got it. TTL is the clock on a key, and eviction is the emergency cleanup when memory is full.
Dr. Wei: Concrete scenario. Suppose maxmemory is 100 megabytes and the eviction policy is all keys least recently used, but keep in mind Redis implements this as an approximation, so this example is an intuition model about what is more likely to be evicted. We currently have three keys: A was accessed 10 seconds ago, B was accessed 3 seconds ago, and C was accessed 1 second ago. You now read A, and immediately after that you write a new key D that pushes memory over the limit. Which key gets evicted?
Sam: I think it evicts B, because after reading A, B becomes the least recently used among A, B, and C.
Dr. Wei: Answer: B is the one that disappears. That matches the least recently used rule of thumb: reading A makes it recently used, and C was already the most recent, so B becomes the least recently used at the moment we need to free space. In real Redis, least recently used is approximate, but the direction still holds: recently accessed keys are less likely to be evicted than cold ones.
Sam: So even a single read can reshuffle what counts as cold, and that changes what is likely to get evicted next.
Dr. Wei: Tradeoff and interview takeaway. TTL adds another policy layer: a key can expire even if it is popular, and a key can be evicted even if it still has time left. So in design answers, explicitly name both: your freshness target, and your eviction behavior under memory pressure.
Persistence: RDB vs. AOF vs. hybrid
Dr. Wei: When Redis is more than a throwaway cache, persistence becomes the big question: what gets saved to disk, and how fast you can recover after a crash.
Sam: So this is basically about what Redis keeps on disk and what happens after a restart?
Dr. Wei: Definition: RDB is periodic snapshotting, AOF is an append-only log of write commands, and a hybrid setup combines them. They differ mainly in write overhead, restart time, and worst-case data loss.
Sam: If I want the smallest possible data loss, I should lean toward the append-only log, right?
Dr. Wei: Concrete scenario. If you snapshot every five minutes and the machine dies, you can lose up to five minutes of writes. With an append-only log that is synced once per second, you usually lose about a second, but you pay more ongoing disk work.
Sam: If I pick append-only log for safety, should I also call out the longer restart time because it has to replay the log?
Dr. Wei: Tradeoff and interview takeaway. Snapshotting is compact and can restart quickly, but has a larger loss window; the append-only log is safer for recent writes, but can slow down writes and can take longer to replay on restart. In interviews, state your acceptable data loss window and your recovery time goal, then pick the persistence mode that matches.
Replication and HA: reads scale, consistency weakens
Dr. Wei: Replication is how we scale reads and improve availability, but it usually means we trade away some consistency. With async replication, updates do not reach replicas instantly, so clients can observe stale data and failover can surface surprising edge cases.
Sam: If I read from replicas for scale, should I assume I might see stale data unless I add something like read your writes behavior in the app?
Sam: So what can go wrong in a concrete request flow?
Dr. Wei: Concrete scenario: a client writes value B to the primary and gets an acknowledgment. Immediately after, the same client reads from a replica. What value might it see: the old value A, or the new value B?
Sam: I would hope it sees B, but if replication is asynchronous, I guess it could still see A for a moment.
Dr. Wei: Exactly. Reading from a replica right after an acknowledged write can return B if the update has arrived, or A if the replica is lagging. That stale A case is the consistency cost of async replication.
Dr. Wei: Now the tradeoff case. If the primary dies before that write replicates and a replica is promoted, that acknowledged write can be missing after failover. The interview takeaway is to say this out loud: read replicas can be stale, and acknowledged writes may not be durable unless you add stronger guarantees.
Redis Cluster: sharding with hash slots
Dr. Wei: Cluster mode is Redis’s built-in way to scale out by sharding keys across multiple nodes, but it comes with important routing constraints you need to design around.
Sam: So in cluster mode, the main idea is that keys get split up across nodes, and that affects what operations you can safely do?
Dr. Wei: Definition: Redis Cluster splits the keyspace into 16,384 hash slots, and each primary owns a range of slots. A key maps to exactly one slot, and that slot maps to exactly one primary at a time.
Sam: If a key maps to a single slot, does the client have to know where that slot lives, or does Redis figure it out?
Dr. Wei: Concrete scenario: you have user colon 42 profile and user colon 42 timeline. If those two keys land on different slots, a multi-key operation cannot run atomically across them, because it would require talking to two primaries.
Sam: So the problem is not that multi-key commands never work, but that they only work when the keys end up in the same slot?
Dr. Wei: Tradeoff and interview takeaway. A client figures out which slot a key maps to, and it must send that command to the node that owns the slot. Multi-key operations only work when the keys land in the same slot, typically by using a shared hash tag pattern in the key names. In interviews, call out this constraint early and show how you shape keys to match required atomicity.
Pattern: rate limiting beyond INCR
Dr. Wei: Rate limiting is about correctness under concurrency, not just counting. When lots of clients hit your service at the same time, a simple counter can drift from the behavior you actually want: allowing a fixed number of requests per user or per key in a defined time window.
Sam: So the risk is that a plain counter might be accurate as a number, but still wrong for the policy we mean to enforce?
Dr. Wei: Definition: a good limiter answers one question atomically: should this request be allowed right now? The limiter also needs a clear time model, like fixed window, sliding window, or token bucket.
Sam: When you say a time model, is that basically choosing how we interpret the minute so it feels fair to users?
Dr. Wei: Concrete scenario: say the rule is 100 requests per minute per user. With a naive counter, retries and clock skew can let some users exceed the limit or get blocked unfairly at window boundaries.
Sam: Ah, so you can get bursts right at the boundary, or accidental blocks if clocks disagree or clients retry.
Dr. Wei: Tradeoff and interview takeaway. Stronger designs make the decision and the update in one step, and choose a windowing strategy that matches product expectations. In interviews, say what fairness means, name your chosen limiter type, and mention one pitfall you are avoiding, like boundary bursts or non-atomic updates.
Pattern: leaderboards and real-time $top-k$
Dr. Wei: In many apps, you need to keep a live leaderboard where rankings change constantly, and you want to fetch the top k results quickly without doing heavy work on every request.
Sam: What makes this hard in practice, is it the sorting, the updates, or both?
Dr. Wei: Definition: a sorted set stores members with scores and keeps them ordered by score, so you can update a single user’s score and still query the top ranks efficiently.
Sam: When scores change a lot, do I need to worry about write amplification or hot keys for top users?
Dr. Wei: Concrete scenario and takeaway. If a user gains 50 points, you update just that member’s score, then read back the top k for the page. The tradeoff is that this is great for rankings, but you must define how you handle ties, resets, and trimming history. In interviews, say why sorted sets beat re-sorting in your database, and name one detail you would clarify, like tie-breaking or time windows.
Pattern: sessions, locks, and messaging
Dr. Wei: Redis is often the "glue" between stateless services and real-time features.
Sam: So Redis sits in the middle, helping services coordinate without each service storing a lot of state?
Dr. Wei: Definition: sessions are time-limited state, locks are short-lived mutual exclusion, and messaging is lightweight event delivery. Redis is attractive here because these are small, fast pieces of state that benefit from low latency and atomic updates.
Sam: If I talk about locks, should I immediately warn about failure cases like timeouts and retries, so it does not sound too hand-wavy?
Dr. Wei: Concrete scenario: you run a job worker fleet and want only one worker to process order colon 123. A lock with a short expiration prevents duplicates, but if the expiration is too short you can double-process, and if it is too long you can stall progress after a crash.
Sam: So the hard part is picking an expiration that does not cause duplicates, but also does not leave the system stuck if a worker dies.
Dr. Wei: Tradeoff and interview takeaway. Locks are deceptively tricky: you must reason about timeouts, retries, and failure, not just mutual exclusion. In interviews, explicitly mention a safety property you want, like no two workers act at once, and a liveness property, like work still completes after failures, and then describe how you use expirations and ownership checks to balance both.
Failure modes and performance gotchas
Dr. Wei: Before we talk about specific settings or commands, we need a mental model for what can go wrong when Redis is under pressure. This slide is about the typical ways performance degrades, and the kinds of failures you might see in production.
Sam: When you say failure modes, do you mean Redis crashing, or just getting slow?
Dr. Wei: Both. Sometimes it is a hard failure like an out of memory kill, a restart, or losing writes you assumed were durable. More often it is a soft failure: latency spikes, timeouts, blocked clients, replication lag, or an event loop that cannot keep up. The key habit is to name these risks up front so you can design around them with limits, backpressure, and realistic expectations.
Redis vs Memcached (and what levels sound like)
Dr. Wei: Choosing the simplest correct tool is a core interview competency.
Sam: If I say “Redis is richer but Memcached is simpler,” what extra details would push that to a senior-level answer?
Dr. Wei: When you compare Redis and Memcached, the goal is not to name features; it is to show judgment about tradeoffs like durability, data structures, operations, and operational complexity.
Sam: So I should mention things like persistence, replication story, and data structure support, but tie each one back to the workload and operational risk?
Dr. Wei: We will also practice what different answer levels sound like: a junior answer states what each tool is, a mid level answer ties that to workload needs, and a senior answer adds failure modes, scaling strategy, and why the simpler option might still win.
Exit ticket: shard a global rate limiter
Dr. Wei: Exit ticket time: imagine you need a single global rate limit for an API, but the traffic volume is too high for one Redis key to handle comfortably. Your goal is to shard the limiter while keeping the user experience consistent.
Sam: Do you want a design that is strictly correct, or one that is approximately fair but operationally simpler?
Dr. Wei: Pick a concrete example, like one million requests per minute globally, and choose a number of shards. Do the quick math to translate that into a per shard budget, and think about how you will distribute requests across shards.
Sam: If I do, say, 100 shards, then each shard is budgeted for about 10,000 requests per minute. But how do I avoid one shard getting hotter than the others?
Dr. Wei: Then sanity-check correctness and operability: what happens when traffic is uneven, when a shard is slow or unavailable, or when you need to change the shard count? State one technique to reduce unfairness and one technique to keep the system easy to monitor and debug.
Thank you for watching!
Thanks for watching. Subscribe and share if you found this useful—see you next time!