Redis
Redis is an in-memory data server that runs commands one by one. That is why it is both very fast and every command is atomic. We use it to build counters, caches, rate limits and distributed locks, as long as we know its limits.
Author: bezzad
The problem: a flash sale
At 12 noon our shop sells 1000 phones at a discount. In the first minute about 200 thousand users come and press “Buy” many times. We have a few needs:
- No more than 1000 items are sold. Even when thousands of requests arrive at the same moment.
- Each user buys only one item.
- Users who click again and again are limited.
- The nightly sales report runs only once on 3 pods.
The main database can do these jobs, but it gets slow under this load. Redis is built for exactly this kind of work.
Why is Redis fast and atomic?
Two reasons:
- All data is in memory. Reading and writing usually take less than one millisecond.
- Commands run one by one. Redis runs commands on one main Thread, one after another. So no two commands run in the middle of each other.
An important result: a command like DECR (decrease by one) is atomic. If 10 thousand requests call it together, each one gets a different number. None of them sees another one’s number.
But this is also the danger. If one command takes a few seconds, for example deleting a Hash with one million fields, all clients wait during this time.
Data structures
Redis is not only “key and string”. Each value is a data structure and has its own commands:
| Structure | Shape | Example in the shop |
|---|---|---|
| String | One value or a number | Product page cache, stock counter with INCR and DECR |
| Hash | Several fields under one key | Shopping cart: each field is an item, its value is the count |
| List | A list in order of arrival | The last items the user has seen |
| Set | A set with no duplicates | Users who bought in the flash sale |
| Sorted Set | A set with scores, sorted | Today’s best-selling items |
| Stream | A log of messages with consumer groups | Simple events between services |
Choosing the right structure is important. For example, for “best sellers”, with a Sorted Set you do not need to read all items and sort them in the app. Redis keeps them sorted.
Reducing stock atomically
Now needs 1 and 2 together: stock goes down and each user buys only once. If you write these two with two separate commands, another request may come in between. For example, the stock goes down, but then it turns out the user was a repeat. One phone is lost.
The solution is a Lua script. Redis runs the whole script like one command, at once and with no break:
public sealed class FlashSale(IConnectionMultiplexer redis)
{
// Returns 1 = bought, 0 = sold out, -1 = this user already bought
private const string BuyScript = """
if redis.call('SISMEMBER', KEYS[2], ARGV[1]) == 1 then return -1 end
if tonumber(redis.call('GET', KEYS[1]) or '0') <= 0 then return 0 end
redis.call('DECR', KEYS[1])
redis.call('SADD', KEYS[2], ARGV[1])
return 1
""";
public async Task<int> TryBuyAsync(int saleId, string userId)
{
var db = redis.GetDatabase();
var tag = "{sale" + saleId + "}";
var result = await db.ScriptEvaluateAsync(BuyScript,
new RedisKey[] { tag + ":stock", tag + ":buyers" },
new RedisValue[] { userId });
return (int)result;
}
}
A few points:
- Create the ConnectionMultiplexer object once and register it as a Singleton. Creating a connection for each request is slow and expensive.
- Redis is not the main source of truth. Also save each successful purchase in the database through a queue. If Redis loses data, the stock can be calculated again.
- The part inside braces in the key names is explained below, in the Cluster section.
Rate limiting
Need 3: each user can make at most 100 requests per minute. The built-in ASP.NET Core rate limiter keeps the counter in the memory of each pod. With 10 pods, the real limit becomes 1000. So the counter must be shared:
var key = $"rate:{userId}:{DateTime.UtcNow:yyyyMMddHHmm}";
var count = await db.StringIncrementAsync(key);
if (count == 1)
await db.KeyExpireAsync(key, TimeSpan.FromMinutes(1));
if (count > 100)
return Results.StatusCode(StatusCodes.Status429TooManyRequests);
This is the simplest way (Fixed Window). It has two weak points:
- It is not atomic. If the app dies between INCR and EXPIRE, the key stays with no expiry. For real code, put both in one Lua script.
- The window edge. A user can send 100 requests in the last second of one minute and 100 requests in the first second of the next minute. The Sliding Window and Token Bucket methods do not have this problem.
Distributed lock
Need 4: the nightly report runs on 3 pods. Each pod has its own scheduler. So each customer gets 3 emails. One way is a lock in Redis:
var token = Guid.NewGuid().ToString();
if (await db.LockTakeAsync("lock:daily-report", token, TimeSpan.FromMinutes(5)))
{
try
{
await reports.SendDailyAsync(ct);
}
finally
{
await db.LockReleaseAsync("lock:daily-report", token);
}
}
Why do we give each lock a unique value (token)? Because the lock is released only when the value still belongs to us. Otherwise we may release a lock that now belongs to another pod.
Why is an expiry time needed? Because if the pod dies in the middle of the work, the lock must free itself. But this same expiry has a danger:
So a Redis lock is not a full guarantee. For important work:
- Make the work itself repeatable (Idempotent). For example, for each customer and each date, save “email sent”. Running it again does no harm.
- Use a Fencing Token. Each time the lock is taken, also get an increasing number. The database rejects a write whose number is lower than the last number it has seen.
- If you can, remove the lock. For example, move the nightly report to a CronJob in Kubernetes, so you do not have three copies at all.
Durability and failure
Redis keeps data in memory. What if it restarts?
- Periodic snapshot (RDB). Every so often, all data is written to disk. Changes after the last snapshot are lost.
- Write log (AOF). Each write command is added to a file. Less data is lost, but it is a bit slower.
- Replica and Sentinel. We have one or more copies. If the main server dies, Sentinel makes a copy the main one. Because copying is asynchronous, the last writes may be lost.
When memory is full, Redis behavior depends on the maxmemory-policy setting. For a cache, people usually choose a policy like allkeys-lru so that rarely used keys are removed. For data that must not be removed (like stock), automatic removal is dangerous.
Redis Cluster
When one server is not enough, Redis Cluster spreads the data across several nodes:
- The key space is split into 16384 parts (Hash Slots). Each node owns some of these parts.
- For each key, Redis takes a hash of its name and finds its part.
- Multi-key commands (like a Lua script, a MULTI transaction and MGET) work only when all keys are in one part. Otherwise you get a CROSSSLOT error.
If part of the key name is inside braces, only that part is used to find the place of the key. This is called a Hash Tag. That is why in the flash sale code, both keys have the shared part “sale42” inside braces.
Two more problems in a Cluster:
- Hot Key. All requests read the “home page settings” key. This key is on one node, so all the load is on that same node. A simple fix: a local cache for a few seconds in each pod.
- Big Key. A Hash or Set with millions of members makes work on it slow and makes everyone wait. Keep keys small. To delete, use UNLINK instead of DEL, which frees memory in the background. To read, use HSCAN instead of HGETALL.
Common mistakes
| Mistake | Result | The right way |
|---|---|---|
| Read the stock, check it in the app, then decrease it | Selling more than the stock | An atomic command or a Lua script |
| Creating a new connection for each request | Slowness and running out of connections | One shared ConnectionMultiplexer |
| A lock with no expiry time | When the pod dies, the lock stays forever | An expiry, together with repeatable work |
| Releasing a lock without checking the value | Releasing another pod’s lock | A unique value for each lock |
| Redis as the only place for important data | Data is lost on restart | The database is the source of truth, Redis is a helper |
| Heavy commands on a big key | All clients wait | Small keys, UNLINK and SCAN |
| Thinking Cluster spreads the load of one key | One hot node and the rest idle | A local cache or several copies of the key |
Summary in six lines
- Redis keeps data in memory and runs commands one by one. So it is fast and atomic.
- One slow command or one big key makes all clients wait.
- Choose the right data structure: counter, Hash, Set, Sorted Set.
- Make several dependent steps atomic with one Lua script.
- A Redis lock has an expiry and is not a full guarantee. Make the work repeatable.
- In a Cluster, the keys of one script must be in one part. Calm a hot key with a local cache.