Levelwise
English
Data and storage

Redis

Redis is an in-memory data server that runs commands one by one. That is why it is both very fast and every command is atomic. We use it to build counters, caches, rate limits and distributed locks, as long as we know its limits.

Not reviewedWritten with AI helpReading time: 16 minOnline shop flash sale exampleC# and StackExchange.Redis code

Author: bezzad

The problem: a flash sale

At 12 noon our shop sells 1000 phones at a discount. In the first minute about 200 thousand users come and press “Buy” many times. We have a few needs:

  1. No more than 1000 items are sold. Even when thousands of requests arrive at the same moment.
  2. Each user buys only one item.
  3. Users who click again and again are limited.
  4. The nightly sales report runs only once on 3 pods.

The main database can do these jobs, but it gets slow under this load. Redis is built for exactly this kind of work.

Why is Redis fast and atomic?

Two reasons:

  1. All data is in memory. Reading and writing usually take less than one millisecond.
  2. Commands run one by one. Redis runs commands on one main Thread, one after another. So no two commands run in the middle of each other.
Pod 1Pod 2Pod 3DECRGETDECRINCRCommand queueExecutorone by oneEach command is atomicTwo commands never run at the same timeOne slow command makes everyone waitFor example, deleting a very big key
This one-by-one way is both the biggest strength of Redis and its biggest danger.

An important result: a command like DECR (decrease by one) is atomic. If 10 thousand requests call it together, each one gets a different number. None of them sees another one’s number.

But this is also the danger. If one command takes a few seconds, for example deleting a Hash with one million fields, all clients wait during this time.

Data structures

Redis is not only “key and string”. Each value is a data structure and has its own commands:

Structure Shape Example in the shop
String One value or a number Product page cache, stock counter with INCR and DECR
Hash Several fields under one key Shopping cart: each field is an item, its value is the count
List A list in order of arrival The last items the user has seen
Set A set with no duplicates Users who bought in the flash sale
Sorted Set A set with scores, sorted Today’s best-selling items
Stream A log of messages with consumer groups Simple events between services

Choosing the right structure is important. For example, for “best sellers”, with a Sorted Set you do not need to read all items and sort them in the app. Redis keeps them sorted.

Reducing stock atomically

Now needs 1 and 2 together: stock goes down and each user buys only once. If you write these two with two separate commands, another request may come in between. For example, the stock goes down, but then it turns out the user was a repeat. One phone is lost.

The solution is a Lua script. Redis runs the whole script like one command, at once and with no break:

public sealed class FlashSale(IConnectionMultiplexer redis)
{
    // Returns 1 = bought, 0 = sold out, -1 = this user already bought
    private const string BuyScript = """
        if redis.call('SISMEMBER', KEYS[2], ARGV[1]) == 1 then return -1 end
        if tonumber(redis.call('GET', KEYS[1]) or '0') <= 0 then return 0 end
        redis.call('DECR', KEYS[1])
        redis.call('SADD', KEYS[2], ARGV[1])
        return 1
        """;

    public async Task<int> TryBuyAsync(int saleId, string userId)
    {
        var db = redis.GetDatabase();
        var tag = "{sale" + saleId + "}";
        var result = await db.ScriptEvaluateAsync(BuyScript,
            new RedisKey[] { tag + ":stock", tag + ":buyers" },
            new RedisValue[] { userId });
        return (int)result;
    }
}

A few points:

  1. Create the ConnectionMultiplexer object once and register it as a Singleton. Creating a connection for each request is slow and expensive.
  2. Redis is not the main source of truth. Also save each successful purchase in the database through a queue. If Redis loses data, the stock can be calculated again.
  3. The part inside braces in the key names is explained below, in the Cluster section.

Rate limiting

Need 3: each user can make at most 100 requests per minute. The built-in ASP.NET Core rate limiter keeps the counter in the memory of each pod. With 10 pods, the real limit becomes 1000. So the counter must be shared:

var key = $"rate:{userId}:{DateTime.UtcNow:yyyyMMddHHmm}";
var count = await db.StringIncrementAsync(key);
if (count == 1)
    await db.KeyExpireAsync(key, TimeSpan.FromMinutes(1));

if (count > 100)
    return Results.StatusCode(StatusCodes.Status429TooManyRequests);

This is the simplest way (Fixed Window). It has two weak points:

  1. It is not atomic. If the app dies between INCR and EXPIRE, the key stays with no expiry. For real code, put both in one Lua script.
  2. The window edge. A user can send 100 requests in the last second of one minute and 100 requests in the first second of the next minute. The Sliding Window and Token Bucket methods do not have this problem.

Distributed lock

Need 4: the nightly report runs on 3 pods. Each pod has its own scheduler. So each customer gets 3 emails. One way is a lock in Redis:

var token = Guid.NewGuid().ToString();
if (await db.LockTakeAsync("lock:daily-report", token, TimeSpan.FromMinutes(5)))
{
    try
    {
        await reports.SendDailyAsync(ct);
    }
    finally
    {
        await db.LockReleaseAsync("lock:daily-report", token);
    }
}

Why do we give each lock a unique value (token)? Because the lock is released only when the value still belongs to us. Otherwise we may release a lock that now belongs to another pod.

Why is an expiry time needed? Because if the pod dies in the middle of the work, the lock must free itself. But this same expiry has a danger:

Pod APod B0 s5303145Has the lockStopped, for example a long pauseWoke up, thinks it has lockTook the lock and worksLock expiredBoth work togetherA longer expiry only lowers the chance. The job itself must handle repeats.
Pod A does not know it was stopped. After it wakes up, it thinks it still has the lock.

So a Redis lock is not a full guarantee. For important work:

  1. Make the work itself repeatable (Idempotent). For example, for each customer and each date, save “email sent”. Running it again does no harm.
  2. Use a Fencing Token. Each time the lock is taken, also get an increasing number. The database rejects a write whose number is lower than the last number it has seen.
  3. If you can, remove the lock. For example, move the nightly report to a CronJob in Kubernetes, so you do not have three copies at all.

Durability and failure

Redis keeps data in memory. What if it restarts?

  1. Periodic snapshot (RDB). Every so often, all data is written to disk. Changes after the last snapshot are lost.
  2. Write log (AOF). Each write command is added to a file. Less data is lost, but it is a bit slower.
  3. Replica and Sentinel. We have one or more copies. If the main server dies, Sentinel makes a copy the main one. Because copying is asynchronous, the last writes may be lost.

When memory is full, Redis behavior depends on the maxmemory-policy setting. For a cache, people usually choose a policy like allkeys-lru so that rarely used keys are removed. For data that must not be removed (like stock), automatic removal is dangerous.

Redis Cluster

When one server is not enough, Redis Cluster spreads the data across several nodes:

  1. The key space is split into 16384 parts (Hash Slots). Each node owns some of these parts.
  2. For each key, Redis takes a hash of its name and finds its part.
  3. Multi-key commands (like a Lua script, a MULTI transaction and MGET) work only when all keys are in one part. Otherwise you get a CROSSSLOT error.
Node 1slots 0 - 5460Node 2slots 5461 - 10922Node 3slots 10923 - 16383stock:p1stock:p2stock:p3stock:{sale42}:p1..p3Three keys on three nodesA multi-key script failsThe part in braces is the sameso all three keys are in one slotCROSSSLOThash tagA key is always on one node. More nodes do not spread the load of one hot key.

If part of the key name is inside braces, only that part is used to find the place of the key. This is called a Hash Tag. That is why in the flash sale code, both keys have the shared part “sale42” inside braces.

Two more problems in a Cluster:

  1. Hot Key. All requests read the “home page settings” key. This key is on one node, so all the load is on that same node. A simple fix: a local cache for a few seconds in each pod.
  2. Big Key. A Hash or Set with millions of members makes work on it slow and makes everyone wait. Keep keys small. To delete, use UNLINK instead of DEL, which frees memory in the background. To read, use HSCAN instead of HGETALL.

Common mistakes

Mistake Result The right way
Read the stock, check it in the app, then decrease it Selling more than the stock An atomic command or a Lua script
Creating a new connection for each request Slowness and running out of connections One shared ConnectionMultiplexer
A lock with no expiry time When the pod dies, the lock stays forever An expiry, together with repeatable work
Releasing a lock without checking the value Releasing another pod’s lock A unique value for each lock
Redis as the only place for important data Data is lost on restart The database is the source of truth, Redis is a helper
Heavy commands on a big key All clients wait Small keys, UNLINK and SCAN
Thinking Cluster spreads the load of one key One hot node and the rest idle A local cache or several copies of the key

Summary in six lines

  1. Redis keeps data in memory and runs commands one by one. So it is fast and atomic.
  2. One slow command or one big key makes all clients wait.
  3. Choose the right data structure: counter, Hash, Set, Sorted Set.
  4. Make several dependent steps atomic with one Lua script.
  5. A Redis lock has an expiry and is not a full guarantee. Make the work repeatable.
  6. In a Cluster, the keys of one script must be in one part. Calm a hot key with a local cache.