Levelwise
English
ASP.NET Core

Rate Limiting

Each customer can send only a set number of requests in a period of time. An extra request is rejected with the 429 code. The built-in ASP.NET Core limiter counts in the memory of each server, so on several servers we need a shared counter.

Not reviewedWritten with AI helpReading time: 14 minExample of the shop's public APIC# and .NET 10 code

Author: bezzad

The problem: one customer slows down everyone

Our shop has a public API. Partner shops read prices and stock with it. By contract, each partner is allowed at most 100 requests per minute.

One day, a script of one of the partners breaks and sends 5000 requests per minute in a loop. The result:

  1. The database gets busy. All queries become slow.
  2. The other partners and the site’s customers also become slow. It is not their fault.
  3. The server cost goes up. For work that has no value.

Rate Limiting means we set a limit for each customer. A request above the limit is rejected before it reaches the main code.

The right answer: the 429 code

When a request is rejected, the right answer is this:

  • The 429 status code, with the name Too Many Requests.
  • The Retry-After header. It tells the client how many seconds later to try again.

With these two, a good client knows what to do. It waits and sends again, instead of trying again and again without a break.

Four algorithms

The ASP.NET Core framework has four ready-made limiters:

  1. Fixed Window. It splits time into one-minute periods. In each period, it counts up to 100. At the start of the next period, the counter goes back to zero. It is the simplest method.
  2. Sliding Window. It splits each period into several small pieces. The count drops the old pieces slowly. It is more exact than the fixed window.
  3. Token Bucket. Each customer has a bucket of tokens. The bucket fills at a fixed speed. Each request takes one token. If the bucket is empty, the request is rejected.
  4. Concurrency. It does not count time. It only says “at most 10 requests at the same time”. It is good for heavy endpoints, like building a report.

The fixed window problem: the window edge

← TimeFirst minuteLimit: 100 requestsSecond minuteLimit: 100 requests100 in the last second100 in the first second200 requests in 2 seconds
The customer sends 100 requests in the last second of one minute and 100 more requests in the first second of the next minute. Both windows followed the rule, but 200 requests arrived in two seconds.

The sliding window and the token bucket do not have this problem. The token bucket has one more benefit: it allows a little burst (Burst). This means if the customer has been quiet for a while, their bucket is full, and they can send several requests one after another.

Live example: token bucket

The bucket in this example has at most 5 tokens. One token is added every second. Send several requests one after another and see what happens:

Token bucket: at most 5 tokens, one token per second
Accepted: 0
Rejected with 429: 0

    Code: a limit for each customer

    The most important decision is the key of the limit. That is, for whom is the count kept separate? We count separately for each partner. We call each key a Partition.

    builder.Services.AddRateLimiter(options =>
    {
        // The default is 503. 429 tells the client "slow down", not "server is broken".
        options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
    
        options.AddPolicy("per-customer", context =>
        {
            // One bucket per customer. Fall back to IP only for anonymous calls.
            var key = context.User.FindFirstValue("sub")
                      ?? context.Connection.RemoteIpAddress?.ToString()
                      ?? "unknown";
    
            return RateLimitPartition.GetTokenBucketLimiter(key, _ => new TokenBucketRateLimiterOptions
            {
                TokenLimit = 20,                               // the biggest burst
                TokensPerPeriod = 10,
                ReplenishmentPeriod = TimeSpan.FromSeconds(6), // 10 tokens every 6 s = 100 per minute
                QueueLimit = 0,                                // reject at once, do not wait
                AutoReplenishment = true
            });
        });
    
        options.OnRejected = (context, ct) =>
        {
            if (context.Lease.TryGetMetadata(MetadataName.RetryAfter, out var retryAfter))
                context.HttpContext.Response.Headers.RetryAfter =
                    ((int)retryAfter.TotalSeconds).ToString();
    
            return ValueTask.CompletedTask;
        };
    });
    
    var app = builder.Build();
    
    app.UseRouting();
    app.UseAuthentication(); // the limiter needs to know the user
    app.UseRateLimiter();
    
    app.MapGet("/products", (ShopDb db, CancellationToken ct) => db.Products.ToListAsync(ct))
       .RequireRateLimiting("per-customer");

    Some points about this code:

    1. The default rejection code is 503. You must change it to 429 yourself. The 503 code means “the server is broken”, and it misleads the client.
    2. The limiter comes after authentication. Otherwise, it does not know yet who the user is, and it counts everyone by IP.
    3. The key is the customer ID, not the IP. Several customers behind one NAT have one IP. One customer may also have several IPs.
    Queue or reject? With the QueueLimit setting, an extra request can wait a little until a token arrives, instead of being rejected right away. For a public API, rejecting right away is usually better. A queue means connections and server memory are held for a longer time.

    The big problem: several servers

    In Production, the service runs on 10 pods. A Load Balancer spreads the requests between them. Now what is the real limit?

    1. The built-in limiter keeps the counter in the memory of that same pod.
    2. Each pod counts up to 100 on its own. It does not know about the others.
    3. So the real limit is 10 × 100 = 1000. On a laptop with one copy everything was correct, but not in Production.
    Counter in each pod's memoryLoad Balancerpod 1100pod 2100pod 10100...Real limit: 1000 per minuteShared counterpod 1pod 2pod 10Rediscustomer:42:10:30 = 87Real limit: 100 per minute
    Right side: each pod has its own counter. Left side: all pods read one shared counter.

    The options:

    1. A shared counter in Redis. The simplest version is a fixed window. The key is made from the customer ID and the current minute. We add one with the INCR command, and the first time we set a 60-second expiration with the EXPIRE command. Put these two in a Lua script so they run together and atomically. If the number goes over 100, return 429.
    2. A limit at the edge of the system. An API Gateway or Ingress in front of all pods does the counting. The app code does not change.
    3. An approximate way. Set each pod’s share to 100 ÷ 10 = 10. It is simple, but it must change when the number of pods changes. It is also not exact if the Load Balancer does not spread the load evenly.

    What if Redis goes down?

    It depends on what the limit is for:

    • To protect the system: usually accept the request (Fail-open). Together with a simple local limit as a backup. We do not want a Redis outage to take down the whole API.
    • For security or cost: maybe rejecting (Fail-closed) is right. For example, a limit that stops password guessing.

    Important rules

    1. Return the 429 code and the Retry-After header. Not 503, not 500.
    2. Choose the key correctly. Customer ID or API Key, not only IP.
    3. On several servers, the counter must be shared. Either in Redis, or in the Gateway.
    4. Choose the algorithm by need. To allow a short burst, token bucket. For a heavy endpoint, concurrency.
    5. Set a separate limit for sensitive endpoints. For example, login and sending SMS codes need a much lower limit.
    6. Measure the rejections. If a good customer gets 429 all the time, maybe the limit is too low.

    Common mistakes

    Mistake Result The right way
    Built-in limiter on several pods The real limit becomes several times the contract. A shared counter in Redis or the Gateway.
    Key is only the IP Customers behind one NAT are rejected together. Customer ID or API Key.
    Default 503 code The client thinks the server is broken. Set the code to 429.
    No Retry-After header The client tries again and again, and the load grows. Fill Retry-After in OnRejected.
    Limiter before authentication The user is not known yet, and everyone is counted by IP. Authentication first, then the limiter.
    INCR and EXPIRE separate from each other If an error happens between the two, the key stays with no expiration. Both in one Lua script.

    Summary in six lines

    1. Rate limiting protects the system and the other customers.
    2. An extra request is rejected with the 429 code and the Retry-After header.
    3. The fixed window is simple, but at the window edge it allows double. The token bucket controls short bursts.
    4. The limit key is the customer ID, not the IP.
    5. The built-in limiter counts in the memory of each pod. On several pods, a shared counter is needed.
    6. If the shared counter goes down, decide between accept and reject based on the goal of the limit.