Rate Limiting
Each customer can send only a set number of requests in a period of time. An extra request is rejected with the 429 code. The built-in ASP.NET Core limiter counts in the memory of each server, so on several servers we need a shared counter.
Author: bezzad
The problem: one customer slows down everyone
Our shop has a public API. Partner shops read prices and stock with it. By contract, each partner is allowed at most 100 requests per minute.
One day, a script of one of the partners breaks and sends 5000 requests per minute in a loop. The result:
- The database gets busy. All queries become slow.
- The other partners and the site’s customers also become slow. It is not their fault.
- The server cost goes up. For work that has no value.
Rate Limiting means we set a limit for each customer. A request above the limit is rejected before it reaches the main code.
The right answer: the 429 code
When a request is rejected, the right answer is this:
- The 429 status code, with the name Too Many Requests.
- The Retry-After header. It tells the client how many seconds later to try again.
With these two, a good client knows what to do. It waits and sends again, instead of trying again and again without a break.
Four algorithms
The ASP.NET Core framework has four ready-made limiters:
- Fixed Window. It splits time into one-minute periods. In each period, it counts up to 100. At the start of the next period, the counter goes back to zero. It is the simplest method.
- Sliding Window. It splits each period into several small pieces. The count drops the old pieces slowly. It is more exact than the fixed window.
- Token Bucket. Each customer has a bucket of tokens. The bucket fills at a fixed speed. Each request takes one token. If the bucket is empty, the request is rejected.
- Concurrency. It does not count time. It only says “at most 10 requests at the same time”. It is good for heavy endpoints, like building a report.
The fixed window problem: the window edge
The sliding window and the token bucket do not have this problem. The token bucket has one more benefit: it allows a little burst (Burst). This means if the customer has been quiet for a while, their bucket is full, and they can send several requests one after another.
Live example: token bucket
The bucket in this example has at most 5 tokens. One token is added every second. Send several requests one after another and see what happens:
Code: a limit for each customer
The most important decision is the key of the limit. That is, for whom is the count kept separate? We count separately for each partner. We call each key a Partition.
builder.Services.AddRateLimiter(options =>
{
// The default is 503. 429 tells the client "slow down", not "server is broken".
options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
options.AddPolicy("per-customer", context =>
{
// One bucket per customer. Fall back to IP only for anonymous calls.
var key = context.User.FindFirstValue("sub")
?? context.Connection.RemoteIpAddress?.ToString()
?? "unknown";
return RateLimitPartition.GetTokenBucketLimiter(key, _ => new TokenBucketRateLimiterOptions
{
TokenLimit = 20, // the biggest burst
TokensPerPeriod = 10,
ReplenishmentPeriod = TimeSpan.FromSeconds(6), // 10 tokens every 6 s = 100 per minute
QueueLimit = 0, // reject at once, do not wait
AutoReplenishment = true
});
});
options.OnRejected = (context, ct) =>
{
if (context.Lease.TryGetMetadata(MetadataName.RetryAfter, out var retryAfter))
context.HttpContext.Response.Headers.RetryAfter =
((int)retryAfter.TotalSeconds).ToString();
return ValueTask.CompletedTask;
};
});
var app = builder.Build();
app.UseRouting();
app.UseAuthentication(); // the limiter needs to know the user
app.UseRateLimiter();
app.MapGet("/products", (ShopDb db, CancellationToken ct) => db.Products.ToListAsync(ct))
.RequireRateLimiting("per-customer");
Some points about this code:
- The default rejection code is 503. You must change it to 429 yourself. The 503 code means “the server is broken”, and it misleads the client.
- The limiter comes after authentication. Otherwise, it does not know yet who the user is, and it counts everyone by IP.
- The key is the customer ID, not the IP. Several customers behind one NAT have one IP. One customer may also have several IPs.
The big problem: several servers
In Production, the service runs on 10 pods. A Load Balancer spreads the requests between them. Now what is the real limit?
- The built-in limiter keeps the counter in the memory of that same pod.
- Each pod counts up to 100 on its own. It does not know about the others.
- So the real limit is 10 × 100 = 1000. On a laptop with one copy everything was correct, but not in Production.
The options:
- A shared counter in Redis. The simplest version is a fixed window. The key is made from the customer ID and the current minute. We add one with the INCR command, and the first time we set a 60-second expiration with the EXPIRE command. Put these two in a Lua script so they run together and atomically. If the number goes over 100, return 429.
- A limit at the edge of the system. An API Gateway or Ingress in front of all pods does the counting. The app code does not change.
- An approximate way. Set each pod’s share to 100 ÷ 10 = 10. It is simple, but it must change when the number of pods changes. It is also not exact if the Load Balancer does not spread the load evenly.
What if Redis goes down?
It depends on what the limit is for:
- To protect the system: usually accept the request (Fail-open). Together with a simple local limit as a backup. We do not want a Redis outage to take down the whole API.
- For security or cost: maybe rejecting (Fail-closed) is right. For example, a limit that stops password guessing.
Important rules
- Return the 429 code and the Retry-After header. Not 503, not 500.
- Choose the key correctly. Customer ID or API Key, not only IP.
- On several servers, the counter must be shared. Either in Redis, or in the Gateway.
- Choose the algorithm by need. To allow a short burst, token bucket. For a heavy endpoint, concurrency.
- Set a separate limit for sensitive endpoints. For example, login and sending SMS codes need a much lower limit.
- Measure the rejections. If a good customer gets 429 all the time, maybe the limit is too low.
Common mistakes
| Mistake | Result | The right way |
|---|---|---|
| Built-in limiter on several pods | The real limit becomes several times the contract. | A shared counter in Redis or the Gateway. |
| Key is only the IP | Customers behind one NAT are rejected together. | Customer ID or API Key. |
| Default 503 code | The client thinks the server is broken. | Set the code to 429. |
| No Retry-After header | The client tries again and again, and the load grows. | Fill Retry-After in OnRejected. |
| Limiter before authentication | The user is not known yet, and everyone is counted by IP. | Authentication first, then the limiter. |
| INCR and EXPIRE separate from each other | If an error happens between the two, the key stays with no expiration. | Both in one Lua script. |
Summary in six lines
- Rate limiting protects the system and the other customers.
- An extra request is rejected with the 429 code and the Retry-After header.
- The fixed window is simple, but at the window edge it allows double. The token bucket controls short bursts.
- The limit key is the customer ID, not the IP.
- The built-in limiter counts in the memory of each pod. On several pods, a shared counter is needed.
- If the shared counter goes down, decide between accept and reject based on the goal of the limit.