Levelwise
English
Distributed systems

Fault tolerance with Polly

In a distributed system, other services are sometimes slow or return errors. Use a Timeout, a retry with growing delays (Retry) and a fuse (Circuit Breaker) to stop the failure from spreading. But if you set them up badly, they break the system themselves.

Not reviewedWritten with AI helpReading time: 17 minExample with an order service and an inventory serviceC# and Polly code on .NET 10

Author: bezzad

The problem: one slow service takes down the whole shop

In an online shop, the order service calls the inventory service for each order to check the stock. One day the inventory service becomes slow. Its answer takes 30 seconds instead of 200 milliseconds. What happens?

  1. Every order request waits. Each one holds a Thread or a connection (Connection).
  2. Resources run out. A few hundred users at the same time fill all the connections.
  3. The order service stops answering too. Even for work that has nothing to do with inventory, like viewing past orders.
  4. The failure spreads. Services that call the order service also get stuck.

This is called a cascading failure (Cascading Failure). A slow service is worse than a service that is down. A service that is down fails right away. A slow service holds on to your resources.

The goal of Resilience: The goal is not that errors never happen. The goal is that an error in one part stays small and contained, and does not spread to the rest of the system.

Temporary errors and permanent errors

First of all, we must know the kind of error:

Temporary error (Transient)

  • A retry may fix it.
  • Examples: a short network outage, code 503 (service temporarily unavailable), code 429 (too many requests), a timeout.

Permanent error

  • Repeating it never fixes it.
  • Examples: code 400 (bad request), code 401 (not allowed), code 404 (not found).
  • A retry only adds extra load.

Tool one: Timeout

No network call should be without a time limit. The HttpClient class waits 100 seconds by default. For a web page, this is far too long.

A simple rule: set the Timeout based on the real response time of the service. For example, if inventory usually answers in 200 milliseconds and in 800 milliseconds in the worst normal cases, a limit of 2 seconds is reasonable.

Time budget: The time limit of the upper layer must be longer than the total time of the lower layer, with all its retries. Otherwise the upper layer gives up, but the lower layer still works on the same request and creates useless load.

Tool two: Retry, with care

A retry is useful for a temporary error. But a bad Retry is more dangerous than no Retry. There are three main traps:

Trap one: retries multiply across layers

UserOrder servicePrice serviceInventory1 click4 requests16 requestsOne try and three retriesFour times fourInventory is slow, so it gets more requestsMore load, more delay, more retries: a bad loop
Three retries in each layer turn one click into 16 requests.
  1. The order service calls the price service up to 4 times: one first try and three retries.
  2. In each of them, the price service calls inventory up to 4 times.
  3. That is 4 times 4, or 16 requests to inventory for one click.
  4. Inventory was already slow. Now it gets more load and becomes even slower. This is called a Retry storm (Retry Storm).

The solution: only one layer should retry. Usually the layer closest to the user, or the layer that understands the work best.

Trap two: everyone retries at the same time

Fixed delay, no randomness: all at onceBig waves at one momentGrowing delay and randomness: spread over time200ms400ms800ms1600msTime runs from right to left
A growing delay with a little randomness spreads the retries over time.

If a thousand clients get an error together and all retry after exactly one second, a wave of a thousand requests hits the service again. The solution:

  1. A growing delay (Exponential Backoff). For example 200, then 400, then 800 milliseconds. The service gets time to breathe.
  2. A little randomness (Jitter). Each client retries a bit earlier or later. The waves spread out.

Trap three: repeating work that must not be repeated

If a “payment” or “reserve item” request failed because of a Timeout, maybe the work was already done on the server and only the answer did not arrive. A retry means a second payment.

So only retry work that is Idempotent. This means repeating it gives the same result. For POST, either send a repeat key (Idempotency Key) or turn Retry off.

Tool three: Circuit Breaker (fuse)

If inventory is completely broken, a retry does not help. It only wastes our resources and gives inventory no chance to recover. The circuit breaker works like the electric fuse in a house:

ClosedClosed: requests goOpenOpen: fast error, no callHalf-OpenHalf-open: one test callError rate passed the limitBreak time is overTest succeededTest failed
The fuse opens when there are too many errors. After a while it makes one test. If the test succeeds, it closes again.
  1. Closed (Closed). The normal state. Requests go through. The fuse counts the error percentage over a time window.
  2. Open (Open). The error percentage passed the limit. For a set time, no request goes to inventory. An error comes back right away.
  3. Half-open (Half-Open). After that time, one test request goes through. If it succeeds, the fuse closes. If not, it opens again.

Why is a fast error better than waiting? Because the Thread and the connection are freed right away. The order service stays healthy for the rest of its work.

Live example

Break the inventory service and send a few requests. See when the fuse opens and how it makes a test after 5 seconds:

Live example: a fuse in front of the inventory service

The rule of this fuse: if at least 3 of the last 4 requests fail, the fuse opens for 5 seconds.

ClosedFuse state
0Calls to inventory
0Fast rejects, no call
0Successful answers

    When the fuse is open, what do we tell the user?

    This decision is called Fallback. The answer depends on the kind of work:

    1. Reading data. Show the last cached data. For example, “The stock may be a little old”.
    2. A non-essential part. Hide that part for a while. For example, the “recommended products” section. The rest of the page works.
    3. Writing and important work. Return a 503 error right away. Never pretend that the order was saved. If the work can be done later, put it in a queue and say “processing”.

    Tool four: limiting concurrency (Bulkhead)

    Even with a Timeout, one slow dependency can take all the connections. So limit the number of concurrent requests to each dependency. It is like the walls inside a ship: if one part takes in water, the rest of the ship does not sink. In Polly, you do this with a concurrency Rate Limiter.

    Code

    The simple way: the standard handler for HttpClient

    The Microsoft.Extensions.Http.Resilience library is built on Polly. With one line, it adds five protection layers with good default settings to HttpClient:

    Rate Limiter1. Limit concurrent requestsTotal Timeout2. Time limit for all work, all triesRetry3. Retry with a delayCircuit Breaker4. Pause callsAttempt Timeout5. Limit per try
    The order of the layers in the standard handler, from outside to inside.
    builder.Services
        .AddHttpClient<InventoryClient>(client =>
            client.BaseAddress = new Uri("https://inventory"))
        .AddStandardResilienceHandler(options =>
        {
            options.AttemptTimeout.Timeout = TimeSpan.FromSeconds(2);
            options.TotalRequestTimeout.Timeout = TimeSpan.FromSeconds(10);
    
            options.Retry.MaxRetryAttempts = 2;
            options.Retry.Delay = TimeSpan.FromMilliseconds(200);
            options.Retry.DisableForUnsafeHttpMethods(); // no retry for POST, PUT, PATCH, DELETE
    
            options.CircuitBreaker.BreakDuration = TimeSpan.FromSeconds(15);
        });

    A few points:

    1. Retry is on by default for POST too. The DisableForUnsafeHttpMethods method turns it off for unsafe methods.
    2. The time budget is correct. Three tries of 2 seconds and two short delays are less than the total limit of 10 seconds.
    3. The delays grow and have randomness. This is the default behavior of the standard handler.

    The custom way: build the layers yourself

    If you need more settings, put the layers together yourself:

    builder.Services
        .AddHttpClient<PriceClient>(client => client.BaseAddress = new Uri("https://price"))
        .AddResilienceHandler("price", pipeline =>
        {
            pipeline.AddTimeout(TimeSpan.FromSeconds(5)); // total budget
    
            pipeline.AddRetry(new HttpRetryStrategyOptions
            {
                MaxRetryAttempts = 2,
                BackoffType = DelayBackoffType.Exponential,
                UseJitter = true,
                Delay = TimeSpan.FromMilliseconds(200)
            });
    
            pipeline.AddCircuitBreaker(new HttpCircuitBreakerStrategyOptions
            {
                FailureRatio = 0.5,                          // open at 50% failures
                SamplingDuration = TimeSpan.FromSeconds(30),
                MinimumThroughput = 20,                      // need enough samples first
                BreakDuration = TimeSpan.FromSeconds(15)
            });
    
            pipeline.AddTimeout(TimeSpan.FromSeconds(1)); // each attempt
        });

    The order matters. The layer you add first is the outermost layer. So the first Timeout is the limit for the whole work, and the last Timeout is the limit for each try.

    For work that is not HTTP

    For other work, for example calling an SDK, use Polly directly:

    ResiliencePipeline pipeline = new ResiliencePipelineBuilder()
        .AddRetry(new RetryStrategyOptions
        {
            MaxRetryAttempts = 3,
            BackoffType = DelayBackoffType.Exponential,
            UseJitter = true,
            ShouldHandle = new PredicateBuilder().Handle<TimeoutException>()
        })
        .AddTimeout(TimeSpan.FromSeconds(2))
        .Build();
    
    await pipeline.ExecuteAsync(async token => await search.ReindexAsync(productId, token), ct);

    Important rules

    1. Every network call has a Timeout. Both for each try and for the whole work.
    2. Only retry temporary errors. A 400 error is not fixed by repeating it.
    3. Only retry Idempotent work. Otherwise send a repeat key.
    4. Only one layer retries. Retries multiply across layers.
    5. Growing delays with randomness. Not a fixed delay.
    6. A fuse for important dependencies. A fast error is better than a long wait.
    7. A clear Fallback for each dependency. And never pretend that the work was done.
    8. Monitor everything. The number of retries and fuse openings are early signs of a problem.

    Common mistakes

    Mistake Result Right way
    Retry in every layer A Retry storm and the last service breaks. Only one layer retries.
    A fixed delay without randomness Waves from all clients at the same time. A growing delay with Jitter.
    Retry for a payment POST The money is taken twice. Turn Retry off or use a repeat key.
    A longer Timeout when the service is slow More resources wait. A real time limit and a fuse.
    The upper layer’s time limit is shorter than the lower layer’s Useless work goes on in the lower layer. A time budget from top to bottom.
    Retry for a 400 or 401 error Extra load with no chance of success. Only temporary errors.

    Which tool for which problem?

    Problem Tool
    The service sometimes fails for a moment Retry with a growing delay
    The service is slow and holds our resources A time limit and limited concurrency
    The service is fully broken for a while A fuse and a Fallback
    The user must not see an empty page A cache and a Fallback

    Summary in six lines

    1. One slow service can take down the whole system by holding resources.
    2. Every network call needs a time limit for each try and for the whole work.
    3. Retry only for temporary errors, only for Idempotent work, and only in one layer.
    4. The delays should grow, with a little randomness.
    5. A fuse cuts the calls when there are many errors, so the broken service has a chance to recover.
    6. In .NET, the standard handler adds most of these with one line.