Background work with Background Service
Some jobs must not wait for an HTTP request, like deleting old carts or sending emails. For these jobs we write a Background Service. Three things matter in it. A separate Scope for each round of work, error handling, and safe shutdown.
Author: bezzad
The problem: jobs with no request
Our shop has a few jobs that no user is waiting for:
- Deleting expired carts. Once every few minutes.
- Sending the order confirmation email. The user must not wait three seconds for the email server.
- Sending Outbox messages. Every few seconds, send the saved messages to the message queue.
We do not put these jobs inside an endpoint. We build a service that starts with the app, works in the background, and shuts down with the app. We call it a Background Service.
The idea: a service with a start and an end
In .NET, the Host runs the app. Every service that implements the IHostedService interface has two methods:
- The StartAsync method: it is called when the app starts.
- The StopAsync method: it is called when the app shuts down.
The ready-made BackgroundService class makes this simpler. You only write one method: ExecuteAsync. This method gets a token called stoppingToken. When the app wants to shut down, this token is cancelled.
Code: deleting expired carts
public sealed class ExpiredCartCleaner(
IServiceScopeFactory scopeFactory,
ILogger<ExpiredCartCleaner> logger) : BackgroundService
{
protected override async Task ExecuteAsync(CancellationToken stoppingToken)
{
using var timer = new PeriodicTimer(TimeSpan.FromMinutes(5));
while (await timer.WaitForNextTickAsync(stoppingToken))
{
try
{
// A new scope (and a new DbContext) for every round.
await using var scope = scopeFactory.CreateAsyncScope();
var db = scope.ServiceProvider.GetRequiredService<ShopDb>();
var deleted = await db.Carts
.Where(c => c.ExpiresAt < DateTime.UtcNow)
.ExecuteDeleteAsync(stoppingToken);
logger.LogInformation("Deleted {Count} expired carts", deleted);
}
catch (Exception ex) when (ex is not OperationCanceledException)
{
// One bad round must not kill the service. Try again next tick.
logger.LogError(ex, "Cart cleanup failed");
}
}
}
}
// Program.cs
builder.Services.AddHostedService<ExpiredCartCleaner>();
There are three points in this code. Each one has its own section below:
- A new Scope for each round.
- Errors are caught inside the loop.
- The stoppingToken is passed everywhere.
Point one: a Scope for each round
A background service is built once and lives until the app ends. So it behaves like a Singleton. But DbContext is Scoped. If you take the DbContext straight from the constructor, in the Development environment the app gives an error at start. In Production this check is off by default, so you see no error, and the problem below happens. If you create one Scope at the start of the method and keep it until the end, you have a worse problem:
- One DbContext stays alive forever.
- Everything read with it stays in its Change Tracker. Even after saving.
- Memory slowly grows. After a few days, the pod restarts with an out-of-memory error.
- It also gets slower. Each save checks all those things again.
Point two: an error must not kill the service
If an error comes out of the ExecuteAsync method, by default the whole app stops (this has been the default since .NET 6). This is good, because the error does not stay hidden. But it means that a short database outage shuts down the whole service.
So:
- Catch the error of each round inside the loop. Log it and try again in the next round.
- Do not catch the cancel error. The OperationCanceledException error means the app is shutting down. Let the loop end.
- If the error keeps repeating, create an alert. A background service that silently fails every time is worse than a service that is off.
In-memory queue: sending emails
For the order confirmation email, the endpoint must not wait for the email server. So the endpoint only puts the job in a queue and replies right away. A background service takes it from the queue and sends the email. The Channel class is made for this job.
public sealed class EmailQueue
{
// Bounded: if the queue is full, writers wait instead of using all memory.
private readonly Channel<OrderEmail> _channel = Channel.CreateBounded<OrderEmail>(1_000);
public ValueTask EnqueueAsync(OrderEmail email, CancellationToken ct) =>
_channel.Writer.WriteAsync(email, ct);
public IAsyncEnumerable<OrderEmail> ReadAllAsync(CancellationToken ct) =>
_channel.Reader.ReadAllAsync(ct);
}
public sealed class EmailSender(
EmailQueue queue, IServiceScopeFactory scopeFactory, ILogger<EmailSender> logger)
: BackgroundService
{
protected override async Task ExecuteAsync(CancellationToken stoppingToken)
{
await foreach (var email in queue.ReadAllAsync(stoppingToken))
{
try
{
await using var scope = scopeFactory.CreateAsyncScope();
var client = scope.ServiceProvider.GetRequiredService<IEmailClient>();
await client.SendAsync(email, stoppingToken);
}
catch (Exception ex) when (ex is not OperationCanceledException)
{
logger.LogError(ex, "Sending email for order {OrderId} failed", email.OrderId);
}
}
}
}
// Program.cs
builder.Services.AddSingleton<EmailQueue>();
builder.Services.AddHostedService<EmailSender>();
Point three: safe shutdown
In Kubernetes, every deploy means the old pods shut down. The order is:
- The SIGTERM signal arrives. The .NET host starts to shut down.
- New requests are not accepted. Running requests get time to finish.
- The stoppingToken is cancelled. The background service loop must see this and end.
- The host waits up to a time limit. This limit is set with the ShutdownTimeout setting in HostOptions. Its default value is 30 seconds.
- If Kubernetes does not wait longer, it sends SIGKILL. The Kubernetes grace period is also 30 seconds by default. After that, the process is killed at once.
builder.Services.Configure<HostOptions>(options =>
{
// Must be shorter than the Kubernetes grace period.
options.ShutdownTimeout = TimeSpan.FromSeconds(20);
});
Even with all of this, a pod sometimes dies suddenly. So every background job must be Idempotent. This means that if it runs twice, the result is not broken.
Several pods, several copies of the same job
A financial report is built every night at 2 AM and emailed to customers. To handle the load, we raise the number of pods from 1 to 3. Now what happens?
- Each pod is a full copy of the app. So each pod has its own background service.
- All three wake up at 2 AM. None of them knows about the others.
- Each customer gets three identical emails.
The options, from simple to complex:
- Move the job out of the service. A CronJob in Kubernetes that runs a short-lived pod every night. Now there are no multiple copies.
- Record the run in the database with a Unique Constraint. Before it starts, each pod adds a “today’s report” row. Only the first write succeeds.
- A distributed lock. With a database lock or Redis. Or a ready-made tool like Hangfire or Quartz.NET in cluster mode.
Important rules
- Create a new Scope for each unit of work. Use the CreateAsyncScope method on IServiceScopeFactory.
- Pass the stoppingToken everywhere. To the database, HttpClient and delays.
- Catch and log the error of each round. But do not catch the cancel error.
- Set the shutdown limit lower than the Kubernetes grace period.
- Write the job to be Idempotent. A restart and a second run are always possible.
- Think about the number of pods. A scheduled job on several pods runs several times.
Common mistakes
| Mistake | Result | Right way |
|---|---|---|
| One Scope and one DbContext for the whole life of the service | Memory keeps growing and speed drops. | A new Scope for each round. |
| Not catching errors inside the loop | One temporary error stops the whole app. | Catch the error, log it, and try in the next round. |
| A loop with Task.Delay and no token | The app shuts down late and work is cut in the middle. | Pass the stoppingToken. |
| An in-memory queue for an important job | Jobs are lost on restart. | A durable queue or Outbox. |
| A built-in scheduler on several pods | The job runs several times. | CronJob, Unique Constraint or a lock. |
| Heavy CPU work in the web service | User requests get slow. | A separate Worker service. |
When to use a Background Service?
Good fit
- Short, periodic jobs, like cleanup.
- Reading from a message queue (Kafka or RabbitMQ).
- Sending Outbox messages.
- Light work that must not make the user wait.
Bad fit
- A nightly job that must run only once while we have several pods. A CronJob is simpler.
- A long, important job that must not be lost on restart, but lives only in memory.
- Heavy work that takes the web service’s resources. Move it to a separate service.
Summary in six lines
- A background service starts and shuts down with the app. The ExecuteAsync method is the main work.
- A background service is like a Singleton. Create a new Scope for each round of work.
- Catch the error of each round. If not, by default the whole app stops.
- Pass the stop token everywhere so shutdown is safe.
- An in-memory queue is lost on restart. An important job needs a durable queue.
- On several pods, each pod runs the job separately. You need coordination and Idempotency.