The Saga pattern
When one job must run in several services and there is no shared transaction, split the job into small steps. Keep a compensating action ready for each step. If a step fails, compensate the earlier steps.
Author: bezzad
The problem: a transaction across several services
A customer makes a purchase. The system must do four things: save the order, reserve the item in the warehouse, take the money, and ship the package.
In a Monolith with one database, this is simple. We put everything inside one transaction. If one step fails, the database undoes everything (Rollback).
But in Microservices, each service has its own database. There is no shared transaction.
Now imagine the money was taken, but the warehouse said “this item is out of stock”. In our system:
- The customer’s money is gone.
- There is no item to ship.
- No database undoes this automatically.
Saga is the solution to this exact problem.
Why not a distributed transaction (2PC)?
An old approach is the two-phase commit (Two-Phase Commit). A coordinator asks all databases “are you ready?”. If all say yes, it tells all of them “commit”. This approach is usually not a good fit for Microservices:
- Long locks. Until the last service answers, all records stay locked.
- Dependency on everyone. If one service or the coordinator goes down, the others wait.
- Little support. Kafka, RabbitMQ, Redis and most NoSQL databases do not take part in 2PC.
The Saga idea: small steps and compensating actions
We split the big job into several local transactions. Each service runs only its own database transaction. For each step, we also write a compensating action (Compensation) that cancels the effect of that step.
There are three kinds of steps:
- Compensatable step. If a problem happens later, it has a compensating action. For example, reserving the item.
- Point of no return (Pivot). If this step succeeds, the Saga does not go back anymore. In our example, taking the money.
- Retriable step. It comes after the point of no return and must succeed in the end. If it fails, we try again. For example, shipping the package.
Live example
Pick a scenario and see what the Saga does, step by step:
Compensation is not Rollback
A database Rollback behaves as if nothing happened. But in a Saga, each step really happened, and others may have seen it.
Database Rollback
- The change was never seen.
- It is automatic.
- No trace is left.
Compensating action in a Saga
- The change happened and maybe was seen.
- We must write it ourselves.
- Its effect stays. For example, the customer gets a text message “order cancelled”.
So a compensating action is a business action, not a simple delete. For example, if an order is cancelled after payment, the compensation for the payment is a “refund”. We do not delete the payment record. We add a refund record.
Two ways to run it: Choreography and Orchestration
Way one: Choreography (each service listens to the previous event)
There is no manager. Each service does its job and sends an event. The next service listens to that event.
Way two: Orchestration (one manager gives commands)
One part, called the Orchestrator, knows the whole flow. It gives a command to each service, gets the answer, and decides the next step.
| Choreography | Orchestration | |
|---|---|---|
| Where do we see the flow? | Spread over all services | In one place, inside the Orchestrator |
| Coupling between services | Very low | All depend on the Orchestrator |
| Adding a new step | Several services change | Only the Orchestrator changes |
| Finding errors | Hard, you must follow the events | Easy, the state of each Saga is saved |
| Good for | Short flow with 2 or 3 steps | Long flow or flow with many conditions |
Code
A simple version to understand the idea
When a step succeeds, we put its compensating action on a stack (Stack). If a step fails, we run the compensating actions in reverse order.
public sealed class PlaceOrderSaga(
IOrders orders, IWarehouse warehouse, IPayments payments, IShipping shipping)
{
public async Task RunAsync(Order order, CancellationToken ct)
{
var compensations = new Stack<Func<Task>>();
try
{
await orders.CreateAsync(order, ct);
compensations.Push(() => orders.CancelAsync(order.Id));
await warehouse.ReserveAsync(order.Id, order.Items, ct);
compensations.Push(() => warehouse.ReleaseAsync(order.Id));
await payments.ChargeAsync(order.Id, order.Total, ct);
// Pivot: after a successful payment we only move forward.
}
catch
{
while (compensations.TryPop(out var compensate))
await compensate();
throw;
}
await shipping.ShipWithRetryAsync(order.Id, ct);
}
}
The real version: state in the database
In Production, the state of each Saga must be saved in the database. The MassTransit library does this with a State Machine. Each order has one state row. Each message moves the Saga from one state to the next:
public sealed class OrderState : SagaStateMachineInstance
{
public Guid CorrelationId { get; set; } // = OrderId
public string CurrentState { get; set; } = "";
}
public sealed class OrderSaga : MassTransitStateMachine<OrderState>
{
public State Reserving { get; private set; } = null!;
public State Paying { get; private set; } = null!;
public State Completed { get; private set; } = null!;
public State Cancelled { get; private set; } = null!;
public Event<OrderCreated> OrderCreated { get; private set; } = null!;
public Event<StockReserved> StockReserved { get; private set; } = null!;
public Event<StockFailed> StockFailed { get; private set; } = null!;
public Event<PaymentDone> PaymentDone { get; private set; } = null!;
public Event<PaymentFailed> PaymentFailed { get; private set; } = null!;
public OrderSaga()
{
// Every message has a CorrelationId (the OrderId),
// so MassTransit finds the right saga row by itself.
InstanceState(x => x.CurrentState);
Initially(
When(OrderCreated)
.Publish(ctx => new ReserveStock(ctx.Saga.CorrelationId))
.TransitionTo(Reserving));
During(Reserving,
When(StockReserved)
.Publish(ctx => new ChargePayment(ctx.Saga.CorrelationId))
.TransitionTo(Paying),
When(StockFailed)
.Publish(ctx => new CancelOrder(ctx.Saga.CorrelationId))
.TransitionTo(Cancelled));
During(Paying,
When(PaymentDone)
.Publish(ctx => new ShipOrder(ctx.Saga.CorrelationId))
.TransitionTo(Completed),
When(PaymentFailed)
.Publish(ctx => new ReleaseStock(ctx.Saga.CorrelationId))
.Publish(ctx => new CancelOrder(ctx.Saga.CorrelationId))
.TransitionTo(Cancelled));
}
}
This code is exactly the picture above. Each During line is a state. Each When is a message that the Saga waits for. You can save the state with EF Core or Redis.
Important rules
- Every step and every compensating action must be Idempotent. A message may arrive twice. If “release the item” runs twice, it must not release twice the items. Use the order ID to check if it was already done.
- A compensating action must not fail for good. If it gives an error, try again after a delay. If it still fails after several tries, send it to the error queue (Dead Letter) and alert a person. Never drop it silently.
- Save the Saga state. After a restart, the Saga must continue from the same step it was on.
- The database change and the message send must happen together. Use the Outbox pattern. Otherwise the database may change, but the message is never sent.
- Set a Timeout. If the warehouse service never answers, the Saga must not wait forever. After a set time, start the compensation.
- A compensating action may arrive before the action itself. For example, “cancel reservation” arrives before “reserve”. The service must record this, so that when “reserve” arrives, it does not run it.
Common mistakes
| Mistake | Result | The right way |
|---|---|---|
| Keeping the Saga state only in memory | After a restart, the Saga stays half done. | State in the database. |
| Compensating by deleting the record | History is lost and auditing is not possible. | Add a compensating record, like a “refund”. |
| Steps that are not Idempotent | A duplicate message takes the money twice. | A unique ID for each step and a duplicate check. |
| Showing “confirmed” before the Saga ends | The customer thinks the purchase is done, but later it is cancelled. | A “processing” status until the Saga ends. |
| Choreography with many steps | Nobody understands the whole flow. | Orchestration. |
| No Timeout | The Saga waits forever for an answer. | A Timeout and starting the compensation. |
When to use a Saga?
Good fit
- One business job across several services with separate databases.
- Each step has a reasonable compensation.
- A few seconds of inconsistency is acceptable for the business.
Bad fit
- All data is in one database. A normal transaction is enough.
- The business does not accept any moment of inconsistency.
- The service boundaries are wrong, and every simple job needs several services. Fix the boundaries first.
Summary in six lines
- In Microservices there is no shared transaction.
- A Saga splits a big job into several local transactions.
- Each step has a compensating action that runs in reverse order.
- Compensation is a business action, not deleting data.
- Choreography for a short flow, Orchestration for a long flow.
- Do not forget saved state, Idempotency, Outbox and Timeout.
Related questions: 5. Lost message and duplicate message · 15. Repeated request in payment · 21. Idempotency Key · 25. Failure in compensation, Saga