Microsoft Orleans
In Orleans, each entity is a Grain, which means a virtual Actor that always exists. Orleans itself activates it on one of the servers, runs its messages one by one, and removes it from memory when it is idle. Its biggest trap is a deadlock in circular calls.
Author: bezzad
The problem: millions of shopping carts
On a sale day, our online shop has one million active users. Each user has a shopping cart. The cart changes every few seconds: adding items, changing the quantity, applying a coupon code.
The usual way is to write each change directly to the database. But:
- The database gets overloaded. Each click means one read and one write.
- Concurrency is trouble. A user adds items in two tabs at the same time, and one of the changes is lost.
- Caching is hard. If we keep the cart in the memory of one server, the next request may go to another server.
The Actor idea helps here: one Actor for each cart. But who creates millions of Actors, spreads them across servers, and creates them again after a restart? Orleans does exactly this.
The main idea: the virtual Actor
In Orleans, we call each Actor a Grain. We call each server a Silo. Several Silos together make a cluster.
The big difference between Orleans and classic Actors is this: a Grain always exists. You do not need to create it or destroy it.
- Each Grain has an ID. For example, Ali’s shopping cart.
- The code only asks for the ID. If the Grain is not in memory, Orleans activates it right then on one of the Silos.
- At any moment, each Grain usually has only one active copy in the whole cluster. So all requests for Ali’s cart go to one place.
- If a Silo dies, the next message activates the Grain on another Silo. Our code sees no change.
This is why we call it a Virtual Actor. Like virtual memory: you have only the address, and the system knows where the data really is.
The code of a Grain
First the Grain interface. All methods are async, because the Grain may be on another server:
public interface ICartGrain : IGrainWithStringKey
{
Task AddItem(string productId, int quantity);
Task<IReadOnlyDictionary<string, int>> GetItems();
}
Then the Grain itself. Its state is read from the database automatically:
[GenerateSerializer]
public sealed class CartState
{
[Id(0)] public Dictionary<string, int> Items { get; set; } = [];
}
public sealed class CartGrain(
[PersistentState("cart", "carts")] IPersistentState<CartState> cart) : Grain, ICartGrain
{
public async Task AddItem(string productId, int quantity)
{
// No lock: Orleans runs one request at a time for this grain.
cart.State.Items[productId] = cart.State.Items.GetValueOrDefault(productId) + quantity;
await cart.WriteStateAsync();
}
public Task<IReadOnlyDictionary<string, int>> GetItems()
=> Task.FromResult<IReadOnlyDictionary<string, int>>(cart.State.Items.AsReadOnly());
}
Setting up the Silo and using the Grain in an API:
var builder = WebApplication.CreateBuilder(args);
builder.UseOrleans(silo =>
{
silo.UseLocalhostClustering(); // for local development only
silo.AddMemoryGrainStorage("carts"); // use a real database in production
});
var app = builder.Build();
app.MapPost("/carts/{customerId}/items", async (
string customerId, AddItemRequest request, IGrainFactory grains) =>
{
var cart = grains.GetGrain<ICartGrain>(customerId); // no "new", no lookup
await cart.AddItem(request.ProductId, request.Quantity);
return Results.NoContent();
});
app.Run();
public sealed record AddItemRequest(string ProductId, int Quantity);
Single-threaded: both a benefit and a trap
By default, each Grain is Non-Reentrant. This means:
- One request starts.
- If it awaits in the middle of the work, the Grain does not start another request. It waits until this request is fully finished.
- So between two awaits, no other request changes the state.
This is great: no lock is needed inside a Grain. But it also has a big trap.
Deadlock: when two Grains wait for each other
At checkout, Ali’s cart asks the discount Grain: “how much is the discount for this cart?”. To calculate, the discount Grain needs the cart items. So it calls Ali’s cart again.
// CartGrain
public async Task<decimal> Checkout()
{
var discounts = GrainFactory.GetGrain<IDiscountGrain>(0);
var discount = await discounts.Calculate(this.GetPrimaryKeyString()); // waits...
return discount;
}
// DiscountGrain
public async Task<decimal> Calculate(string cartId)
{
var cart = GrainFactory.GetGrain<ICartGrain>(cartId);
var items = await cart.GetItems(); // queued behind Checkout: deadlock
return items.Count >= 3 ? 0.1m : 0m;
}
Step by step:
- The cart starts the Checkout request and waits for the discount answer.
- For Orleans, Checkout is not finished yet. So no other request is given to the cart.
- The discount Grain sends the GetItems request to the cart. This request stays in the queue, behind Checkout.
- Now we have a cycle. Checkout waits for the discount, the discount waits for GetItems, and GetItems waits for Checkout to finish.
- None of them moves forward. After the default wait time (thirty seconds), the call fails with a Timeout error.
Ways out, from better to worse
- Remove the cycle in the design. The cart sends its items along with the request. Then no call back is needed.
var discount = await discounts.Calculate(cart.State.Items.AsReadOnly());
- Allow re-entry only for this call chain. The AllowCallChainReentrancy method of the RequestContext class lets in only the requests that come back from this same chain.
using var scope = RequestContext.AllowCallChainReentrancy();
var discount = await discounts.Calculate(this.GetPrimaryKeyString());
- Mark the read-only methods. With the ReadOnly or AlwaysInterleave attributes on the GetItems method in the interface, this method can run in the middle of other work.
- Make the whole Grain Reentrant. This is the simplest way, but also the most dangerous. Between two awaits, another request can change the state. The same Race Condition we ran away from comes back.
Hot Grain
Imagine we have a Grain that counts the total number of today’s orders. Every order calls it. Adding Silos does not make it faster either. Why?
- This Grain has only one active copy, on one Silo.
- It runs the requests one by one.
- So more Silos do not help a single Grain.
Ways out:
- Split the counter. For example, sixteen partial counters based on the order ID, and one Grain that sums them every few seconds.
- Send changes in batches, instead of one call for each order.
- For work without state, use StatelessWorker. This kind of Grain can have several copies at the same time.
A few other Orleans tools
- Timer and Reminder. A timer works only while the Grain is active. A reminder is saved and wakes up even an inactive Grain.
- Stream. For sending events from one Grain to several other Grains, without waiting.
- One-way methods (OneWay). The sender does not wait for an answer. For jobs that need no answer.
Common mistakes
| Mistake | Result | Right way |
|---|---|---|
| Two Grains that call each other with await | Deadlock and a Timeout error. | Send the needed data with the request. |
| Making the whole Grain Reentrant to fix a deadlock | Race Conditions come back inside the Grain. | Only the call chain or read-only methods. |
| One global Grain for everyone | A bottleneck, and more Silos do not help. | Make the state finer or split it. |
| Blocking work inside a Grain | All requests of that Grain get slow. | Only short async code. |
| Not saving the state after a change | After deactivation or a crash, the change is lost. | Save the state after every important change. |
| A query over all Grains, like “all carts above one million” | Orleans is not built for this. | Also write the data separately to a database for reports. |
When to use Orleans?
Good fit
- A large number of entities with their own state: shopping cart, player, device.
- Many reads and writes on each entity.
- The team uses .NET and wants a cluster without complex code.
Poor fit
- A normal CRUD app with low load.
- The main logic is queries and reports over all the data.
- All jobs need one shared global state.
Summary in six lines
- Each Grain is a virtual Actor with an ID. It always exists, and Orleans activates it by itself.
- Each Grain usually has one active copy in the cluster and runs requests one by one.
- No lock is needed inside a Grain.
- If two Grains wait for each other with await, they create a deadlock.
- First remove the cycle in the design. Making the whole Grain Reentrant is the last option.
- A global Grain is a bottleneck. Split the state.