Levelwise
English
ASP.NET Core

gRPC and SignalR

We use gRPC for fast calls between services. It has a strict contract, binary messages, and it runs on HTTP/2. We use SignalR to send live news from the server to the user. Both keep a long-lived connection, so load balancing and multiple servers need special care for them.

Not reviewedWritten with AI helpReading time: 16 minExample with the shop's orders and inventoryC# code on .NET 10

Author: bezzad

Two different problems

Our shop has two new needs:

  1. The order service keeps asking the inventory service for stock. Thousands of times per minute. It works with JSON and REST, but the messages are big, and the contract between the two teams lives only in a document. One team renames a field, and the other service silently breaks.
  2. The customer wants to see the order status live. “Packed”, “Shipped”. Right now the page asks the server every 5 seconds: “Do you have news?”. Most answers are “No”.

For the first problem we have gRPC. For the second one, SignalR. This lesson explains both, because they share one feature: a long-lived connection.

Part one: gRPC for calls between services

The idea

In gRPC, we write the contract first. It is a proto file that says what methods the service has and what fields each message has. The server code and the client code are generated automatically from this file.

RESTSeparate document, maybe outdatedOrder serviceInventory service{ "productId": 7, "available": 12 }Each field name repeats in every messagegRPCinventory.protoOrder serviceInventory serviceClient codeServer code08 07 10 0COnly field number and value
In REST the contract is separate from the code. In gRPC the code of both sides is generated from one contract file.

Three things make gRPC fast and safe:

  1. Binary messages (Protobuf). Messages are smaller than JSON and faster to read.
  2. The HTTP/2 protocol. Several calls at the same time go over one connection. There is no need to open a new connection for each call.
  3. A strict contract. If one side breaks the contract, the code of the other side does not compile. It does not silently break in Production.

Code: the contract

syntax = "proto3";
option csharp_namespace = "Shop.Inventory";

service Inventory {
  rpc GetStock (StockRequest) returns (StockReply);
}

message StockRequest {
  int32 product_id = 1;
}

message StockReply {
  int32 product_id = 1;
  int32 available = 2;
}

The number next to each field is its number in the binary message. The field name does not go into the message. So never change the number of a field, and do not reuse the number of a deleted field.

Code: server and client

// Inventory service: implement the generated base class.
public sealed class InventoryService(ShopDb db) : Inventory.InventoryBase
{
    public override async Task<StockReply> GetStock(StockRequest request, ServerCallContext context)
    {
        var available = await db.Stock
            .Where(s => s.ProductId == request.ProductId)
            .Select(s => s.Available)
            .FirstOrDefaultAsync(context.CancellationToken);

        return new StockReply { ProductId = request.ProductId, Available = available };
    }
}

// Program.cs of the inventory service
builder.Services.AddGrpc();
app.MapGrpcService<InventoryService>();
// Order service: register the generated client.
builder.Services.AddGrpcClient<Inventory.InventoryClient>(o =>
    o.Address = new Uri("https://inventory"));

// Use it with a deadline. Never wait forever for another service.
var reply = await inventory.GetStockAsync(
    new StockRequest { ProductId = 7 },
    deadline: DateTime.UtcNow.AddSeconds(2),
    cancellationToken: ct);
Do not forget the Deadline. If the inventory service gets stuck and you set no deadline, the calls of the order service pile up until it fails too. With a deadline, the call ends with an error after 2 seconds, and the server also knows it should not continue.

Four kinds of calls

  1. A simple call (Unary). One request, one answer. Like above.
  2. A stream from the server (Server streaming). One request, several answers one after another. For example, “send me the price changes”.
  3. A stream from the client (Client streaming). Several messages from the client, one answer at the end. For example, a batch upload of stock.
  4. A two-way stream (Bidirectional). Both sides send messages whenever they want.

The load balancing trap

Interviews ask about this a lot. We run the inventory service on 5 pods. But we see one pod under heavy load and the rest idle.

Order service1000 calls a secondLayer 4 balancerSees only the connectionOne connectionpod 1pod 2pod 3pod 4All the loadIdle
A layer 4 Load Balancer only spreads connections. Because all calls are on one HTTP/2 connection, they all reach one pod.

Why? Step by step:

  1. The gRPC client opens one HTTP/2 connection and keeps it.
  2. All calls go over that one connection.
  3. A layer 4 Load Balancer (like a normal Service in Kubernetes) sees only the connection, not the calls. So it gives the connection to one pod, once.
  4. The result: all calls from this client reach that same pod.

Solutions:

  • A layer 7 proxy that understands HTTP/2. Like Envoy, or a Service Mesh like Linkerd. This proxy spreads each call separately.
  • Client-side load balancing. The gRPC client in .NET can get the addresses of all pods (for example from the DNS of a Headless Service) and spread the calls between them by itself.
Browsers do not speak gRPC directly. Browsers do not give JavaScript code enough control over HTTP/2. To call gRPC from a browser, you need the gRPC-Web version. For a public API and browsers, REST and JSON are usually simpler.

Part two: SignalR for live news to the user

The idea

Instead of the browser asking every few seconds, one connection stays open. Whenever there is news, the server itself sends it.

Asking every 5 secondsBrowserServerNoNoNoShippedThree useless requests, news up to 5 seconds lateOpen connectionBrowserServerConnected once and stayed openShipped, at that momentNo extra requests
With constant asking, most requests are useless and the news arrives late. With an open connection, the news arrives at once.

The SignalR library does the hard work:

  1. It picks the best way to connect. First WebSocket. If that fails, Server-Sent Events. If that fails too, Long Polling.
  2. It reconnects a dropped connection. If you turn this on in the client.
  3. It makes sending to groups simple. To one user, to one group, or to everyone.

Code: the Hub and sending news

The center of communication in SignalR is a Hub. With an interface, we define the client-side methods with exact types:

public interface IOrderClient
{
    Task OrderStatusChanged(Guid orderId, string status);
}

[Authorize]
public sealed class OrderHub : Hub<IOrderClient>;

// Program.cs
builder.Services.AddSignalR();
app.MapHub<OrderHub>("/hubs/orders");

The news usually does not start inside the Hub. It comes from somewhere else, for example when the “shipped” event arrives from the shipping service. For this, we use IHubContext:

public sealed class OrderShippedHandler(IHubContext<OrderHub, IOrderClient> hub)
{
    public Task HandleAsync(OrderShipped e) =>
        hub.Clients.User(e.CustomerId.ToString())
           .OrderStatusChanged(e.OrderId, "Shipped");
}

And in the browser:

const connection = new signalR.HubConnectionBuilder()
  .withUrl("/hubs/orders")
  .withAutomaticReconnect()
  .build();

connection.on("OrderStatusChanged", (orderId, status) => showStatus(orderId, status));
await connection.start();
Send to the user, not to the connection. A customer may have the site open in two tabs and on a phone. Sending to the user delivers the news to all of their connections. The user ID comes from authentication.

Several servers: the main problem of SignalR

The site runs on 3 pods. The customer’s browser is connected to pod number 1. The “shipped” event reaches pod number 2. Pod number 2 does not know this customer, because their connection is on another pod. So the news never arrives.

Shipping eventOrderShippedpod 2pod 3pod 1Redis BackplaneGives the message to all1. Publish2. To allCustomer browserConnected to this pod3. Sent
With a Backplane, each pod publishes the message, and the pod that holds the user's connection delivers it to the browser.

Solutions:

  1. A Backplane like Redis. Each pod publishes the message to Redis. All pods receive it, and whichever pod holds the user’s connection sends it. You turn it on with the AddStackExchangeRedis method after AddSignalR.
  2. The managed Azure SignalR Service. The service itself holds the connections. Our servers only send messages.
You also need a Sticky Session. The first step of a SignalR connection (negotiation) and fallback methods like Long Polling send several separate requests. All of them must reach the same pod. So turn on sticky sessions in the Load Balancer. You only do not need it if all clients use only WebSocket and skip negotiation, or if you use Azure SignalR Service.

Which one to use where?

REST and JSON gRPC SignalR
Best use Public API, browser, partners Internal calls between services Live news from the server to the user
Message format JSON text Binary Protobuf JSON text or binary MessagePack
Contract Separate from code (for example OpenAPI) A proto file and generated code Hub methods
Direction The client asks All four kinds, with streams The server can start too
From the browser Yes Only with gRPC-Web Yes

Important rules

  1. Always set a deadline in gRPC. And pass the cancellation token.
  2. Do not change the field numbers in proto. Add a new field with a new number.
  3. Balance gRPC load per call. A layer 7 proxy or client-side balancing.
  4. In SignalR, send to a user or a group, not to a connection ID. The connection ID changes on every reconnect.
  5. SignalR on several servers needs a Backplane and sticky sessions. Or a managed service.
  6. A SignalR message may not arrive. If the user was offline, the message is lost. So keep the real state in the database and read it again after reconnecting.

Common mistakes

Mistake Result Right way
A gRPC call without a deadline When one service gets stuck, other services get stuck too. A short deadline for each call.
Layer 4 load balancing for gRPC One pod gets all the load. A layer 7 proxy or client-side balancing.
Changing a field number in proto Old services read the data wrong. Only add fields with new numbers.
SignalR on several pods without a Backplane Some users never get the news. Redis or Azure SignalR Service.
Fully trusting the SignalR message A user who was offline never sees the status. The database is the source of truth. The message only notifies.
Using gRPC for a public browser API Clients cannot call it easily. REST and JSON outside, gRPC inside.

Summary in six lines

  1. In gRPC the contract is in a proto file, and the code of both sides is generated from it.
  2. Binary messages and HTTP/2 make calls between services fast.
  3. All gRPC calls go over one connection. Layer 4 load balancing is not enough.
  4. The SignalR library keeps a connection open so the server can send news by itself.
  5. On several servers, SignalR needs a Backplane and sticky sessions.
  6. A live message is only news. The real state stays in the database.