Microservices Architecture
Split the system into several small services. Each one owns one part of the business, has its own database and is released on its own. This gives teams independence, but it has a high technical cost. So use it only when you have a real reason.
Author: bezzad
The problem: one big app for many teams
We have an online shop. At the start, a team of 4 people built one app (a Monolith). Orders, inventory, payment and shipping are all in one project and one database.
Two years later, the company has grown. Now 5 teams work on the same app. The problems start:
- Releases get slow. The payment team has a small change, but it must wait until the inventory team’s work is ready too. Because everything is released together.
- One bug breaks everything. A memory leak in the reports part takes down the whole shop.
- The load is not equal. Product search has a hundred times more load than placing orders. But to scale search, we must run more copies of the whole app.
The idea of Microservices is this: each part of the business is a separate app. Each team has its own service and its own database, and releases it whenever it wants.
Three ways to build a system
Between “one tangled app” and “dozens of separate services” there is a middle way: the Modular Monolith.
- One app (Monolith). One app and one database. If the boundaries are not clear, all the code becomes dependent on all the other code.
- Modular app (Modular Monolith). It is still one app and is released once. But each module has its own tables. Another module talks to it only through a public interface.
- Separate services (Microservices). Each module is a separate app with a separate database. Calls go over the network.
The cost of Microservices
In one app, placing an order and decreasing the stock are in one transaction. Either both happen, or neither does. But in Microservices these two jobs are in two services and two databases. So:
- There is no shared transaction. If the order is saved and the call to inventory fails, the data becomes inconsistent. To fix this, we need Outbox, Saga and compensating actions.
- The network is not reliable. Every call may be slow, may fail or may arrive twice. So you need Retry, Timeout, Circuit Breaker and Idempotency.
- Finding bugs is hard. One request goes through several services. Without Distributed Tracing and structured logs, you do not know where it broke.
- Operations work multiplies. Instead of one release pipeline, you have ten. Each one needs its own monitoring, alerts and versioning.
- Data is not the same right away. When services sync through messages, it takes a few seconds until all of them see the same thing (Eventual Consistency).
Where does a service boundary come from?
The biggest mistake is to build one service for each table: a customer service, a product service, a price service. The result: for every simple job, several services must call each other.
The right boundary comes from parts of the business. In DDD language, we call each part a Bounded Context. To find the boundary, ask these questions:
- Which team owns this rule? The “maximum discount” rule belongs to the sales team. So it stays in the sales service.
- Does this word mean the same thing everywhere? “Product” in sales means name, picture and price. In inventory it means quantity and shelf. So these are two separate models.
- Which data changes together? Data that always changes together, in one transaction, must be in one service.
Each service, its own database
The main rule is: no service touches another service’s database directly. Not for reading, not for writing.
Why is a shared database bad? Step by step:
- The order service and the inventory service use the same table.
- The inventory team renames a column.
- The order service breaks, because it still reads the old name.
- From now on, the two teams must coordinate every change and release together.
This is called a Distributed Monolith: you have all the pain of the network, but no independence.
So if the order service needs inventory data, it has two ways:
- Call an API. When you need the answer right now. But be careful: for a list of 50 orders, do not call the service 50 times in a row. Collect all the ids and get them with one batch request.
- Keep a local copy. The inventory service publishes an event on every change. The order service keeps a small table with only the data it needs. The data is a few seconds behind, but reading it is fast and does not depend on inventory.
Synchronous calls or messages?
Services have two ways to talk:
Synchronous (HTTP or gRPC)
- When you need the answer right now.
- Example: the checkout page must know the final price right now.
- If the target service is down, your request fails too.
With messages (Kafka or RabbitMQ)
- When the work can happen a few seconds later.
- Example: sending a text message after an order is placed.
- If the target service is down, the message stays in the queue and is processed later.
Why is a long synchronous chain dangerous?
- Imagine each service is healthy 99.9 percent of the time.
- One request goes through 5 services in a row.
- The probabilities multiply. So the total chance of success is about 99.5 percent.
- This means the whole system fails about 5 times more than one service. The longer the chain, the worse it gets.
Code
The right start: Modular Monolith
Each module has a public contract. The rest of the code sees only this contract, not the internal tables and classes.
// Inventory module: the only public surface.
public interface IInventoryModule
{
Task<bool> ReserveAsync(Guid orderId, IReadOnlyList<OrderLine> lines, CancellationToken ct);
}
// Everything else in the module is internal.
internal sealed class InventoryModule(InventoryDbContext db) : IInventoryModule
{
public async Task<bool> ReserveAsync(
Guid orderId, IReadOnlyList<OrderLine> lines, CancellationToken ct)
{
// Uses only the inventory tables (its own schema).
// ...
await db.SaveChangesAsync(ct);
return true;
}
}
If one day the inventory module really needs to be separated, only the implementation of this interface changes. The order code does not change.
After the split: calls with protection
When inventory becomes a separate service, the same contract is implemented with HTTP. Because the network is not reliable, we also add a Resilience layer:
builder.Services
.AddHttpClient<IInventoryModule, InventoryHttpClient>(client =>
client.BaseAddress = new Uri("https://inventory"))
.AddStandardResilienceHandler(); // timeout, retry, circuit breaker
internal sealed class InventoryHttpClient(HttpClient http) : IInventoryModule
{
public async Task<bool> ReserveAsync(
Guid orderId, IReadOnlyList<OrderLine> lines, CancellationToken ct)
{
var response = await http.PostAsJsonAsync(
$"reservations/{orderId}", lines, ct);
return response.IsSuccessStatusCode;
}
}
Non-urgent work: events
To send a text message, the order service does not wait for the notification service. It only publishes an event:
public sealed record OrderPlaced(Guid OrderId, Guid CustomerId, decimal Total);
This event must be sent together with saving the order, using the Outbox pattern. Otherwise the order may be saved but the event never goes out.
Key rules
- Start with a Modular Monolith. Make the boundaries clear in the code. Delay the split until you have a real reason.
- Split only for a real reason. Good reasons: several teams that do not want to wait for each other, a part with a very different load, or a part with a separate technical need.
- Take the boundary from the business, not from the tables.
- No shared database. Each service owns its own data.
- Wherever you can, use messages instead of a synchronous chain. Keep synchronous chains short.
- Have observability tools from day one. A trace id (Trace Id) in all services, structured logs and monitoring.
Common mistakes
| Mistake | Result | Right way |
|---|---|---|
| Ten services for a small team | All the team’s time goes to infrastructure, not the product. | Start with a Modular Monolith. |
| One service for each table | Every simple job calls several services. | Boundaries based on business parts. |
| Shared database | Teams must release together (Distributed Monolith). | Separate databases, and talk through an API or events. |
| Long chain of synchronous calls | One slow service makes all of them slow. | Messages for non-urgent work, a local copy of data. |
| A transaction across services, assuming “it always works” | Data becomes inconsistent and nobody notices. | The Outbox and Saga patterns. |
| No Tracing and no structured logs | Finding one bug takes several days. | A shared trace id from day one. |
When to use Microservices?
Good fit
- Several independent teams work on one big product.
- There are parts with very different loads that must scale separately.
- The boundaries of the business parts are clear and stable.
- The team has the experience and tools for a distributed system.
Poor fit
- A small team and a new product that does not know its domain well yet.
- The boundaries are not clear yet and change a lot.
- The main reason is “it is in fashion” or “big companies do it”.
- The business does not accept any temporary inconsistency.
Summary in six lines
- Microservices architecture means each part of the business is a separate service with a separate database.
- Its main benefit is team independence, not speed or cleaner code.
- Its cost is high: the network, no shared transaction, tracing and operations work.
- Start with a Modular Monolith and split only for a real reason.
- Take the boundary from the parts of the business, and never share a database.
- Use messages for non-urgent work, and keep synchronous chains short.