1Service lifetimes in DIEasyDependency Injection
Question: You see this code in code review. The ReportCache service is registered as Singleton. The AppDbContext service is registered with the AddDbContext method.
builder.Services.AddDbContext<AppDbContext>(o => o.UseSqlServer(cs));
builder.Services.AddSingleton<ReportCache>();
public class ReportCache(AppDbContext db)
{
public Task<Report> LoadAsync(CancellationToken ct)
=> db.Reports.FirstAsync(ct);
}
This service is used in every request. What is the difference between Singleton, Scoped and Transient? What do you think about this code? Does it have a problem in Production?
Short answer: A Singleton service lives until the end of the app. So the DbContext inside it also stays until the end of the app. As a result, all requests use this one DbContext at the same time. But DbContext is not built for this.
Hint (if the candidate says there is no problem): In the Development environment, the app gives this error at run time: “Cannot consume scoped service from singleton”. In Production this error does not appear. But after a few hours, this error sometimes shows up in the log: “A second operation was started on this context instance before a previous operation completed”. Users also sometimes see old data. Why?
First: the three lifetimes
| Type | How many instances are made? | Good example |
|---|---|---|
| Singleton | Only one instance for the whole app | Settings, in-memory cache |
| Scoped | One instance for each request | DbContext |
| Transient | A new instance every time it is asked for | A light service with no state |
Why does this happen? (step by step)
- The DI system builds ReportCache only once. At that moment it also builds one AppDbContext and gives it to it.
- The ReportCache service is never destroyed. So that AppDbContext is never disposed either. This mistake is called a Captive Dependency.
- Now all requests share one ReportCache, so they also share one AppDbContext.
- When two requests run a query at the same moment, DbContext gives the “second operation” error. The reason is that DbContext is not thread-safe.
- Every record that DbContext reads stays in its own memory (the change tracker). If another service changes that record in the database, this DbContext still returns the old version. So the user sees old data. Memory also grows slowly.
- Why does only Development give an error? Because ASP.NET Core turns on the ValidateScopes check by default only in Development.
Simple rule: A service can only depend on a service whose lifetime is equal to or longer than its own.
Solution: In the Singleton service, each time we need it, we create a short-lived DbContext and then dispose it:
public class ReportCache(IDbContextFactory<AppDbContext> dbFactory)
{
public async Task<Report> LoadAsync(CancellationToken ct)
{
await using var db = await dbFactory.CreateDbContextAsync(ct);
return await db.Reports.AsNoTracking().FirstAsync(ct);
}
}
Another way is to create a separate scope with IServiceScopeFactory. This way is common in a BackgroundService.
Follow-up question: How do you use HttpClient? Why is creating a new HttpClient in every request bad?
Follow-up answer:
- Each new HttpClient opens its own TCP connections. Even after Dispose, the operating system keeps that socket in the TIME_WAIT state for a few minutes.
- Under heavy load, the free ports run out and we get a SocketException error. This is called Socket Exhaustion.
- There is also the opposite mistake: one static HttpClient forever. In this case ports do not run out, but the app never sees DNS changes.
- First right way: IHttpClientFactory. This tool pools the connections and refreshes them every few minutes.
- Second right way: one shared HttpClient with SocketsHttpHandler and the PooledConnectionLifetime setting.
Red flag: The candidate only recites the definition of the three lifetimes, but cannot explain what problem a wrong mix of them creates.
2Working with async/await and the Thread PoolEasyasync/await and the thread pool
Question: In an ASP.NET Core API, one place in the code is written like this:
var price = _httpClient.GetStringAsync(url).Result;
You see this code in code review. On campaign day, this endpoint gets several times its normal traffic. What do you think about this code? Does it have a problem? And what does async/await actually do? Does it make code faster?
Short answer: Async code does not make a single request faster. Its job is to free the thread while we wait for I/O (network, database, file). But the Result property blocks that same thread. Under heavy load, all threads of the thread pool sit and wait. So new requests get stuck in the queue. This is called Thread Pool Starvation.
Hint (if the candidate says there is no problem): With low traffic, everything is fine. On campaign day, response time goes from 50 milliseconds to more than 10 seconds. But the server CPU is only at 15%. The number of threads in the process also grows slowly and steadily. Why?
A simple example: Threads are like waiters in a restaurant.
- In sync code (or with Result): the waiter gives the order to the kitchen and stands right there until the food is ready. With 8 waiters, only 8 tables get served at the same time.
- In async code: the waiter gives the order and goes to take the order of the next table. When the food is ready, any free waiter carries it.
Why does this happen? (step by step)
- Each request runs on a thread from the thread pool.
- With await: when we reach the network call, the method returns for a while and the thread goes back to the pool. When the network answer arrives, the rest of the method runs on a free thread. For this, the compiler turns the method into a state machine.
- With Result: the thread stays blocked until the answer arrives and does no other work.
- Suppose the pool starts with 8 threads (usually the number of CPU cores). Each outside call also takes 1 second. So with Result, at most 8 requests per second get an answer. The rest wait in the queue.
- The thread pool sees the queue has grown and adds new threads. But it does this slowly (about one or two threads per second). That is why the number of threads grows slowly.
- These threads only wait and do no computing work. So CPU stays low. “Low CPU but a slow system” is the sign of this exact problem.
In apps that have a SynchronizationContext (like WinForms and old ASP.NET), using Result can even cause a deadlock.
Solution: Make the code async from start to end (async all the way):
public async Task<decimal> GetPriceAsync(string url, CancellationToken ct)
{
var text = await _httpClient.GetStringAsync(url, ct);
return decimal.Parse(text);
}
Note: for heavy CPU work (for example, a long calculation), async does not help. Because there the thread is really busy working, not waiting.
Follow-up question: If we want to send 100 HTTP requests, but at most 10 at the same time, what do you do?
Follow-up answer: One way is the Parallel.ForEachAsync method with a limit of 10:
await Parallel.ForEachAsync(urls,
new ParallelOptions { MaxDegreeOfParallelism = 10, CancellationToken = ct },
async (url, token) => await CallAsync(url, token));
Another way is to use a SemaphoreSlim with a capacity of 10.
Why is running all 100 requests with Task.WhenAll not good? Because all 100 requests go out together. The other service may get overloaded, or we may hit its rate limit.
Red flag: The candidate says “async means the code runs on a new thread” or “async makes code faster”.
3The N+1 problem in EF CoreMediumEntity Framework Core
Question: We have an endpoint that returns the list of the last 50 orders, together with the customer name of each order. Lazy Loading is turned on in the project, and the code is this:
var orders = await db.Orders
.OrderByDescending(o => o.CreatedAt)
.Take(50)
.ToListAsync();
return orders.Select(o => new OrderDto(o.Id, o.Total, o.Customer.Name));
You see this code in code review. What do you think? Does it have a problem? How many queries does this code send to the database?
Short answer: This problem is called N+1: 1 query for the list of orders, and then N queries (one for each order) for the customer. The cause is Lazy Loading. This means EF Core reads the customer of each order at the exact moment the code touches it for the first time, and it does this with a separate query.
Hint (if the candidate says there is no problem): The tables are small, but the answer of this endpoint takes about 2 seconds. We turn on the SQL log in EF Core (or use SQL Server Profiler or an APM tool). We see that for one call of this endpoint, 51 queries go to the database: one query for the orders table, and 50 separate queries for the customers table (one customer each time). The API output is correct and returns the same 50 orders. So the problem is not the number of records, it is the number of queries. Why?
What Lazy Loading means:
- When Lazy Loading is on, EF Core reads only the columns of the orders table in the first query. The customer of the order stays empty.
- Whenever the code reaches the customer of an order, EF Core sees it is not loaded yet. So at that moment it sends a query for that one customer.
- Without Lazy Loading, the customer of the order would just be null and we would get a NullReferenceException error. That is why some teams turn on Lazy Loading and do not see the N+1 problem.
Why does this happen? (step by step)
- The ToListAsync method runs one query: 50 orders, without customers.
- Then Select runs on the in-memory list. For each order, it reads the customer name.
- Each time, Lazy Loading sends a separate query. So 50 orders means 50 queries.
- Each query, even if it takes 5 milliseconds, has one network round trip. 50 round trips one after another make the endpoint slow. With a list of 1000 it gets even worse.
The same problem also happens without Lazy Loading, if the developer runs a query inside a foreach loop.
Solution:
- With Include, the customer comes with a JOIN in the same first query. So 51 queries become 1 query.
var orders = await db.Orders.Include(o => o.Customer)
.OrderByDescending(o => o.CreatedAt).Take(50).ToListAsync();
- A better way is Projection: we write Select before ToListAsync. In this case EF Core builds the JOIN itself and reads only the needed columns. Lazy Loading and change tracking are not involved either.
var result = await db.Orders
.OrderByDescending(o => o.CreatedAt).Take(50)
.Select(o => new OrderDto(o.Id, o.Total, o.Customer.Name))
.ToListAsync();
- If we really need the entity itself and we only read, we add the AsNoTracking method.
- Many teams turn Lazy Loading off completely, so this problem does not stay hidden.
One more important point: As long as the query is an IQueryable, it runs in the database (SQL). If we call the AsEnumerable or ToList method in the middle, the later filters and Selects run in memory. This means all the data comes from the database first.
Follow-up question: Now only one query goes out, but it is still slow. What do you do next?
Follow-up answer:
- The candidate takes the real SQL query from the log.
- They look at its execution plan.
- For example, if the whole table is scanned to sort by creation date, they add an index on the CreatedAt column of the orders table.
- They measure the time before and after the change.
Red flag: The candidate has never looked at the SQL that EF Core builds, or does not know when Lazy Loading runs a query.
4Caching with RedisMediumRedisCaching patterns
Question: The user profile page is read from the database in every request, and the database CPU is high. We decided to cache the profile in Redis. Three questions:
- How do you design reading and writing the cache?
- When the user edits their profile, how do you make sure old data is not shown?
- If Redis goes down, what happens?
We also have another key: “home page settings”. This key is read in every home page request (about 2000 requests per second) and its TTL is 10 minutes. How do you see this design? What happens at the moment this key expires?
Short answer:
- We use the Cache-Aside pattern and always set a TTL.
- After an edit, we delete the cache key.
- If Redis goes down, we read directly from the database.
- The name of the last problem is Cache Stampede: when a popular key expires, all requests go to the database at the same time.
Hint (if the candidate says there is no problem): In monitoring we see that the “home page settings” key expires every 10 minutes. At exactly that moment, the database CPU reaches 100% for a few seconds. Why?
Part 1: How does the Cache-Aside pattern work?
- First we look for the key in Redis (for example, the profile key of user number 42).
- If it is there, we return it.
- If it is not there, we read from the database. Then we put it in Redis with a TTL (for example, 10 minutes) and return it.
Why is a TTL needed? If something goes wrong somewhere and the cache is not deleted, the TTL makes sure old data stays only a few minutes, not forever.
Part 2: When the data changes
- First we save in the database, then we delete the cache key. The next request reads fresh data from the database and caches it again.
- Why delete and not update the cache? Suppose we have two edits at the same time:
- The database may be updated in the order A then B.
- But the cache may be updated in the order B then A.
- Now the cache has wrong data until the TTL ends.
- Deleting the cache does not have this problem.
Part 3: If Redis goes down
- The cache must be “optional”. This means we catch the Redis error and read directly from the database.
- We set a short timeout for Redis. Otherwise each request waits a few seconds for a dead Redis, and the whole site gets slow.
- We must keep in mind that the database must also handle the load without the cache, or we must have a plan for this case.
Part 4: Why does the database CPU go up at the moment of expiry? (Cache Stampede)
- While the key is in the cache, for example 2000 requests per second get their answer from Redis.
- At the moment the key expires, all these requests see the cache empty.
- So they all send the same query to the database together, until one of them fills the cache again.
Ways to fix it:
- Only one request builds the data, and the others wait for its result (one lock for each key). In .NET, the HybridCache library does this inside each server with the GetOrCreateAsync method.
- A little before expiry, we refresh the cache in the background.
- If many keys expire together, we add a small random number to the TTL (jitter). This way they do not all expire at one moment.
What should we not cache? Data that must always be exact (like an account balance), or data that changes very quickly.
Follow-up question: What is the difference between a local (In-Memory) cache and a Redis cache? If we use a local cache instead of Redis and we have 5 servers, what problem happens?
Follow-up answer:
- A local cache is in the memory of the same server. It is very fast, because there is no network. But each server has its own separate copy.
- The problem: the user edits their profile. Only the server that got this request deletes its own cache. The other 4 servers show old data until the TTL ends. So with each refresh, the user may see old or new data.
- First way: a short TTL for the local cache.
- Second way: send a “delete this key” message to all servers (for example, with Redis Pub/Sub).
- Third way: combine two levels, meaning a local cache in front of Redis.
Red flag: The candidate has not thought at all about deleting the cache after a data change, or about Redis going down.
5Messaging with Kafka: lost messages and duplicate messagesMedium to hardKafkaIdempotencyOutbox and Inbox
Question: After an order is placed, we send the “order placed” message (OrderCreated) to Kafka. The Notification service reads this message and sends an SMS to the customer. The order code is this:
db.Orders.Add(order);
await db.SaveChangesAsync();
await producer.ProduceAsync("orders", new Message<string, string> { Key = order.Id, Value = json });
The services run on Kubernetes with several pods, and we deploy several times each week. You see this code and this design in code review. What do you think? Does the customer always get exactly one SMS? Why?
Short answer:
- Problem 1 (the message gets lost): writing to the database and writing to Kafka are two separate jobs and do not share a transaction. The fix is the Outbox Pattern.
- Problem 2 (duplicate message): Kafka delivers a message “at least once”, not “exactly once”. The fix is to make the consumer Idempotent.
Hint (if the candidate says there is no problem): We have two complaints from customers. First, some customers placed an order but got no SMS. The order is in the database, but its message is not in the topic. Second, some customers got two SMS messages for one order. Why?
Problem 1: Why does the message get lost? (Dual Write)
- Saving to the database succeeds and the order is stored.
- Right after that, the pod restarts (for example, because of a deploy), or Kafka is not reachable for a few seconds.
- Sending the message to Kafka does not run or gives an error. The order exists, the message does not, and nobody notices.
- If we change the order (Kafka first, then the database), the problem flips: a message goes out for an order that is not in the database.
Solution: Outbox Pattern
- In the same database transaction, we save both the order and a row in the Outbox table (the message text). The database makes sure that either both are saved, or neither.
- A BackgroundService reads the unsent Outbox rows and sends them to Kafka. Then it marks them as “sent”.
- Even if Kafka is down for a few minutes, the messages stay in the Outbox and are sent later. So no message is lost.
Note: if the worker sends the message and crashes before marking it, next time it sends the same message again. So the Outbox also creates duplicate messages. This brings us to problem 2.
Another way instead of the Outbox is a CDC tool (like Debezium) that reads the database changes.
Problem 2: Why are two SMS messages sent?
- The Notification service reads the message and sends the SMS.
- Before it tells Kafka “this message is done” (that is, commits the offset), the pod dies or a rebalance happens.
- So Kafka thinks the message was not processed and gives it again. A second SMS is sent.
- This behavior is not a bug in Kafka. Delivery in Kafka is at-least-once. So we must always expect duplicate messages.
Solution: an Idempotent consumer
Idempotent means this: if one message is processed two times, the result is the same as processing it once.
- Each message has a unique ID (for example, MessageId or OrderId).
- Before sending the SMS, we record this ID in the ProcessedMessages table. This table has a unique constraint on the ID.
- If the insert fails with a unique error, it means the message was processed before. So we do not send the SMS.
- We commit the offset after successful processing, not before it. Because if we commit before processing and then crash, the message is lost.
Follow-up question: If the order of the messages of one order matters (for example, “placed” before “canceled”), how does Kafka keep the order?
Follow-up answer: Kafka guarantees order only inside one partition. If we set the message Key to the order ID, all messages of one order go to one partition. So they are also read in the same order.
Red flag: The candidate says “Kafka is exactly-once by itself and we have no duplicate messages”, or their only fix for problem 1 is “try/catch and retry”.
6Fixing slowness in ProductionHardDistributed tracingStructured logging
Question: At 10 AM we deployed a new version of the API. Since then, the monitoring dashboard shows that the API response time went from about 100 milliseconds to 3 seconds. But server CPU is only 20%, memory is normal, and there is no special error in the log. What do you do, step by step?
Short answer: First we save the user (rollback). Then we find the cause with data, not with guesses. The main clue is this: “low CPU but slow” means the app is not busy computing; it is waiting for something.
Steps and the reason for each one:
- Roll back the version (Rollback). The problem started exactly with the deploy. So the fastest way for the user is to go back to the previous version. Finding the cause is still possible later.
- See what changed. Look at the diff of the new version: a new query, a new call to a service, a new library, or a settings change.
- Find where the time is spent, with tracing. A distributed tracing tool (like OpenTelemetry) shows how much of these 3 seconds is in the database, how much in another service, and how much inside the app itself.
- See what the app is waiting for. We have four main candidates:
- Thread pool starvation (like question 2): in the dotnet-counters tool, the thread pool queue length is high and the number of threads grows slowly.
- A slow query or a lock in the database: in the database itself, look at the list of running queries and locks.
- A slow outside service: in tracing we see that the HTTP call to that service took most of the time.
- The connection pool runs out, for the database or HttpClient: requests wait for a free connection. For example, when the new code does not dispose the connection.
- Take a dump, if the cause is still not clear. Take a dump of the process with the dotnet-dump tool. Then see on which line of code all threads are waiting. If hundreds of threads are waiting on Result or on a lock, the cause is found.
- After finding the cause: fix, load test, deploy again, and a short report (postmortem). The goal is to find out sooner next time (for example, with an alert on latency).
Follow-up question: How do you know the problem is from GC (garbage collection)?
Follow-up answer: In the dotnet-counters tool, the candidate looks at two numbers:
- The percent of time in GC. If it is very high, the app keeps pausing for GC.
- The number of generation 2 GCs. This is the most expensive kind of GC.
If memory also keeps growing after each GC, we probably have a memory leak (question 35). Of course, in this scenario CPU is low, so GC is less likely. Because heavy GC usually pushes CPU up.
Red flag: Their first answer is “we add more servers”, or they start changing code based on guesses, without data.
7Choosing between Monolith and MicroservicesHardArchitecture stylesMicroservices
Question: A team of 4 people wants to build a new online shop. One team member suggests building 10 separate microservices from the start (user, product, inventory, order, payment, shipping, and so on). Each one has its own database. What do you think? When do you split a system into microservices, and when not?
Short answer: For a small team and a new product, a Modular Monolith is usually better. The reason is simple:
- Microservices solve an organizational problem: several teams that want to work independently.
- In return, they have a large technical cost.
- A team of 4 does not have this organizational problem. So it only pays the cost and gets no benefit.
The cost of microservices, with a real example: Placing an order must also reduce the inventory stock.
- In a Monolith: both jobs are done in one database transaction. Either both succeed, or neither. It is only a few lines of code.
- In microservices: order and inventory are two separate services with two separate databases. We have no shared transaction. If the order is saved and then the call to the inventory service fails, the data becomes inconsistent. So we must build messaging, an Outbox, a Saga and compensating actions (questions 5 and 25).
Other costs of microservices:
- Every call goes over the network. So it is slower and can fail.
- To find one bug, you must look at the logs of 10 services. So tracing is needed.
- Instead of one pipeline and one deploy, we have 10.
- DevOps work grows a lot.
What does Modular Monolith mean?
- We have one app and one deploy. But the code is split into separate modules (for example, each module is one project).
- Each module has its own tables. Another module is not allowed to touch these tables directly. It talks to the module only through a public interface.
- First benefit: today the system is simple.
- Second benefit: if one day a module really needs to be split out, its boundary is already clear. So splitting it is easy.
When do we split out a service? Only when we have a real reason:
- We have several separate teams that do not want to wait for each other to deploy.
- One part has a very different load and must scale separately. For example, product search has 100 times the load of the rest.
- One part has a different technical need. For example, it needs a different technology, or it must work even when the rest is broken.
Where do service boundaries come from? From the Bounded Context in DDD, not from database tables. Example:
- In the catalog part, “product” means name, picture and description.
- In the inventory part, “product” means count and shelf.
- So these are two separate models with two separate owners.
A common mistake: a shared database between services.
- Suppose two services use one table.
- If one column changes, both break.
- So we must deploy both together.
- Result: we have all the costs of microservices and none of the benefits. This is called a Distributed Monolith.
Follow-up question: When two services talk to each other, when do you choose sync communication (like HTTP or gRPC) and when async communication (messages with Kafka or RabbitMQ)?
Follow-up answer:
- Sync communication is for when we need the answer right now. Example: the payment page must know the final price right now.
- Async communication is for when the job can be done a few seconds later. Example: sending an email after an order. If the email service is down for a few minutes, placing the order must not break.
- Why is a long sync chain bad?
- Suppose each service is healthy 99.9% of the time.
- One request passes through 5 services one after another.
- The probabilities are multiplied. So the total success probability becomes about 99.5%.
- The longer the chain, the more fragile the system.
Red flag: The candidate thinks microservices are always better and cannot name any specific cost for them.
8System design: Flash SaleDesignWeight ×2System designRedis
Question: At 12 noon, an online shop sells 1000 phones at a discount. In the first minute, about 200 thousand users come in and press the buy button many times. Design the order system. We have three conditions:
- No more than 1000 units are sold.
- The system does not go down.
- Each user buys only one unit.
Ask the candidate to draw it on the whiteboard. If they ask questions before designing (for example, “is payment also in this same flow?”), that itself is a plus.
Short answer: We have three main ideas:
- Most users (199 thousand people) will not buy anything anyway. So we must reject their requests cheaply and early.
- The most important part is reducing the stock with an atomic operation. This way two people do not buy the last phone together.
- Heavy jobs (creating the order, the invoice, the email) are done later with a queue.
First, let’s see the main problem (overselling). This code is wrong:
var stock = await db.Products.Where(p => p.Id == id).Select(p => p.Stock).FirstAsync();
if (stock > 0)
{
await db.Database.ExecuteSqlAsync($"UPDATE Products SET Stock = Stock - 1 WHERE Id = {id}");
}
Why is it wrong? Step by step:
- Suppose only one phone is left.
- Two requests arrive at the same moment.
- Both read the stock and see the number 1.
- Both see the condition as true, and both reduce the stock.
- The stock becomes -1 and one extra phone is sold.
This is called a Race Condition. There is a gap between “read” and “write”, and another request comes in the middle of this gap.
A good design (step by step):
- Make the needs clear. Is payment in this same flow? How much time does the user have to pay?
- Reduce the load before it reaches the server. The product page and pictures should be on a CDN. Set a limit on the number of requests (Rate Limiting) for each user and each IP. Why? Because most requests are refreshes and repeated clicks, and they must not reach the database.
- Reduce the stock atomically (the most important part). We have two ways:
- First way, with Redis: the stock is kept in Redis and reduced with the DECR command. Redis runs commands one by one. So two DECR commands never run in the middle of each other. If the number goes negative, we answer “sold out”.
- Second way, with the database: checking the stock and reducing it are done in one UPDATE statement. The condition “stock greater than zero” is inside that same statement (code below). The database runs this one statement atomically. If the number of changed rows is zero, it means the stock is finished.
UPDATE Products SET Stock = Stock - 1
WHERE Id = @id AND Stock > 0;
- One purchase per user. Put a unique constraint on the pair of user and flash sale. Or in Redis, create a key for each user that is created only if it does not exist (the SET command with the NX option). Important point: the user check and the stock reduction must be done together (with a Lua script in Redis or one database transaction). Otherwise the stock may be reduced, but the purchase is rejected because the user is a duplicate. In this case one phone gets lost.
- Separate the reservation from the heavy work. After a successful reservation, only send one message to the queue (Kafka or RabbitMQ) and tell the user “reserved”. Other services build the order, the invoice and the email at their own speed. Why? A reservation in Redis takes a few milliseconds, but building a full order is slower. The queue smooths out the peak load.
- Payment with a deadline. A reservation is valid for, for example, 10 minutes. If it is not paid, the phone goes back to the stock.
- Handle duplicate requests. The user may click twice. The client sends an Idempotency Key. For a repeated key, the server returns the same earlier answer (question 21).
- Monitoring and testing. Before the sale day, run a load test. Have a dashboard for stock and errors. Also add a switch to turn off the sale quickly (feature flag).
Follow-up question 1: If Redis restarts in the middle of the sale, what happens to the stock?
Answer: If Redis has no persistence, the stock number is lost. We have two ways:
- The database is the main source of truth. Each reservation is also recorded in the database. This way we can calculate the stock again (1000 minus the number of reservations).
- Or we run Redis with persistence and a replica.
Follow-up question 2: If the order “first come, first served” must be exact, what do you do?
Answer: We build an entry queue (virtual waiting room). Each user gets a turn number and is let into the buy page in order.
Red flag: The candidate checks the stock with two separate commands (first read, then reduce), or their only answer is “more servers and pods”.
9Changing one record at the same time: Lost UpdateMediumTransactions and isolation levels
Question: In the support panel, several staff members work together. When saving, the order edit page sends all of the order data with one PUT request. The server writes all fields to the database. Suppose two staff members open the page of one order at the same time. The first one changes the address and saves. The second one, a few seconds later, changes the status to “shipped” and saves. How do you see this design? What data stays in the database at the end? Does it have a problem?
Short answer: The form of the second staff member still had the old address. The PUT request sends all fields. So the old address was written over the new address. The solution is Optimistic Concurrency: each record has a version number. If the version has changed, the save is rejected.
Hint (if the candidate says there is no problem): We have seen this in Production: the first staff member changed the address and got the message “saved”. But after the second staff member saved, the new address was gone. The package was sent to the old address. There is no error in the log either.
What happened? (in time order)
| Time | Staff A | Staff B | Data in the database |
|---|---|---|---|
| 1 | Opens the form | Address: Tehran, Status: New | |
| 2 | Opens the form (address Tehran) | Address: Tehran, Status: New | |
| 3 | Changes the address to Tabriz and saves | Address: Tabriz, Status: New | |
| 4 | Changes the status to “shipped” and saves (the form still has address Tehran) | Address: Tehran, Status: Shipped |
The last write won (last write wins). The change of A was lost silently.
Why does “we add a transaction” not help?
- These are two separate requests.
- There are a few minutes between opening the form and saving.
- Each save on its own is successful and correct.
- The problem is that B saves with stale data.
Solution: Optimistic Concurrency with EF Core
- We add a version column to the table. With each change to the row, SQL Server changes its value automatically:
public class Order
{
public int Id { get; set; }
public string Address { get; set; } = "";
[Timestamp] public byte[] RowVersion { get; set; } = [];
}
- When the form opens, it also gets the version value (RowVersion). When saving, it sends it back.
- When saving, EF Core adds a condition to the UPDATE statement: “only if the version is still the same old version”.
- In the example above, the save of A changed the version. So the condition is not true for B. No row is updated, and EF Core gives a DbUpdateConcurrencyException error.
- In this case the API returns a 409 Conflict response. The form tells B: “Someone else changed this order, please open it again.” What we do next (reload or merge the changes) is a business decision.
Other ways and their limits:
- Pessimistic Lock: lock the record while the form is open. This is not good for a form that the user keeps open for several minutes. For example, if the user closes the browser, what happens to the lock?
- Sending only the changed fields (PATCH): in this example it solves the problem. But if both people change the same field, one change is still lost.
Follow-up question: In a REST API, with what standard way do the server and the client know that the data the client has is stale?
Follow-up answer: With the ETag header. Steps:
- In the GET response, the server sends the version in the ETag header.
- On PUT, the client sends the same value back in the If-Match header.
- If the version has changed, the server returns a 412 Precondition Failed response.
Red flag: The candidate says “we add a transaction” or “we add a lock” and does not notice that these are two separate requests with time between them.
10Paging on a big tableMediumSQL and indexes
Question: The orders table has 50 million records. The order list API pages with OFFSET and FETCH (in EF Core, the same as Skip and Take). Each page has 20 records, and the newest orders come first. Several new orders are placed every second. The code is this:
var page = await db.Orders
.OrderByDescending(o => o.CreatedAt)
.Skip((pageNumber - 1) * 20)
.Take(20)
.ToListAsync(ct);
How do you see this paging method? With this amount of data and this number of new orders, what happens in Production?
Short answer: For OFFSET, the database must read all records before that page and throw them away. The duplicates happen because, between two requests, a new order was added and all records moved down by one place. The solution is Keyset Pagination (or Cursor Pagination).
Hint (if the candidate says there is no problem): We have seen two things in Production. First, the first pages are fast, but page 50000 takes several seconds. Second, users say that when they open the next page, sometimes the last order of the previous page shows up again at the top of the new page.
Problem 1: Why are the last pages slow?
- Page 50000 means “skip 999980 rows, then give 20 rows”.
- The database cannot jump directly to row 999980. It must start from the beginning of the index.
- It counts about one million rows and throws them away. Then it returns 20 rows.
- So the bigger the page number, the more work and the slower it is.
Problem 2: Why is a duplicate record shown?
- The user sees page 1: orders 1 to 20.
- At that moment a new order is placed and sits at the top of the list. Now all orders have moved down by one place.
- The user asks for page 2 (skip 20 rows). Row 21 is now the old order 20. So it is shown again.
- The other way around, if a record is deleted, one record is never shown at all.
Solution: Keyset Pagination
Instead of “which page number?”, we say “after the last record I saw, give me the next 20”:
var page = await db.Orders
.Where(o => o.CreatedAt < lastCreatedAt
|| (o.CreatedAt == lastCreatedAt && o.Id < lastId))
.OrderByDescending(o => o.CreatedAt).ThenByDescending(o => o.Id)
.Take(20)
.ToListAsync(ct);
Why is this way better?
- With an index on the creation time and the order ID, the database jumps directly to the right place in the index (seek). It reads only 20 rows. So page 1 and page 50000 have the same speed.
- A new order at the top of the list does not move our starting point. So duplicates do not happen.
- Why is the order ID also needed? Because two orders may have the same creation time. Without the ID, the order is not unique, and a record may still be skipped or repeated.
- We give the client a “cursor”. For example, the same two values (time and ID) as a base64 string. The client sends it back in the next request.
Another point: counting all rows (to show “page 1 of 2.5 million”) is also expensive on a big table. Usually an approximate number is enough.
Follow-up question: With your method, can the user go directly to page 47? How do you explain this to the Product Owner?
Follow-up answer: No. The Keyset method only has “next page” and “previous page”. It is great for infinite scroll and APIs. We tell the Product Owner: “Nobody wants page 47 because of its number. They are looking for a specific order. It is better to give a date filter and search.” If page numbers are really needed, we keep OFFSET only for the first few pages (for example, at most 100 pages).
Red flag: The candidate only says “we add an index” and does not know why OFFSET is slow even with an index.
11Safe shutdown in KubernetesMedium to hardKubernetesBackground services
Question: Our service runs in Kubernetes. It both answers HTTP requests and reads messages from Kafka. Traffic is high, and we deploy several times a day with a rolling update. We have no special settings for app shutdown. Everything is on the default settings of Kubernetes and ASP.NET Core. During a deploy, when an old pod shuts down, what happens to the requests and messages that are being processed? How do you design a safe shutdown (graceful shutdown)?
Short answer: When Kubernetes shuts down a pod, the app must finish its half-done work and then close. The main problem is this:
- At one moment, it tells the app “shut down”.
- At the same time, it tells the load balancer “do not send traffic to this pod anymore”.
- But the second one takes effect a few seconds later.
Hint (if the candidate says there is no problem): We have seen this in Production: each time we deploy, in the first few seconds some users get a 502 error. Some Kafka messages are also processed twice.
First, let’s see how Kubernetes shuts down a pod:
- It sends the SIGTERM signal to the process. It means: “Please finish your work and close.”
- At the same time, it removes the pod from the service’s Endpoints list. This change takes a few seconds to reach kube-proxy and the ingress.
- It has a deadline called the termination grace period (default 30 seconds). If the app has still not closed after this deadline, it sends SIGKILL and the process is killed at once.
Why do we see a 502 error?
- Steps 1 and 2 start at the same time.
- With SIGTERM, the app starts to close.
- But the load balancer still sends new requests to this pod for a few seconds.
- These requests reach an app that no longer accepts requests.
Why do we see duplicate messages? A message that was in the middle of processing did not finish, or its offset was not committed, and the process was killed. So another pod gets the same message again.
Solution (step by step):
- Wait a few seconds before closing. We add a preStop hook. Kubernetes runs this first, then sends SIGTERM. In these few seconds the load balancer has time to remove the pod from the list:
lifecycle:
preStop:
sleep:
seconds: 5
- Finish the current requests. With SIGTERM, the ASP.NET Core app no longer takes new requests and waits for the running requests. The maximum time for this wait is set by the ShutdownTimeout setting in HostOptions.
- Do the time math. The preStop time plus ShutdownTimeout must be less than the termination grace period. For example, 5 + 30 = 35, so we set the grace period to 45 seconds. Otherwise SIGKILL arrives in the middle of the work.
- Background work. In a BackgroundService, we pass the stop token (stoppingToken) to all jobs. This way, with SIGTERM, no new job starts.
- The Kafka consumer. First the current message finishes. Then its offset is committed. Then the consumer is closed with the Close method. This way Kafka quickly gives the partitions to another pod.
Even with all of this, a pod sometimes dies suddenly (for example, OOM or a node failure). So message processing must be idempotent (question 5).
Follow-up question: What is the difference between a readiness probe and a liveness probe? If the liveness probe also checks the database, what problem happens?
Follow-up answer:
- The readiness probe asks: “Can I take traffic right now?” If the answer is no, traffic just does not go to this pod.
- The liveness probe asks: “Is the app alive, or is it stuck?” If the answer is no, Kubernetes restarts this pod.
- Now suppose liveness checks the database and the database becomes slow for a few seconds. In this case all pods restart together. A small database problem turns into an outage of the whole service.
- So the liveness probe must check only the process itself.
Red flag: The candidate does not know what SIGTERM is, or thinks Kubernetes waits by itself until all work is finished.
12Hot Path code and memoryHardPerformance and SpanMemory and the Garbage Collector
Question: A service receives price messages as text and parses them. Each message has three parts separated by commas: symbol, price and volume (for example, AAPL and 123.45 and 100). This code runs 50 thousand times per second:
var parts = line.Split(',');
var quote = new Quote(parts[0], decimal.Parse(parts[1]), int.Parse(parts[2]));
You see this code in code review. What do you think? With this load, what happens in Production? Would you change anything in it?
Short answer: Step by step:
- Each run of this code creates several new objects on the heap (one array and three strings).
- 50 thousand times per second means hundreds of thousands of objects per second that become garbage at once.
- The GC system must keep collecting them. Some GCs pause the threads.
- Solution: reduce allocation with tools like Span.
Hint (if the candidate says there is no problem): We have seen this in Production: processing time is good most of the time. But every few seconds, everything stops for a few hundred milliseconds. The dotnet-counters tool also shows that the number of GCs is very high.
Why does a lot of allocation cause pauses? (step by step)
- Each time we create a new object (and each time we call the Split or Substring method), an object is created on the heap.
- A new object goes into generation 0 (Gen0). Generation 0 is small, and at this speed it fills up very quickly.
- When it is full, GC runs. A generation 0 GC is fast. But if it runs very often, its total time becomes significant.
- Objects that are still alive during GC move to generations 1 and 2. When generation 2 fills up, a full and expensive GC runs. This GC can take a few hundred milliseconds. These are the same “pauses” we saw.
- So the less garbage we create, the less often GC runs.
Solution:
- Measure first. Find where most of the allocation is with the dotnet-trace tool or PerfView. Compare before and after with BenchmarkDotNet and its MemoryDiagnoser feature.
- Parse without creating strings and arrays. For this, use ReadOnlySpan. A Span is only a “window” on the same original string and copies nothing:
ReadOnlySpan<char> s = line;
int i = s.IndexOf(',');
var symbol = s[..i];
s = s[(i + 1)..];
int j = s.IndexOf(',');
var price = decimal.Parse(s[..j], CultureInfo.InvariantCulture);
var volume = int.Parse(s[(j + 1)..]);
- Also look at other common sources of allocation:
- Joining strings in a loop.
- Using LINQ and lambdas in hot code.
- Turning a struct into an object (boxing).
- Creating a List without an initial capacity.
- Reuse buffers. With ArrayPool, borrow an array (Rent) and give it back after the work (Return). This is especially important for arrays bigger than about 85 kilobytes. Because these arrays go directly to the Large Object Heap, and collecting them is expensive.
- Check GC settings last. Look at settings like Server GC after fixing the code, not before.
Follow-up question: When do you use ValueTask instead of Task? What things must you not do with ValueTask?
Follow-up answer:
- The Task type is a class. Usually, each call of an async method creates an object on the heap.
- Sometimes a method finishes without waiting most of the time (for example, it answers from the cache 99% of the time) and it is in hot code. In this case ValueTask removes this allocation.
- Limits of ValueTask:
- Await it only once.
- Do not await it from two places at the same time.
- Do not read its result (Result) before the work is finished.
- If you need any of these, first turn it into a Task with the AsTask method.
- The default is still Task, because it is simpler and safer.
Red flag: The candidate starts “optimizing” without measuring, or cannot explain why a lot of allocation causes slowness.
13A Job that runs again on several PodsHardBackground servicesRedis
Question: Every night at 2 AM, a BackgroundService builds a financial report and emails it to customers. To handle more load, we want to raise the number of pods of this service from 1 to 3. We do not change the Job code. Does this change cause a problem? Why?
Short answer: Each pod runs a full copy of the app. So each pod has its own BackgroundService. 3 pods means 3 schedulers. Each one starts the job at 2 AM. Solution: a shared mechanism between the pods that lets only one of them run.
Hint (if the candidate says there is no problem): We have seen this in Production: since the day after we made it 3 pods, each customer gets 3 identical emails every night.
Ways, from simple to complex:
- Separate the Job from the service. Create a CronJob in Kubernetes. This CronJob runs a short-lived pod every night. Set its concurrencyPolicy to Forbid, so two runs do not start together. This is the simplest way, because we no longer have several copies.
- Record the run in the database with a unique constraint. We have a table for runs. We put a unique constraint on “Job name and date”. Before starting, each pod adds a row for today to this table:
INSERT INTO JobRuns (JobName, RunDate) VALUES ('daily-report', '2026-10-04');
-- UNIQUE (JobName, RunDate)
- Only the first write succeeds. The other two pods get a unique error and drop the work. This way is simple and safe, because the database itself guarantees it.
- A database lock. In SQL Server the sp_getapplock feature, and in PostgreSQL the advisory lock feature, do this. This lock is tied to the connection. If the pod dies, the connection closes and the lock is released automatically.
- A Redis lock with the SET command and the NX option (only if the key does not exist) and an expiry time. Or a ready-made tool like Hangfire or Quartz.NET in cluster mode.
- In all cases, it is better if sending the email itself is also idempotent. This means for each customer and each date we record “sent”. If the Job fails in the middle and runs again, nobody gets an email twice.
Follow-up question: Suppose you took a lock with Redis and the lock TTL is 30 seconds. The first pod stops for 40 seconds in the middle of the work (for example, because of a long GC or a network problem). What happens?
Follow-up answer:
| Time | pod A | pod B |
|---|---|---|
| 0 | Takes the lock until second 30 and starts the work | Waiting |
| 5 | Stops (GC pause) | |
| 30 | Still stopped; the lock expires | |
| 31 | Takes the lock and starts the work | |
| 45 | Wakes up and thinks it still has the lock; continues | Working at the same time |
- Now two pods work at the same time.
- A longer TTL only lowers the chance. It does not solve the problem.
- Renewing the lock during the work also helps. But the renewal itself may not happen because of this same pause.
- The right solution is a Fencing Token. Step by step:
- Each time someone takes the lock, they also get an increasing number. For example, A gets the number 33 and B the number 34.
- Each write to the database carries this number with it.
- The database keeps the biggest number it has seen. It rejects a write whose number is smaller.
- So after B starts (number 34), the writes of A (number 33) are rejected.
- If the candidate knows the debate about Redlock and the critique by Martin Kleppmann, that is a sign of very good depth.
Red flag: The candidate only says “we take a lock with Redis” and has never thought about the lock expiring in the middle of the work.
14Retry stormHardResilience with Polly
Question: Service A calls service B, and B calls service C. In both layers (A to B and B to C), this is set with Polly: 3 retries and a timeout of two seconds for each try. The normal response time of C is about 200 milliseconds. You see this setting in code review. What do you think? Does it have a problem? If C gets slow and its response time reaches about 2 seconds, what happens? What is the right setting?
Short answer: The number of retries in the layers is multiplied. One user request can send up to 16 requests to C. When C is slow, these extra requests make it even slower. A bad cycle starts.
Hint (if the candidate says there is no problem): One day C got a little slow, and its response time went from 200 milliseconds to about 2 seconds. In monitoring we saw that the number of incoming requests to C grew several times, while the number of users was the same. After a few minutes, C went down completely. Then A and B also stopped answering, even for work that has nothing to do with C. Why?
Why does this happen? (step by step)
- Retries multiply. Service A calls B up to 4 times (1 original call and 3 retries). In each of these 4 calls, B calls C up to 4 times. This means 4 × 4 = 16 requests to C for one user click.
- The timeout is right at the edge. Service C answers in about 2 seconds, and the timeout is also 2 seconds. So most tries time out and a retry starts. Important point: C is still processing the earlier request. This means it does repeated work.
- The bad cycle. More load on C means C gets slower. Slower C means more timeouts. More timeouts mean more retries. And more retries mean even more load on C. This goes on until C goes down completely.
- Cascading failure. All threads and connections of A and B wait for C. So A and B do not answer other users either. Even for work that has nothing to do with C.
The right setting:
-
Only one layer retries, not all layers.
-
The gap between retries grows step by step. This is called exponential backoff: 200 milliseconds, then 400, then 800. Also add a small random number (jitter), so thousands of clients do not retry at the same moment.
-
Retry only for temporary errors: timeouts and the codes 503 and 429. Do not retry for an error like 400, because repeating does not fix it.
-
The right time budget. The timeout of the upper layer must be longer than the total time of the lower layer (with its retries). Otherwise A gives up, but B is still working on the same request.
-
Use a Circuit Breaker. It is like an electric fuse:
- If a large percent of requests to C fail, the circuit opens.
- For example, for 30 seconds we do not call C at all and return an error at once.
- During this time C gets a chance to recover.
- Then we send one test request. If it succeeds, we send requests to C again.
-
Use a Bulkhead, meaning limit the number of concurrent requests to C. This way one slow dependency does not take all threads and connections of B.
-
In .NET, the Microsoft.Extensions.Http.Resilience library (which is built on Polly) has most of these with good default settings:
builder.Services.AddHttpClient<ServiceCClient>()
.AddStandardResilienceHandler();
Follow-up question: When the Circuit Breaker is open, what answer do you give the user?
Follow-up answer: It depends on the kind of work. This is called a fallback:
- For reading data (for example, a price list): we show the last cached data and say “it may be a little old”.
- For non-core parts (for example, product suggestions): we hide that part for a while, but the rest of the page works.
- For writes and important work (for example, placing an order): we return a 503 error at once, with a Retry-After header. We do not pretend the work is done.
- The main point: a fast error is better than a long wait, because threads and connections stay free.
Red flag: Their solution is “we increase the number of retries” or “we make the timeout bigger”.
15Retrying requests in a payment systemHardIdempotencyResilience with Polly
Question: For each purchase, our payment service calls the bank gateway (PSP). The timeout of this call is 10 seconds. If no answer comes, we send the same request again with Polly. If no answer comes after all tries, we mark the order as “failed”. You see this design in code review. What do you think? Does it have a problem? On a busy day when the bank answers slowly, what happens? How do you design calling the payment gateway again?
Short answer: A timeout means “I got no answer”, not “the payment did not happen”. Maybe the bank took the money and only its answer did not reach us. So a blind retry means paying again. The right way: one unique ID for each payment, check the status before trying again, an “unknown” status in the database, and a reconciliation job with the bank report.
Hint (if the candidate says there is no problem): A few days after a busy day when the bank was slow, support reports two kinds of complaints. Some customers say money was taken from their account twice, but they have only one order. Some other customers say money was taken, but their order is “failed” in our system. In the bank transaction report we also see two successful transactions for some purchases. Why?
Why was the money taken twice? (step by step)
- We send the payment request to the bank.
- The bank takes the money, but because it is busy, its answer arrives after 10 seconds.
- After 10 seconds we get a timeout and think the payment did not happen.
- Polly sends the same request again. The bank sees it as a new payment and takes the money again.
- Result: two successful transactions in the bank, but one order in our system.
Why was the money taken but the order is “failed”?
- The bank took the money, but all our tries timed out.
- After the last try, our code marked the order as “failed”.
- This means we treated the answer “I don’t know” as the answer “no”.
The right design:
- A unique ID for each payment. Before calling the bank, we create a payment ID and save it in the database. In all tries, we send this same ID to the bank. Most gateways, if this ID is a repeat, do not create a new payment and return the same earlier result.
- Three statuses, not two. A payment is not only “successful” and “failed”. An “unknown” status is also needed. A timeout means “unknown”, not “failed”.
- First check, then try again. After a timeout, we do not send the payment request again. First we ask the bank with the same ID: “What is the status of this payment?”
- If the bank says it was successful, we mark the order as successful.
- If it says it has not seen this payment, only then do we send it again.
- If the bank still does not answer, the payment stays in the “unknown” status.
- A reconciliation job. Every few minutes, a background job checks the “unknown” payments with the bank. Every day it also compares the bank transaction report with our database. If money was taken but we have no order, it either completes the order or gives the money back (refund).
- Few retries, with gaps. For payments, the number of tries is small, and the gaps grow step by step (exponential backoff). We do not retry in several layers together (like question 14).
- The right message to the user. In the “unknown” status, we do not tell the user “the payment failed”. We say “the payment is being checked”, so they do not pay again themselves.
An important point: We must save the unique ID in the database before calling the bank. If our service restarts in the middle of the work, after it comes back up it knows which payments were started and must be checked.
Follow-up question: If the bank gateway does not support a unique ID and status checks at all, what do you do?
Follow-up answer:
- In this case, we turn off automatic payment retry. The risk of paying twice is bigger than the risk of one failed purchase.
- After a timeout, the payment stays in the “unknown” status.
- Only the daily reconciliation job with the bank report decides the final status.
- If money was taken twice, we give it back with an automatic refund.
- We also tell the product team about this risk, because it affects the user experience.
Red flag: The candidate says “we make the timeout longer” or “we add more retries”, or marks the payment as “failed” after a timeout.
16A database with a Read ReplicaHardSQL and indexesConsistency and CAP
Question: To reduce the database load, we added a Read Replica. All writes go to the main database (Primary) and all reads go to the Replica. How do you see this design? Does it work correctly for all parts of the system? For example, a user edits their profile and presses save. Then the profile page loads again. What happens? If there is a problem, what solutions do you have, and what does each one cost?
Short answer: Changes on the Primary reach the Replica with a small delay. This delay is called Replication Lag. Right after saving, the user reads from the Replica. That is, before the change has reached it. But the user expects to see their own change at once. This need is called Read-Your-Writes.
Hint (if the candidate says there is no problem): Users complain: “I edited my profile, pressed save and saw the success message. But the page still shows the old information.” After a few refreshes it is fixed. In busy hours this complaint is more common. Why?
What happens? (in time order)
- At time 0, the user presses save. The change is stored on the Primary and the API returns the “success” message.
- At 50 milliseconds, the profile page loads again and reads from the Replica.
- The change has not reached the Replica yet. Because replication is usually async and takes from a few milliseconds to a few seconds (more under heavy load). So old data comes back.
- A few seconds later, the change reaches the Replica and the next refresh is correct.
Solutions and the cost of each one:
| Solution | How it works | Cost |
|---|---|---|
| Show from the save response | After saving, the page shows the data from the response of that same save request and does not read again | Almost nothing; but it only fixes that one page |
| Read from the Primary for a few seconds | For, say, 10 seconds after each write, reads of the same user go to the Primary (with a timed flag in a cookie or cache) | Low and simple; the most common way |
| “My own data” pages from the Primary | Pages where users edit their own data always read from the Primary | Part of the load stays on the Primary |
| Wait for the Replica | After a write, we keep the log position (like the LSN in PostgreSQL). We read only from a Replica that has reached this position | Exact, but complex |
| Synchronous (sync) replication | The Primary server waits until the Replica confirms | Writes get slower; if the Replica goes down, writes also get stuck |
- We must also monitor the replication lag and set an alert for it.
Follow-up question: Which parts of the system must never read from the Replica?
Follow-up answer: Any place where we make a decision with the data we read and then write.
- Example: checking an account balance before a withdrawal. If we read from the Replica, we may not see a withdrawal from a few seconds ago. So we may allow a withdrawal bigger than the balance.
- Other cases: permissions and password changes, creating unique numbers, and any read that is inside a write transaction.
Red flag: The candidate does not know the idea of replication lag, or says “the problem is the browser cache”.
17Changing the database schema without DowntimeDesignWeight ×2Entity Framework CoreCI/CD
Question: The customers table has a “full name” column (FullName). We want to turn it into two columns: “first name” (FirstName) and “last name” (LastName). The service has 4 pods and deploys are rolling. This means the pods are replaced one by one, and for a few minutes the old and new versions of the code run together. A developer wrote a migration that does three things at once: it creates the two new columns, copies the data, and drops the full name column. You see this migration in code review. Does this migration cause a problem during the deploy? Why? If there is a problem, how do you do it step by step?
Short answer: During a rolling deploy, the old and new code work together. So at every moment, the database must be compatible with both versions of the code. The solution is the Expand and Contract pattern:
- First, add the new thing.
- Keep both for a few steps.
- At the end, remove the old thing.
Hint (if the candidate says there is no problem): We did this same thing in an earlier deploy. In the middle of the deploy, some requests got a 500 error. In the log of the old pods we saw the error “column FullName does not exist”. Then the new version had a bug and we wanted to roll back. But the old version gave the same error and did not work. Why?
Why does a one-shot migration give errors?
- When the migration runs, the full name column is dropped.
- 3 pods still have the old code and read the full name column. The database gives the error “column does not exist” and the user gets a 500 error.
- If we do it the other way around, deploying the code first and then the migration, the new pods look for the first name and last name columns. But these columns do not exist yet.
- If the new version has a bug, rollback is also not possible. Because the old code needs the full name column, and that column is dropped.
Solution: Expand and Contract in 6 steps
| Step | Change | Where does the code read from? | Where does the code write? |
|---|---|---|---|
| 1 | Add the new columns as nullable (the code does not change) | FullName | FullName |
| 2 | Deploy code version 1 | FullName | Both |
| 3 | Fill old records in small batches (backfill) | FullName | Both |
| 4 | Deploy code version 2 | New columns | Both |
| 5 | Deploy code version 3 | New columns | Only the new columns |
| 6 | After a few days and when sure, drop the FullName column | New columns | Only the new columns |
- Why is this way safe? Because at every step, the previous version of the code still works with the database. So the rolling deploy has no errors. And going one step back (rollback) is always possible.
A few important points:
- Fill the data (backfill) in small batches, for example 1000 rows each time. One big UPDATE on millions of rows locks the table for a long time and fills the transaction log.
- Run the migration separately from app startup, for example with a Job in the pipeline. If all 4 pods call the Migrate method at startup, they may run at the same time and run into problems.
- Be careful with heavy ALTERs. Some changes (like changing a column type) on a big table rewrite and lock the whole table.
Follow-up question: After step 4, a bug was found and we must go back to version 1. Does a problem happen? What if it is after step 6?
Follow-up answer:
- After step 4 there is no problem. Version 1 reads from the full name column. This column still exists and is still being filled.
- After step 6, the old column is dropped. So going back to version 1 is not possible.
- That is why step 6 is the last step and is done separately, after some time.
Red flag: The candidate writes a migration with rename or drop and says “we deploy quickly, a few seconds of errors does not matter”.
18Exporting a large amount of dataMediumEntity Framework CorePerformance and Span
Question: In the admin panel, a user clicks “Download Excel” to get 2 million transaction records. The current endpoint does three things: it reads all records at once with the ToListAsync method, builds the Excel file in memory, and returns it in the response. The memory limit of each pod is one gigabyte. You see this code in code review. What do you think? In Production, with this amount of data, what happens? If needed, how do you redesign this feature?
Short answer: We have two separate problems:
- All the data is in memory at once.
- A job of several minutes is done inside one HTTP request.
- Solution: move the work to the background, read and write the data piece by piece, and give the ready file to the user later.
Hint (if the candidate says there is no problem): After about 100 seconds, the user sees a timeout error. Sometimes the pod also restarts with the OOMKilled status (out of memory). At that same moment, the requests of other users are also cut. Why?
Why do we get OOM?
- The ToListAsync method builds all 2 million records as C# objects in memory. If each object with its strings is about 500 bytes, that is about 1 gigabyte. The change tracker also keeps extra information for each object.
- Then the Excel file is also built fully in memory.
- If the pod memory limit is, for example, 1 gigabyte, Kubernetes kills it. All other requests that were on the same pod are also lost.
Why do we get a timeout? There are several layers between the browser and the app: the load balancer, the ingress and the reverse proxy. Each of them has a maximum time for one request. A job that takes several minutes should not be inside a request at all.
The right design (step by step):
- First, the endpoint only records an “export Job”. Then it returns the code 202 (Accepted) at once, together with a Job ID.
- A worker in the background (a BackgroundService or a queue consumer) runs this Job.
- It reads the data piece by piece, not all at once. With the AsNoTracking and AsAsyncEnumerable methods, each record is read, written and released from memory:
await foreach (var row in db.Transactions.AsNoTracking().AsAsyncEnumerable().WithCancellation(ct))
{
writer.WriteRow(row);
}
- The file is also written as a stream (for example, with OpenXmlWriter). If the business accepts it, the CSV format is much simpler and lighter.
- The file is saved in a storage (like S3 or MinIO), not in the memory or disk of the pod.
- So that the load does not go to the main database, we read from a Read Replica.
A good point (a plus if the candidate says it): Each sheet in Excel has at most 1,048,576 rows. So 2 million rows do not fit in one sheet at all.
Follow-up question: How does the user know their file is ready? How do you make the download link secure?
Follow-up answer:
- To notify: the page asks the server for the Job status every few seconds (polling). Or the server notifies with SignalR, email or a notification.
- For security: the link is a Pre-signed URL with a short life (for example, 15 minutes). This link is created only after checking that this same user owns this Job.
- The files are deleted automatically after a few days, because financial data is sensitive.
Red flag: Their answer is “we make the timeout longer” or “we give the pod more memory”.
19Request limiting (Rate Limiting) on several serversMediumRate limitingRedis
Question: We have a public API. By contract, each customer is allowed at most 100 requests per minute. We set this limit with the built-in ASP.NET Core middleware (the AddRateLimiter one). In Production, the service runs on 10 pods, and a load balancer spreads the requests between them. How do you see this design? In Production, is the same limit of 100 requests per minute kept? If not, how do you fix it?
Short answer: The built-in middleware keeps the counter in the memory of the same pod. So:
- Each pod counts up to 100 on its own and knows nothing about the other pods.
- The load balancer service also spreads the requests between 10 pods.
- So the real limit becomes 10 × 100 = 1000.
- Solution: the counter must be in a shared place.
Hint (if the candidate says there is no problem): In the test on a laptop, everything worked correctly. But in Production monitoring we see that some customers send about 1000 requests per minute. These customers get no 429 error. Why?
Ways:
- A shared counter in Redis. The simplest version is the Fixed Window:
- The key is built from the customer ID and the current minute. For example, “customer 42 at 10:30”.
- With the INCR command we add one to the counter. The first time, with the EXPIRE command, we set an expiry time of 60 seconds. It is better to put both in one Lua script, so they run atomically.
- If the number goes above 100, we return the code 429.
- A limit at the edge of the system. An API Gateway or Ingress is in front of all pods. It counts right here, in one place.
- An approximate way. The share of each pod becomes 100 ÷ 10 = 10. This way is simple, but it has two problems: it must change when the number of pods changes. And if the load balancer does not spread the load evenly, it is not exact.
The difference between algorithms (an important point):
- In the Fixed Window, a customer can send 100 requests in second 59 of the first minute and 100 requests in second zero of the next minute. That is 200 requests in two seconds.
- The Sliding Window and Token Bucket algorithms do not have this problem.
- In the Token Bucket, each customer has a bucket of tokens. The bucket fills at a fixed speed. Each request takes one token. If there is no token, the request is rejected.
The right details:
- The right answer is the code 429 (Too Many Requests) with a Retry-After header. With this header, the client knows when to try again.
- The limit key must be the API Key or the customer ID, not the IP. Because several customers behind one NAT share one IP. And one customer may also have several IPs.
Follow-up question: If the Redis that keeps the counters becomes unavailable, do you accept or reject the requests? Why?
Follow-up answer: It depends on why the limit was set.
- If the goal is to protect the system, we usually accept the request (fail-open). Together with a simple local limit as a backup. Because we do not want Redis going down to take down the whole API.
- If the limit is financial or for security, maybe rejecting (fail-closed) is right. For example, when each request costs us money, or the limit stops password guessing.
- A good candidate says “it depends” and explains why.
Red flag: The candidate does not know that the built-in limiter counts separately in each pod.
20Revoking access with JWTMediumAuthentication and authorization
Question: Our system uses JWT for authentication. Each token lives 1 hour. The security manager deactivates the account of an abusive user in the admin panel (the user’s active field becomes false in the database). In your opinion, after this, what happens to this user’s access to the API? Does this design have a problem? If yes, what ways do you have, and what does each one cost?
Short answer: To check a JWT, the server does not look at the database at all. It only checks the signature and the expiry date of the token itself. So deactivating the user in the database has no effect on a token that was already issued. This token works until it expires.
Hint (if the candidate says there is no problem): In API monitoring we see that after being deactivated, this user still calls the API. The responses are also successful. This goes on until about one hour later. Why?
How does JWT work?
- At login, the server creates a token. The user information (ID and roles) and the expiry time are inside the token. The server signs the token with its own key.
- In each request, the server only checks the signature and the expiry. It does not connect to any database. That is why it is fast. This is called stateless.
- This same benefit is the problem here. The token is valid until it expires, no matter what happens in the database.
Ways and the cost of each one:
| Way | How it works | Cost |
|---|---|---|
| Short life + Refresh Token | The access token is valid for only 5 to 15 minutes. For a new token, the client sends a Refresh Token. At that moment the server checks the user status in the database | Simple, but there is still a gap of a few minutes |
| Denylist | We keep the token ID (the jti field) or the user ID in Redis until expiry. We check it in each request | One fast Redis check in each request |
| Security stamp | A version number is inside the token. When the user is deactivated, this number changes in the database. Old tokens are rejected | Must be checked in each request (with a short cache) |
| Reference Token | The token is only an ID. The server asks the Identity Server each time | Exact, but has more load and delay |
The usual choice is a mix of the first and second ways. Most of the time a short life is enough. For emergencies (like an abusive user), we use the denylist.
Follow-up question: Where do you keep the Refresh Token? If it is stolen, how do you find out?
Follow-up answer:
- On the server, we keep it as a hash in the database (like a password). If the database leaks, the tokens themselves do not leak.
- In the browser, we put it inside a cookie, with the HttpOnly, Secure and SameSite settings. Not in localStorage. The reason is simple:
- If the site has an XSS weakness, the attacker’s JavaScript code runs.
- This code can read localStorage.
- But it cannot read a cookie with HttpOnly.
- To detect theft, we use Refresh Token Rotation:
- Each time a Refresh Token is used, a new Refresh Token is given and the old one is revoked.
- If an attacker has stolen the token, sooner or later one of the two sides (the user or the attacker) sends a revoked token.
- When the server sees a revoked token, it knows theft has happened. So it revokes all tokens of that session.
- The real user only logs in one more time.
Red flag: The candidate thinks that logout, or deleting the token in the browser, also revokes the token on the server.
21Two concurrent requests with one Idempotency KeyHardIdempotency
Question: We built an Idempotency Key for a payment API. This is how it works now: (1) We look in the database to see if this key has a result. (2) If it has one, we return that result. (3) If it does not, we make the payment and save the result with the key. The service runs on several pods. You see this design in code review. What do you think? What happens if two requests with the same key arrive at almost the same moment?
Short answer: The “check first, then do the work, then save” pattern has a race condition. Both requests check before the other one saves anything. The right way is this: first record the key (with a unique constraint), then do the work.
Hint (if the candidate says there is no problem): Because of a bug, the mobile app sometimes sends two requests with the same key at almost the same moment. In these cases, the customer’s account is charged twice. Now, why?
What happens? (in time order)
| Time | Request 1 | Request 2 |
|---|---|---|
| 0 ms | Checks the key: not found | |
| 2 ms | Checks the key: not found | |
| 5 ms | Starts the payment | Starts the payment |
| 800 ms | Saves the result | Saves the result |
Neither request found anything at check time. This is because the result is saved only after 800 ms. So both of them made the payment.
Solution: record the key first
- The keys table has a unique constraint on two columns: the user ID and the key.
- The first request inserts a row with the status “in progress”. Then it makes the payment.
- The second request also tries to insert the same key. The database gives it a unique error. The database guarantees that only one succeeds, even if both arrive in the same millisecond and on two different pods.
- The second request returns a 409 (Conflict) response. This means “it is in progress, ask again later”. Or it waits a little and returns the result of the first request.
- When the payment ends, the status “done” and the result (status code and response body) are saved.
db.IdempotencyKeys.Add(new IdempotencyKey(userId, key, Status.InProgress, requestHash));
try
{
await db.SaveChangesAsync(ct);
}
catch (DbUpdateException ex) when (IsUniqueViolation(ex))
{
return Results.Conflict("Request is already being processed.");
}
// Only one request reaches this point
Two more points:
- Also save a hash of the request body. If the same key comes with a different amount, return an error (for example, code 422). This is a sign of a bug in the client.
- Delete keys after some time (for example, 24 hours).
Follow-up question: What happens if the first request crashes in the middle and the key stays “in progress” forever?
Follow-up answer: The user can no longer pay with this key. The solution:
- Also save the start time. If a key stays “in progress” longer than a limit (for example, 5 minutes), the work has stopped halfway.
- But before repeating the payment, we must find out if the first payment really happened. So we ask the payment gateway with the same ID. This is called reconciliation.
- If the payment happened, save the result and return it. If not, do it again.
Red flag: Suggests using the lock statement in C#. This statement works only inside one process. It does not stop two requests on two different pods.
22Event order from several serversHardKafkaConsistency and CAP
Question: Three services on three different servers create the events of an order: “order created”, “payment done”, “shipped”. Each service puts a timestamp on the event using its own server clock. All servers sync with NTP. A reporting service sorts these events by timestamp and then runs the report calculations. What do you think of this design? Is this order always reliable?
Short answer: Server clocks are not exactly the same. This difference is called Clock Skew. When two events happen a few milliseconds apart on two different servers, their timestamps are not reliable. To decide the order, you must use something other than the clock.
Hint (if the candidate says there is no problem): Sometimes in the report output, “payment done” appears before “order created”. In these cases, the report calculations are wrong. Now, why?
A numeric example:
- The order server clock is 100 ms ahead. The payment server clock is exact.
- The order is created at the real time “10:00, millisecond 0”. But its timestamp is recorded as “10:00, millisecond 100”.
- The payment happens 50 ms later, at “10:00, millisecond 50”. It also gets this as its timestamp.
- Now, if we sort by timestamp, the payment comes before the order.
Even with NTP, clocks differ by a few milliseconds (sometimes more). A server clock may even jump backward during sync.
Options:
- A version number from a single source: The order has a version number in its own database (1, 2, 3 and so on). Each event carries the order’s version number. This number is created in only one place, so its order is correct.
- Order in Kafka: Send all events of an order to one topic with the order ID as the key. They all go to one partition. The offset number shows the real order of arrival.
- An explicit cause and effect link: The “payment” event carries the ID of the event that caused it (the order). This ID is called a causation id. The report knows this event comes after that one, whatever timestamp it has.
- A logical clock (Lamport Clock or Vector Clock): It is enough for the candidate to know this option exists.
Rule: A timestamp is good for display and rough analysis. But it is not right for deciding the exact order between servers.
Follow-up question: How do you store time in the database? If a Job must run “at 00:00 in each customer’s local time” for customers in several countries, what should you watch for?
Follow-up answer:
- Store time in UTC. The DateTimeOffset type is better than DateTime. This is because the offset is kept with the time, so nothing is ambiguous.
- For “local midnight”, keep each customer’s time zone as a standard IANA ID (for example, the Madrid zone). Then calculate with the TimeZoneInfo class, not with a fixed offset. The reasons:
- With daylight saving time, the offset changes.
- Sometimes one hour repeats twice.
- Sometimes one hour does not exist at all.
- To make the code testable, use the TimeProvider class instead of reading the current system time directly.
Red flag: Says “we sync all servers with NTP and the problem is solved”.
23Hot partition in KafkaHardKafka
Question: We have a topic with 12 partitions. A service with 12 pods (in one consumer group) reads it. The message key is the customer ID. One big customer creates about 70% of all messages. What do you think of this design? What happens in Production with this load? If the load grows, is adding more partitions and pods enough?
Short answer: Kafka sends all messages with the same key to one partition. Each partition is also given to only one consumer. So 70% of the work is on one pod. More partitions do not help, because this customer’s key is still one key. We must change the key or spread this customer’s load in a different way.
Hint (if the candidate says there is no problem): In monitoring, we see one pod is several hours behind (consumer lag is high). The other eleven pods are almost idle. The team raised the number of partitions and pods to 24, but nothing changed. Now, why?
Why does this happen? (step by step)
- Kafka has a simple formula to choose a partition: it takes the hash of the key and calculates the remainder of dividing it by the number of partitions. So one customer always goes to one partition. This is on purpose, to keep the order of each customer’s messages.
- In a consumer group, each partition is given to only one consumer (one pod) at any moment.
- So all messages of the big customer, which is 70% of all work, reach only one pod. This problem is called a Hot Partition.
- With 24 partitions, the big customer is still one key. So it still has only one partition and one pod.
Options:
- A finer key: Maybe order is needed only inside one order, not for the whole customer. In this case, use the order ID as the key. The big customer’s orders spread across all partitions. This is the best and simplest option.
- Salting: Use the customer ID plus a number from 0 to N as the key (for example, “customer 42, number 3”). This customer’s load spreads over N partitions. But the overall order of this customer’s messages is lost.
- A separate topic for big customers, with its own consumers.
- Parallel processing inside the same pod: The pod spreads messages between several internal workers by order ID. This option is harder, because committing the offset gets complex. You can commit only up to a message whose earlier messages are all finished.
- Faster processing of each message: For example, writing to the database in batches instead of one by one.
For monitoring, look at the lag of each partition separately, not only the total for the whole group. The total hides this problem.
Follow-up question: Does your solution break the order of this customer’s messages? Does it matter or not?
Follow-up answer:
- With Salting, yes. The overall order is lost.
- With the order ID as the key, the order of messages in each order is kept. But the order between different orders of one customer is not kept.
- A good candidate first asks: “Where is order really needed?” For example, for the customer’s account balance, the overall order may matter. But for shipping, it does not.
Red flag: Their only solution is “more partitions and consumers”.
24Data that lives in another serviceDesignWeight ×2MicroservicesDomain-Driven Design
Question: The “order list” page shows 50 orders. For each order, the customer name and the product name are also needed. Orders are in the Order service, customers in the Customer service, and products in the Catalog service. Each one has its own database. Right now, for each order, the Order service calls the Customer service once and the Catalog service once over HTTP. What do you think of this design? Does it have a problem? If you want to improve it, what options do you have, and which one do you choose?
Short answer: This is the same N+1 problem (question 3), this time with HTTP calls instead of queries. 50 orders times 2 calls is 100 network calls, one after another. If each one takes about 40 ms, the total is 4 seconds. The simplest improvement: instead of 100 calls, make 2 batch calls in parallel.
Hint (if the candidate says there is no problem): Loading this page takes about 4 seconds. Each HTTP call to another service takes about 40 ms.
Options, from simple to complex:
1. Batch calls in parallel (API Composition):
- We collect the customer IDs and remove duplicates. We do the same for products.
- We get all customers with one request. We also get all products with one request.
- We send these two requests at the same time, not one after another.
var customerIds = orders.Select(o => o.CustomerId).Distinct();
var productIds = orders.Select(o => o.ProductId).Distinct();
var customersTask = customerClient.GetByIdsAsync(customerIds, ct);
var productsTask = catalogClient.GetByIdsAsync(productIds, ct);
await Task.WhenAll(customersTask, productsTask);
We go from 100 calls one after another to 2 calls at the same time. The time drops from 4 seconds to about 50 ms. A short cache also helps. This is usually the first step.
2. Copy the data at save time (Snapshot): When the order is created, the customer name and the product name and price are saved in the order itself. Then no service is needed to show the list. For invoices and financial documents, this is often the most correct choice. This is because the price at purchase time matters.
3. A local copy with events:
- The Customer service publishes a “customer name changed” event on every name change.
- The Order service has a small table: customer ID and name.
- The Order service updates its table with these events.
Reading is fast and does not depend on the Customer service. But the data is a few seconds late (eventual consistency) and it needs more code.
4. A separate read model (Read Model in CQRS): We have a database or index just for this page (for example, Elasticsearch). This index is filled from the events of all three services. It is good for search and complex filters. But it is the most expensive option.
Which one should we choose?
- Usually option 1 first.
- Option 2 for data that must stay “historical”.
- Option 3 or 4 when the load is high or complex search is needed.
- If these three pieces of data are almost always needed together, maybe the service boundaries are drawn wrong. A candidate who says this has good depth.
Follow-up question: If a customer changes their name, should their old orders show the new name or the old name? Who makes this decision?
Follow-up answer: This is a business decision, not a technical one. You must ask the Product Owner. Usually, an invoice must keep the name and price at purchase time (option 2). For display in a panel, the new name may be better. The main point is that the candidate should not decide alone.
Red flag: Suggests that the Order service connects directly to the Customer service database, or makes a JOIN between databases.
25When compensation fails in a SagaDesignWeight ×2The Saga pattern
Question: Placing an order in our microservice system is a three-step Saga: (1) the warehouse service reserves the item, (2) the payment service takes the money from the customer’s card, (3) the shipping service books a courier. One day step 3 fails (for example, the address is outside the delivery area). So the Saga must give the money back to the customer (refund) and release the warehouse reservation. But the payment service is also down right now, and the refund fails. What should happen? How do you design the Saga so that no order is lost in an unknown state and no customer’s money is lost?
Short answer: Compensation can also fail. So we must have a plan for it:
- The Saga state is saved in the database.
- The compensation is retried until it succeeds.
- If it still fails, the work goes to manual review.
Best of all is to design the order of the steps so that a refund is needed less often.
First: what is a Saga?
- In microservices, we have no shared transaction between services.
- A Saga is several steps, one after another. Each step is a local transaction in its own service.
- Each step has a “compensating action”. For example, the compensating action of “reserve” is “release the reservation”.
- If a step fails, the compensating actions of the earlier steps run in reverse order.
When the refund fails, what should happen?
- The Saga state is in the database, not only in memory. For example, the order is in the state “waiting for refund”. If the service restarts, it knows where the work stopped.
- Retry with growing delays: The refund is tried again after 1 minute, 5 minutes, 30 minutes and so on, even if it takes hours.
- The compensation must be idempotent: Each refund has a unique ID. Maybe one attempt succeeded but its response was lost. With this ID, the next attempt does not refund the money twice.
- Manual review: If it still has not succeeded after a set limit (for example, 24 hours), the order goes to the state “needs review”. The message goes to a Dead Letter Queue, and the finance team gets an alert.
- A time limit (timeout) for each step: Sagas that have stayed in one step for a long time are found and reported.
- The real state is shown to the customer: “Order canceled, refund in progress”.
- A nightly comparison (reconciliation) between our system and the payment gateway, as the last line of safety.
A better order of steps (the golden point):
- In step 2, we do not take the money. We only hold it (Authorize).
- After step 3 succeeds, we take the money (Capture).
- If shipping fails, only the held money is released. This is much simpler than a refund. It is usually released automatically after a few days too.
General rule: the step that is hardest to undo should come last.
Follow-up question: Do you build the Saga with Orchestration or Choreography? What is the difference?
Follow-up answer:
- Orchestration: A central service (the orchestrator) is like a project manager. It tells each service what to do and keeps the state. Seeing the state, handling timeouts and compensation are simpler. For multi-step Sagas like this one, it is usually better.
- Choreography: There is no central manager. Each service reacts to the previous event. For example, the warehouse reserves on “order placed”, and payment takes the money on “reserved”. There is less coupling. But it gets hard to know “where is this order in the process right now?”.
- In .NET, tools like MassTransit (with State Machine Saga) have this pattern ready.
Red flag: Suggests a distributed transaction (2PC) between services, or has no plan for when the compensation itself fails.
26Scaling in Kubernetes: horizontal and verticalMedium to hardKubernetes
Question: We have a .NET API in Kubernetes. These are the settings of each pod:
- The CPU request is 500m (half a core).
- The CPU limit is 1 core.
- The Memory request and the Memory limit are both 512Mi.
We also have an HPA. When CPU use reaches 70%, it adds a pod. At peak hours, we see these three things in monitoring:
- Response time goes up. But the HPA does not create any new pod, because CPU is about 40%.
- Some pods restart with the status OOMKilled.
- Sometimes new pods are created, but they stay in the Pending status for a few minutes.
What is the difference between horizontal and vertical scaling? What exactly do request and limit do? How do you explain each of these three events, and what do you change?
Short answer: Horizontal means more pods. Vertical means more resources for each pod. The request is a reservation and the limit is a cap on use. The three events happen for these reasons:
- Event 1: CPU is not the bottleneck. So the HPA, which looks at CPU, does nothing.
- Event 2: Memory use went over the limit.
- Event 3: No node has enough free space for the request.
First: horizontal and vertical
| Horizontal (HPA) | Vertical (VPA) | |
|---|---|---|
| What it does | Adds more pods | Gives each pod more CPU and memory |
| Condition | The app must be stateless (session in memory or a local file causes problems) | Usually the pod must restart |
| Limit | Almost no cap (up to cluster capacity) | Has a cap: the size of one node |
Note: Do not use HPA and VPA together on the same metric (for example, both on CPU). They conflict with each other.
Second: request and limit
- The request is what Kubernetes reserves for the pod. The pod is placed only on a node that has this amount free.
- Important point: the CPU percent in the HPA is calculated relative to the request, not relative to the limit.
- The limit is the cap on use. CPU and memory behave differently at this cap:
- When CPU reaches the limit: throttling happens. The pod is not killed, it only gets slower.
- When memory goes over the limit: OOMKilled happens. The pod is killed.
Explaining all three events
Event 1 (slow with CPU at 40%):
- When CPU is low but the app is slow, the app is waiting for something. For example, the database, another service, or the thread pool (like question 6).
- The HPA does not see this waiting, because it looks only at CPU.
- If the bottleneck is the database, more pods only add more load to the database.
- Another cause can be throttling. The average CPU is low, but in short moments the pod reaches the limit and is paused.
- To see this, check the throttling metric in Prometheus (the cfs throttled periods metric for the container).
- Solution: first find the real cause. If load goes up with the number of requests, put the HPA on a better metric: requests per second, latency, or queue length (for example, with KEDA).
Event 2 (OOMKilled):
- In .NET, the GC sees the container memory limit. By default, it uses about 75% of it for the heap.
- The rest of the memory is for native memory, stacks and other things.
- First check for a memory leak (question 35).
- If there is no leak and the app really needs more memory, raise the limit based on real use.
Event 3 (Pending):
- The Pending status means the scheduler found no node with this request amount free.
- One option: the Cluster Autoscaler adds a new node. This takes a few minutes.
- Another option: make the request realistic. A very large request wastes capacity. A very small request makes many pods pile up on one node.
Some .NET points and a common setup
- In .NET, the number of processors (ProcessorCount) comes from the CPU limit. With a limit of 1, the thread pool and the GC think there is only one core.
- Common setup: the Memory request equals the Memory limit, so behavior is predictable. The CPU request is based on real use.
- Some teams set the CPU limit high, or do not set it at all, to avoid throttling. This is debated. A good candidate explains its costs and benefits.
- There are three QoS classes: Guaranteed (when the request equals the limit), Burstable and BestEffort. When a node is short on memory, BestEffort pods are evicted first, then Burstable pods.
Follow-up question: When do you prefer vertical scaling over horizontal? If the traffic peak is every day at 9 AM and the HPA reacts late, what do you do?
Follow-up answer:
- Choose vertical for things that cannot have several copies. For example, a database, a stateful service, or an old app that keeps state in memory.
- For a predictable peak: a little before 9, raise the minimum number of pods on a schedule (for example, with the cron scaler in KEDA).
- Reduce pod startup time: a smaller image, ReadyToRun, and warm-up before readiness turns positive.
- The reason: the HPA always reacts with some delay by design.
Red flag: Mixes up throttling and OOM. Or thinks the HPA percent is relative to the limit. Or their only answer is “more pods”.
27How Kafka partitions relate to podsMediumKafkaKubernetes
Question: We have a topic with 6 partitions. A consumer service with 4 pods reads it. All pods are in one consumer group. Monitoring shows two pods are busier than the others. How do partitions and pods relate? How is the load split between these 4 pods? What happens if the HPA raises the number of pods to 10? What if only 1 pod is left?
Short answer: The main rule is this:
- In a consumer group, each partition is given to only one consumer at any moment.
- But one consumer can have several partitions.
- So the largest useful number of pods equals the number of partitions.
Three cases
| Case | What happens |
|---|---|
| 6 partitions and 4 pods | Two pods get 2 partitions each, and two pods get 1 partition each. The load is uneven. That is why two pods are busier. |
| 6 partitions and 10 pods | Six pods get one partition each, and four pods are idle. Only when a pod dies does an idle one take its place. It does not get faster. |
| 6 partitions and 1 pod | That one pod gets all 6 partitions. It works, only slower. |
Practical results
- The maximum number of pods in the HPA (the maxReplicas setting) should not be more than the number of partitions. Extra pods only waste resources.
- For an even split, choose a partition count that divides evenly by common pod counts. For example, 12 partitions split evenly with 2, 3, 4, 6 or 12 pods.
- To scale consumers, the right metric is consumer lag, which means the number of unread messages. CPU is not the right metric. With KEDA, you can scale based on lag.
- You can add partitions later, but you cannot remove them.
- Adding partitions moves some keys to other partitions. So the order of messages for those keys may get mixed up at that moment.
- So leave some room for growth from the start.
The idea of rebalance
- Every time the number of pods changes (scale, deploy or crash), Kafka splits the partitions between the pods again. This is called a rebalance.
- During this time, consuming messages slows down or stops.
- To reduce its effect, use the CooperativeSticky strategy. With it, only the needed partitions move, not all of them.
- For short restarts, turn on static membership. This means each pod has a fixed ID in the group.
A .NET point: in the Confluent.Kafka library, you cannot use one consumer from several threads at the same time. If we create several consumers in one pod, each one counts as a separate member of the group and gets its own partitions.
Follow-up question: In the middle of processing a message, a rebalance happens and the partition is taken from this pod. What problem comes up, and how do you handle it?
Follow-up answer:
- The problem: a message that was processed but whose offset was not yet committed is given to another pod. So it is processed twice.
- First solution: make processing idempotent (question 5).
- Second solution: Kafka calls a handler before taking the partition (in Confluent.Kafka, with SetPartitionsRevokedHandler). In this handler, finish the current work and commit the offset.
Red flag: Thinks two pods in one group can read one partition together and split the load in half. Or thinks more pods always means more speed.
28Using Task or ThreadMediumasync/await and the thread pool
Question: Think about two cases:
- Case A: An API gets a very large number of requests (for example, 10 thousand concurrent requests). Each request does some I/O work (database, HTTP). Is it better to use Task (that is, async/await) or to create a Thread for each request? Why?
- Case B: A service has 10 background jobs that are always running. For example, listening to a socket or reading a price stream. Are 10 Tasks better for these, or 10 always-running Threads?
Short answer:
- For case A, always async/await. This is because no thread is held while waiting for I/O.
- For case B, the answer depends on whether the code inside the loop is async or blocking.
- If the code is async: a normal Task.
- If the code is blocking: a dedicated thread, not the thread pool.
The difference between Thread and Task, in simple words
- A Thread is an operating system resource. Each thread has its own stack memory (usually about 1 MB reserved). Creating it has a cost. Switching between threads (context switch) also takes CPU time.
- A Task is a “job”, not a thread. Tasks run on threads from the thread pool. When a Task is waiting for I/O (await), it holds no thread.
Case A: why async/await?
- If we create a thread for each request, 10 thousand requests means 10 thousand threads.
- This means about 10 GB of stack space. The CPU also spends most of its time switching between threads.
- Most of these threads are only waiting for the database and do nothing.
- With async/await, when a request is waiting for the database, the thread is freed and moves another request forward.
- So a few dozen threads can serve thousands of concurrent requests.
- ASP.NET Core itself uses the thread pool. Our job is only to write the code as async and not block it by reading Result (question 2).
Case B: it depends on whether the code is async or blocking
If the code is async (like the ReceiveAsync methods on a socket, ReadAsync on a stream, or the System.IO.Pipelines library):
- Ten normal Tasks (for example, 10 BackgroundServices) are the best choice.
- When no data has arrived, no thread is held. So these 10 Tasks cost almost nothing.
If the code is blocking (a library with only sync methods, a blocking Read method, or an endless loop of heavy calculation):
- You should not run it as a normal Task on the thread pool. The reason, in a few steps:
- These 10 jobs hold 10 threads of the pool forever.
- If the pool has, for example, 8 initial threads, nothing is left for API requests.
- On the other hand, the thread pool adds each new thread slowly.
- The result: the whole app gets slow. This is called thread pool starvation.
- The right way: create a dedicated thread for each job. There are two ways:
// Way 1: a background Thread
var t = new Thread(Listen) { IsBackground = true };
t.Start();
// Way 2: a separate thread, outside the pool
Task.Factory.StartNew(Listen, TaskCreationOptions.LongRunning);
If the work is CPU-heavy:
- Having more active threads than cores does not help. Ten heavy threads on 2 cores only create more context switches.
- A better pattern has three parts: one part that reads the data, an in-memory queue (System.Threading.Channels), and a number of workers equal to the number of cores.
In all cases: A CancellationToken is needed for a safe shutdown (question 11).
Follow-up question: Why is this code a common mistake?
Task.Factory.StartNew(async () => await ListenAsync(ct), TaskCreationOptions.LongRunning);
Follow-up answer:
- The LongRunning option creates a dedicated thread. But the code runs on that thread only until the first await.
- After the first await, the rest of the code moves to the thread pool. The dedicated thread ends up useless.
- So LongRunning makes sense only for blocking code.
- Second problem: StartNew with an async lambda returns a nested Task (a Task that has another Task inside it).
- If we do not open it with the Unwrap method, await waits only until the first inner await. Later errors are also lost.
- For async code, it is enough to call the method directly, or to use the Run method of the Task class. This method unwraps the nested Task by itself.
Red flag: Says “Task is faster than Thread” or “Task always means a new thread”. Or suggests one thread for each request.
29Moving to Redis ClusterHardRedis
Question: We are moving from a single Redis to a Redis Cluster with 3 primary nodes. The current system has these two things:
- We have a Lua script that lowers the stock of several products together, atomically (the stock keys of products 1, 2 and 3). It works correctly on a single Redis.
- We have one very popular key: the home page settings that every request reads.
After moving to the Cluster, how do these two behave? Will there be a problem? If yes, what is the cause and how do you fix it?
Short answer: In a Cluster, each key lives on one specific node.
- Problem 1: the keys of this script are on different nodes. So Redis cannot run it atomically.
- Problem 2: all requests for one key go to that same node. The Cluster does not spread the load of one key.
Hint (if the candidate says there is no problem): After the move, the Lua script returns a CROSSSLOT error. The error text says the keys of this request are not in one slot. In monitoring, we also see the CPU of one node has reached 100%, but the other two nodes are almost idle.
How Redis Cluster spreads keys
- The key space is split into 16384 parts. Each part is called a hash slot. Each node owns some of the slots.
- For each key, Redis takes a hash of the key name (with CRC16) and calculates its remainder when divided by 16384. This number is the key’s slot.
- So the stock key of product 1 and the stock key of product 2 usually have different slots and different nodes.
- Multi-key commands (like MGET, a MULTI transaction and a Lua script) run only on one node.
- If the keys are not in one slot, Redis gives a CROSSSLOT error.
Solution to problem 1: Hash Tag
If part of the key name is inside curly braces, only that part is used to calculate the slot:
stock:{sale42}:p1
stock:{sale42}:p2
stock:{sale42}:p3
- Now all three keys are in one slot, and the script works again.
- Be careful: a hash tag gathers all those keys on one node. If you use it too much, it creates a hot node by itself.
- Another option: design the data so that the atomic operation is on only one key. For example, one Hash with several fields, one field for each product.
Problem 2: Hot Key
One key is always on one node. If all requests read the same key, they all go to that node. Adding nodes does not help either. We have three options:
- A short local cache in each pod (for example, 5 seconds): most reads never reach Redis. For data that changes rarely, this is the simplest and best option.
- Several copies of the key: for example, we create 8 settings keys numbered 1 to 8 with the same value. Each request reads one of them at random. So the load spreads between nodes.
- Reading from a replica: in StackExchange.Redis, with the PreferReplica flag. We must accept a small delay in the data.
To find a hot key: the redis-cli tool with the hotkeys option (if the eviction policy is LFU), or comparing the metrics of each node.
Follow-up question: What is a “Big Key”? Why is deleting a Hash with one million fields using the DEL command dangerous?
Follow-up answer:
- In Redis, commands run on one main thread, one after another.
- Work on a very large key (like DEL, HGETALL or SMEMBERS) may take a few seconds.
- During this time, all clients wait.
- Solution: keep keys small.
- To delete, use UNLINK instead of DEL. UNLINK deletes in the background.
- To read, use HSCAN and SSCAN instead of a full read, so the data is read piece by piece.
Red flag: Thinks the Cluster spreads the load of each key between nodes by itself.
30Deadlock in the databaseMedium to hardTransactions and isolation levels
Question: We have a money transfer service. Each transfer does two things in one transaction:
- It lowers the balance of the source account.
- It raises the balance of the destination account.
At peak hours, many transfers between accounts run at the same time. What do you think of this design? What happens under this concurrent load? Does it have a problem?
Short answer: Each transaction has locked one of the two accounts and is waiting for the other one. Neither can move forward. This is called a deadlock. The main solution: all transactions take the locks in a fixed order.
Hint (if the candidate says there is no problem): At peak hours, the API returns this error for some transfers: “Transaction was deadlocked on lock resources with another process and has been chosen as the deadlock victim”. The investigation shows this error usually happens when a transfer from account A to B and a transfer from B to A run at the same time.
What happens? (in time order)
| Time | Transaction 1 (A to B) | Transaction 2 (B to A) |
|---|---|---|
| 1 | Updates account A and takes the lock on A | Updates account B and takes the lock on B |
| 2 | Wants to update B, so it waits for the lock on B | Wants to update A, so it waits for the lock on A |
| 3 | Both wait for each other forever. The database detects this and picks one as the victim |
Solution
- A fixed lock order: always update the account with the smaller Id first, whether it is the source or the destination. Now both transactions go to A first. The second one only waits until the first one finishes. No deadlock happens.
- Short transactions: slow work (an HTTP call, heavy calculation) should not be inside the transaction. The shorter a lock is held, the lower the chance of a clash.
- The right index: without an index, the database reads and locks more rows to find a row.
- Retry: deadlocks never go fully to zero. For a deadlock error, retry the whole transaction a few times. The number of this error is 1205 in SQL Server, and the code is 40P01 in PostgreSQL. In EF Core, the execution strategy does this (for example, with the EnableRetryOnFailure option).
- Finding the exact cause: in SQL Server, use the deadlock graph (with Extended Events and the default system_health session). In PostgreSQL, the database log records it. This report shows which two queries and which locks were involved.
Follow-up question: Do you know Isolation Levels? What problem does Read Committed Snapshot (RCSI) solve, and what problem still remains?
Follow-up answer:
- In RCSI, reads use the previous committed version of the data (row versioning). So a read does not wait for a write lock.
- The result: reads and writes do not block each other. Most deadlocks between reads and writes go away.
- But two writes on the same row still block each other. So the deadlock in this question is not solved by RCSI.
- The problem that remains: Write Skew. Two transactions read old data. Each one makes a decision based on it and writes.
- Example: the balance is 100 and two withdrawals of 80 toman arrive at the same time. Each one sees a balance of 100 and accepts.
- First solution: a conditional, atomic update. That means lowering the balance and checking that it is enough, in one statement:
UPDATE Accounts
SET Balance = Balance - @x
WHERE Id = @id AND Balance >= @x;
- If no row was updated, the balance was not enough.
- Second solution: an explicit lock when reading. In SQL Server with UPDLOCK, and in PostgreSQL with FOR UPDATE.
Red flag: Their solution is “we add NOLOCK”. With NOLOCK, the database returns half-done, uncommitted data (dirty read). It may even return duplicate or missing rows. For money transfers, this is dangerous.
31When the user cancels a requestMediumasync/await and the thread poolMiddleware and the pipeline
Question: We have a heavy report. Its query takes about 2 minutes. Users usually do not wait. They close the page or press Refresh several times. You see this endpoint in code review:
app.MapGet("/reports/sales", async (AppDbContext db) =>
await db.Sales.Where(...).GroupBy(...).ToListAsync());
What do you think? Does it have a problem? When the user closes the page, what happens to the query?
Short answer: When the user closes the page, ASP.NET Core notices. But it only sends a “signal”. This signal is called a CancellationToken. If the code has not given this token to the query, nobody tells the database “stop”. So the query runs to the end. Solution: pass the token through all layers.
Hint (if the candidate says there is no problem): The database administrator (DBA) sees in the monitoring tool that at peak hours, dozens of queries for this same report are running at the same time. Most users of these queries have left. The database has become slow for the other users.
What happens? (step by step)
- The user asks for the report. The 2-minute query starts.
- After 10 seconds, the user presses Refresh. The previous connection is closed. A new request starts.
- The ASP.NET Core framework sees that the first connection was closed. So it cancels the token of that request (the RequestAborted token).
- But the query code did not get any token. So the first query runs to the end. Then its result is thrown away.
- Now two heavy queries are running for one user. With a few impatient users, we have dozens of useless queries.
Solution:
- Pass the cancellation token (CancellationToken) from start to end. In a controller or minimal API, it is enough to take a parameter of this type. ASP.NET Core fills it by itself:
app.MapGet("/reports/sales", async (AppDbContext db, CancellationToken ct) =>
await db.Sales.Where(...).GroupBy(...).ToListAsync(ct));
- When the token is canceled, EF Core and ADO.NET send a cancel command to the database. So the query also stops inside the database. Give the same token to HttpClient and the other async methods too.
- Other protections:
- Set a timeout for the query (the CommandTimeout setting). This way, no query runs forever.
- If this same report is already running for this same user, do not run it again.
- For a very heavy report, make it a background job or cache the result (question 18).
- An OperationCanceledException caused by closing the page is normal. It should not be logged as an error. It should not create an alert.
Follow-up question: Where should you ignore request cancellation and finish the work to the end?
Follow-up answer: When we are in the middle of a multi-step job and leaving it half-done is bad. Example: money was taken by the payment gateway, and only saving the result in the database is left. The user has left. But if we cancel here, the money is taken and no order is saved. For these steps, we pass an empty token (CancellationToken.None). Or we hand the work to a queue or an Outbox. A good candidate separates these two: “the user no longer wants the answer” and “the work must be completed”.
Red flag: Never passes a CancellationToken. Or thinks closing the browser stops the database query by itself.
32Changing message structure (Schema Evolution)DesignWeight ×2Schema evolution
Question: Five different services read the OrderCreated event from Kafka. Each service has its own team and its own deploy time. Right now, the message looks like this:
{ "orderId": "o-1", "amount": 250000 }
Because of moving to multiple currencies, the Order team wants to change the amount field to this shape:
{ "orderId": "o-1", "amount": { "value": 250000, "currency": "IRR" } }
If they deploy this change tomorrow, what happens? How do you make this change so that no service breaks?
Short answer: Consumers that still expect a number do not understand the new message and get stuck. A message is a contract. Changing a field’s type, removing a field or renaming a field is a breaking change. The right way: add the new field next to the old field. Remove the old field last of all.
Why does it break? (step by step)
- The Notification service still has the old code. This service treats the amount field as a decimal number.
- The new message arrives. Now amount is an object. So reading the message (deserialize) fails with a JsonException.
- The consumer tries the message again (retry) and fails again. In a partition, order matters. So the next messages also get stuck behind this message. Or the message goes to the Dead Letter, and the SMS never reaches the customer.
- The reverse problem also exists. If a service moves to the new code earlier, it does not understand the old messages. These messages are still in the topic.
- Making five teams deploy at the same moment is practically impossible.
The right way (like Expand and Contract in question 17):
- A new field with a new name is added. The old field is still filled:
{ "orderId": "o-1", "amount": 250000, "money": { "value": 250000, "currency": "IRR" } }
- The consumers move to reading the money field one by one, whenever they want.
- When all have moved and the old messages are no longer needed, the amount field is removed.
Simple rules:
- Adding an optional field is safe.
- Removing, renaming or changing the type of a field is a breaking change. Either do it with the method above, or create a new version of the event (for example, OrderCreated version 2) and publish both versions for some time.
- On the consumer side, use the Tolerant Reader pattern. This means ignore unknown fields and read only the fields you need.
- Use a Schema Registry tool with Avro or Protobuf. With this tool, an incompatible schema is rejected at registration time, not in Production.
Follow-up question: A new service wants to read all events of the past year from the start (replay). What problem comes up?
Follow-up answer: Messages from one year have several different schema versions. So we have two options:
- The new service understands all versions.
- Or a conversion layer (upcaster) converts each old version to the new version.
That is why each message must carry its version number (in a header or with a schema id). Another point: the data retention time in Kafka must be long enough. If the data was not kept for one year, a replay is not possible.
Red flag: Says “we tell all teams and everyone deploys together”.
33Deadlock in Orleans GrainsHardMicrosoft Orleans
This question is for a candidate who has worked with Orleans or another Actor Model. If they have not, only ask about the idea of a “single-threaded actor” with this example.
Question: In Orleans, we have two grains: the account grain (AccountGrain) and the risk grain (RiskGrain). When the account grain processes a withdrawal, it calls the risk grain to check the risk. For this check, the risk grain needs the current balance. So it calls the GetBalance method on the account grain. What do you think of this design? When a withdrawal comes in, what happens step by step? Does it have a problem?
Short answer: Each grain is single-threaded by default. This means it does not start the next request until its current request is finished. Here, the account grain waits for the risk grain, and the risk grain waits for the account grain. Neither moves forward. The best solution: remove this cycle in the design.
Hint (if the candidate says there is no problem): No withdrawal ever finishes. After 30 seconds, all withdrawals fail with a TimeoutException.
What happens? (step by step)
- First, the account grain starts the “withdraw” request.
- In the middle of the work, it waits for the answer of the risk check (with await). Important point: for Orleans, the “withdraw” request is not finished yet. So no other request is given to the account grain. This is called non-reentrant.
- Then the risk grain sends a GetBalance request to the account grain. This request waits in the queue, behind the “withdraw” request.
- Now we have a cycle:
- “Withdraw” waits for risk.
- Risk waits for GetBalance.
- And GetBalance waits for “withdraw” to finish.
- This is a deadlock. After the default timeout (30 seconds), the request fails.
Why is Orleans single-threaded? Because of this, we do not need a lock inside a grain. The state is never broken by two concurrent requests. This is the main benefit of the Actor Model.
Options, from better to worse:
- Remove the cycle in the design: In this option, the account grain sends the current balance with the request to the risk grain. A call back is no longer needed:
var ok = await riskGrain.Check(amount, currentBalance);
- Allow re-entry only for this call chain: In newer versions of Orleans, the AllowCallChainReentrancy method does this. A request that comes back from the same chain is allowed to run.
- Mark read-only methods: For example, mark the GetBalance method with the ReadOnly or AlwaysInterleave attributes. This way, it also runs in the middle of other work.
- Make the whole grain Reentrant: This is the simplest option, but also the most dangerous. The single-thread safety is lost. Between two awaits, another request can change the state. Example: two withdrawals both see a balance of 100 and both are accepted.
General rule: reduce cycles between grains. Do work that needs no answer with one-way messages (the attribute named OneWay) or with streams.
Follow-up question: We have a global grain, for example a counter of all orders today. All orders call it. Adding silos does not make it faster. Why, and what do you do?
Follow-up answer: Each grain is alive on only one silo and works single-threaded. So all requests are processed one after another. More silos do not help a single grain. Options:
- Split the counter into several grains. For example, 16 partial counters (by the hash of the order ID) and one grain that adds them up every few seconds.
- Collect changes in each silo and send them in batches, instead of one call for each order.
- For work without state, use StatelessWorker. This kind of grain has several copies at the same time.
Red flag: Right away suggests making the grain Reentrant and does not know what protection is lost by doing this.
34Following one request across several services (Observability)MediumDistributed tracing
Question: An order goes along this path: the API, then a Kafka message, then a Worker, then an HTTP call to the payment service. A customer calls and says “my order has been waiting for payment for 2 hours”. The logs are in 4 separate services. Each service has several copies (pods). Each one writes thousands of log lines per minute. Right now, finding where this order got stuck takes several hours. How do you design the system so that you can find the path of this one order in a few minutes?
Short answer: With Distributed Tracing. Each request gets an ID called a Trace Id. This ID travels with the request through all services, even through Kafka. With this one ID, we see the whole path and the time of each part on one page. Next to it, we collect structured logs in one central place.
Why is it hard now?
- The logs are scattered. No shared ID links them together.
- We must guess which pod processed this order. Then we must match the logs by their clock times.
The right design:
- Distributed tracing with OpenTelemetry: Each request has a Trace Id. Each part of the work (for example, a query or an HTTP call) is a Span. In .NET, ASP.NET Core and HttpClient themselves send and read the standard traceparent header. This is done with the Activity class.
- Carrying the Trace Id through Kafka (important point): Kafka does not do this automatically. So two things are needed:
- When producing a message, we put traceparent in the message header.
- When consuming a message, we create a new Activity from that header. Or we use a ready-made instrumentation library.
- If we do not do this, the trace breaks at Kafka.
- Structured Logging: A log must be stored as fields, not only as text:
logger.LogInformation("Payment requested for {OrderId} amount {Amount}", order.Id, order.Amount);
With this method, the order ID and the Trace Id are searchable in the central tool (Elasticsearch, Loki or Seq).
- A business ID in the trace: Also put the order ID on the spans. Support does not know the Trace Id, but they know the order number. So from the order number we reach the trace, and from the trace we reach all the logs.
- An alert before the customer complains: Create a metric that shows the number of stuck orders in each status. When an order stays in one status longer than a limit (for example, 10 minutes), send an alert.
Security note: logs must not contain sensitive data, like passwords, card numbers or tokens.
Follow-up question: Storing a trace for every request is very expensive, because the data volume is large. What do you do?
Follow-up answer: We use Sampling. There are two methods:
- Head-based: We decide at the start of the request. For example, only 10% of requests are stored. It is simple. But the trace of a failed request that we need may be in the dropped 90%.
- Tail-based: We decide after the trace finishes (for example, in the OpenTelemetry Collector). All failed and slow traces are kept. Only a small percent of normal traces is kept.
Red flag: Their solution is this: “we search the logs of all services by order number”. And they have not thought about carrying the Trace Id through Kafka.
35Memory leak in .NETMedium to hardMemory and the Garbage CollectorBackground services
Question: On the monitoring dashboard, we see the memory of a service go up slowly and steadily over several days. In the end, the pod restarts with the status OOMKilled. After the restart, the same cycle repeats. A colleague says: “.NET has a GC, so a memory leak is not possible. Let’s just raise the memory limit.” What do you think? How do you find the cause, step by step?
Short answer: The GC only cleans up objects that nothing points to anymore. Sometimes a forgotten reference still points to an object, for example from a static object. That object is never cleaned up. So a memory leak in .NET is fully possible. Raising the limit only delays the OOM.
A real example of a leak:
// PriceFeed is registered as Singleton
public class OrderHandler // Scoped: one instance per request
{
public OrderHandler(PriceFeed feed)
{
feed.PriceChanged += OnPriceChanged; // never unsubscribed
}
private void OnPriceChanged(object? s, PriceEventArgs e) { /* ... */ }
}
Why is this a leak? (step by step)
- Each event keeps a reference to all its subscribers. So the PriceFeed service points to every OrderHandler.
- The PriceFeed service is a Singleton. This means it lives until the end of the app’s life.
- So every OrderHandler created in each request also stays alive forever.
- After millions of requests, we have millions of handlers in memory.
A second real example: one DbContext forever in a BackgroundService
// Wrong: one scope and one DbContext for the whole life of the service
protected override async Task ExecuteAsync(CancellationToken ct)
{
using var scope = _scopeFactory.CreateScope();
var db = scope.ServiceProvider.GetRequiredService<AppDbContext>();
while (!ct.IsCancellationRequested)
{
var items = await db.Outbox.Where(x => !x.Sent).Take(100).ToListAsync(ct);
foreach (var item in items) item.Sent = true;
await db.SaveChangesAsync(ct);
await Task.Delay(1000, ct);
}
}
Why is this a leak? (step by step)
- A BackgroundService starts once and lives until the end of the app’s life. So that scope and that DbContext also stay alive until the end.
- Each DbContext has a Change Tracker. Every entity read or saved with it stays in the Change Tracker. It does not leave even after SaveChanges.
- The loop reads 100 new rows every second. So after one day, we have millions of entities in the Change Tracker.
- The GC cannot clean them up, because the DbContext still points to all of them.
- Memory goes up slowly until the pod restarts with OOMKilled.
- Speed also drops. Before saving, the SaveChanges method runs DetectChanges. This means it checks all entities in the Change Tracker one by one. The more there are, the slower each loop round gets.
// Right: a new scope (and a new DbContext) for each batch
protected override async Task ExecuteAsync(CancellationToken ct)
{
while (!ct.IsCancellationRequested)
{
using (var scope = _scopeFactory.CreateScope())
{
var db = scope.ServiceProvider.GetRequiredService<AppDbContext>();
var items = await db.Outbox.Where(x => !x.Sent).Take(100).ToListAsync(ct);
foreach (var item in items) item.Sent = true;
await db.SaveChangesAsync(ct);
} // scope and DbContext are disposed here
await Task.Delay(1000, ct);
}
}
- With this, each loop round has a fresh, empty DbContext. At the end of the round, the DbContext is disposed and the GC cleans up all the entities.
- For queries that only read and change nothing, use the AsNoTracking method. Then the entity never enters the Change Tracker.
- If for some reason the same DbContext must stay, call the ChangeTracker.Clear method after each batch. But the first option (a new scope) is cleaner.
- Instead of a scope, you can also use IDbContextFactory and create a fresh DbContext in each round.
Other common causes:
-
A DbContext that lives forever (for example, in a BackgroundService or inside a Singleton). All entities pile up in its Change Tracker.
-
A static cache or Dictionary that only gets new items and is never cleaned. Or an IMemoryCache with no expiration time and no size limit (SizeLimit).
-
Timers that are not disposed.
-
A disposable Transient service taken from the root container, not from a scope. In this case, the container keeps it until the end of the app’s life, so it can dispose it at the end.
-
Unmanaged resources, like streams and connections, that are not disposed.
How to find the cause (step by step):
- With the dotnet-counters tool, see which one goes up: the GC Heap memory (.NET objects) or only the total process memory. If the GC Heap is steady but total memory goes up, we probably have a native leak.
- Take two memory snapshots some time apart (for example, one hour). Use the dotnet-gcdump or dotnet-dump tool for this.
- Compare the two snapshots (in Visual Studio, PerfView or dotnet-dump). The question is: which object type keeps growing in count? For example, you see 2 million OrderHandlers in memory.
- For one of those objects, find the GC Root (with the gcroot command). This command shows the chain of references. That means it tells you what keeps the object alive. Here, the answer is the event inside PriceFeed.
- Fix the code and measure again. In this example, unsubscribe from the event in the Dispose method. Or change the design so that events on a Singleton are not used.
Follow-up question: Suppose we have no leak. But after a load peak, the service memory stays high and does not come down. Is this normal? How do you know if it is a leak or not?
Follow-up answer: Yes, it may be normal. The reason:
- The GC does not always give freed memory back to the operating system right away. It keeps it for later use.
- Server GC mode also uses more memory by default.
The right measure: look at memory after each full GC (Gen2). If it returns to a steady level each time, there is no leak. If this level goes higher each time, we have a leak. If memory use matters more than speed, you can tune the GC. For example, with the GCConserveMemory setting, or with Workstation GC in small pods.
Red flag: Says “in .NET we have no memory leaks because we have a GC”. Or their only solution is “we raise the memory limit”.
36Aggregates and keeping domain rulesMedium to hardDomain-Driven Design
Question: In the order system, we have a business rule: the total amount of each order must not be more than the customer’s credit limit. The Order class checks this rule in the AddLine method. Now a new endpoint was written to “change the quantity of one order line”:
public async Task ChangeQuantity(int lineId, int newQty)
{
var line = await _db.OrderLines.FindAsync(lineId);
line.Quantity = newQty;
await _db.SaveChangesAsync();
}
You see this code in code review. The tests are green too. What do you think? Does it have a problem? If you wrote it yourself, how would you write it?
Short answer: This code goes around the Aggregate root. The credit limit rule is in Order, but the code changes OrderLine directly. So the rule is never checked. In DDD, only the Aggregate Root is changed from outside. The root keeps all the rules inside its boundary. The whole Aggregate is loaded together and saved together.
Hint (if the candidate says there is no problem): A few weeks later, the finance team reports that some orders have an amount higher than the customer’s credit limit. There is no error in the log of the place-order endpoint. All these orders were changed from the “edit quantity” page after they were placed. Now, why?
Why is this code dangerous? (step by step)
- The credit limit rule is about the whole order, not one line. To check it, we must see all lines.
- The OrderLine class does not know about the other lines. So it cannot check this rule.
- The code takes one line directly from the lines DbSet and changes its quantity. The AddLine method is not called.
- So the rule is bypassed. Any new developer can also repeat the same thing somewhere else.
- The tests are green, because the tests only tested the AddLine path.
Solution: all changes go through the root
public class Order
{
private readonly List<OrderLine> _lines = new();
public IReadOnlyList<OrderLine> Lines => _lines;
public void ChangeQuantity(int lineId, int qty, Money creditLimit)
{
var line = _lines.Single(l => l.Id == lineId);
var newTotal = Total() - line.Total() + line.Price * qty;
if (newTotal > creditLimit)
throw new DomainException("Credit limit exceeded");
line.SetQuantity(qty); // internal setter
}
}
var order = await _db.Orders.Include(o => o.Lines)
.SingleAsync(o => o.Id == orderId);
order.ChangeQuantity(lineId, newQty, customer.CreditLimit);
await _db.SaveChangesAsync();
Some important rules for an Aggregate:
- Only the root is reachable from outside. For OrderLine, we make no separate repository and no public DbSet. The list of lines is also exposed as read-only.
- The consistency boundary is the transaction boundary. Everything inside the Aggregate must stay correct together in one transaction. One transaction changes only one Aggregate.
- Keep it small. A large Aggregate means a heavy load and many clashes between users (locks or concurrency errors). Put inside it only the things that a real rule ties together.
- Point to other Aggregates by Id. The order keeps only the customer ID, not the customer object itself. This way, the boundaries do not get mixed.
- Think about concurrency too. If two users edit one order at the same time, with a version column (RowVersion) on the root, EF Core rejects the second change.
Follow-up question: When an order is confirmed, the warehouse stock must also go down. Order and Inventory are two separate Aggregates. Do you change both in one transaction or not?
Follow-up answer:
- The basic rule: one transaction, one Aggregate. So we usually use Eventual Consistency.
- When confirmed, the order records a Domain Event named “order confirmed”.
- This event is saved with the Outbox pattern in the same transaction as the order. So it is not lost.
- A separate handler takes the event and lowers the stock. If the stock is not enough, it sends a compensating event, and the order is canceled or put on hold.
- When is a shared transaction acceptable? When both are in one service and one database, the load is low, and the business does not accept even one moment of inconsistency. But it must be a conscious choice. It adds more clashes and locks.
- A good question for the business: “If the stock goes down a few seconds later, what is the harm?” Often the answer is “none”.
Red flag: Makes a repository for each table and says “we check the rules in the service layer”. Or makes a very large Aggregate that holds the customer, all orders and the warehouse together.
37Anemic or Rich domain modelMediumDomain-Driven Design
Question: In the order project, the Order class and the OrderService are written like this:
public class Order
{
public int Id { get; set; }
public string Status { get; set; }
public decimal Total { get; set; }
public DateTime? ConfirmedAt { get; set; }
}
public class OrderService
{
public async Task Confirm(int id)
{
var order = await _db.Orders.FindAsync(id);
if (order.Status != "Draft") throw new Exception("Invalid");
if (order.Total <= 0) throw new Exception("Empty order");
order.Status = "Confirmed";
order.ConfirmedAt = DateTime.UtcNow;
await _db.SaveChangesAsync();
}
// Cancel, Ship, Refund ... the same style, 900 lines
}
This structure repeats across the whole project. What do you think of this design? If it were your project, would you change anything?
Short answer: This is an Anemic Domain Model. The Order class has only data, and all the rules are in the service. Because all setters are public, any code can change the state without checking the rules. In a rich model, the rules are inside the entity itself. For example, the Confirm and Cancel methods. The setters are private. Important values are Value Objects too, like Money. But for a simple CRUD, an anemic model is fully acceptable.
Hint (if the candidate says there is no problem): In Production, we have some orders with the status “confirmed” but no confirmation date. A nightly job also set some canceled orders back to “shipped”. In one month, three bugs of this same kind were reported. Each time in a new place in the code. Now, why?
Why does this design cause problems in a complex domain? (step by step)
- The rule “only a draft order can be confirmed” is only in one service method.
- All setters are public. So a job, a handler or test code can change the state directly.
- Anyone who does this has bypassed the rule. For example, they change the state but forget the confirmation date.
- The rules are repeated in several services and slowly become different from each other.
- The status is a text string. A typo like “confirmed” with a lowercase letter still compiles.
- The service grows to 900 lines. To understand order behavior, you must read all of it.
Solution: move the behavior into the entity
public class Order
{
public int Id { get; private set; }
public OrderStatus Status { get; private set; } = OrderStatus.Draft;
public Money Total { get; private set; }
public DateTime? ConfirmedAt { get; private set; }
public void Confirm(DateTime now)
{
if (Status != OrderStatus.Draft)
throw new DomainException("Only draft orders can be confirmed");
if (Total.Amount <= 0)
throw new DomainException("Order is empty");
Status = OrderStatus.Confirmed;
ConfirmedAt = now;
}
}
public record Money(decimal Amount, string Currency)
{
public static Money operator +(Money a, Money b) =>
a.Currency == b.Currency ? a with { Amount = a.Amount + b.Amount }
: throw new DomainException("Currency mismatch");
}
- Now the service only coordinates: it loads the order, calls the Confirm method and saves.
- The rule is in only one place. No code can bypass it.
- Testing the rule is simple. You create an Order object and call Confirm. No database and no mock are needed.
- With EF Core, private setters and private fields are no problem. For a Value Object, we use an Owned Type or a Complex Property.
When is an anemic model good?
- When the app is only CRUD and has no real rules. For example, an admin panel for a cities table.
- When the class is only for moving data, like a DTO or a read model.
- Here, building a rich model is only extra complexity (YAGNI).
Follow-up question: What is the difference between a Value Object and an Entity? Which one do you choose for a customer address?
Follow-up answer:
- An Entity has an identity. Two orders with the same values are two separate orders if their IDs differ. It also changes over time but stays the same order.
- A Value Object has no identity. It is known by its value. Two amounts of “100 toman” are fully equal.
- A Value Object is immutable. To change it, we create a new object.
- A Value Object checks itself when it is created. For example, an invalid Email is never created. So the rest of the code does not need to check it again.
- A customer address is usually a Value Object. If the address changes, a new address replaces it. But if the business manages addresses separately (for example, an address book with IDs), then it is an Entity.
- In C#, the record type is a very good fit for a Value Object. This is because it builds value-based equality by itself.
Red flag: Says “an entity must only hold data and logic always belongs in the service”. Or the opposite: insists on building a rich model and Value Objects even for a simple CRUD.
38Bounded Context and shared languageDesignDomain-Driven Design
Question: In the company, we have three teams: sales, support and finance. All three use one shared Customer class and one shared customers table. This class is in a Shared project, and all services reference it:
public class Customer
{
public int Id { get; set; }
public string Name { get; set; }
public string SalesStage { get; set; } // Sales
public decimal? DiscountPercent { get; set; } // Sales
public int OpenTickets { get; set; } // Support
public string SlaLevel { get; set; } // Support
public string TaxCode { get; set; } // Billing
public decimal CreditLimit { get; set; } // Billing
// ... 60 more fields
}
Each team adds fields for its own needs. What do you think of this design? What happens as the company and the teams grow?
Short answer: This model wants everything for everyone. The word “customer” has a different meaning in each team. So one shared class ties all teams together. The DDD solution is to split the system into Bounded Contexts. Each context has its own model of the customer, with its own language (Ubiquitous Language). Contexts connect to each other only through the customer ID and through events. A Context Map shows how they relate.
Hint (if the candidate says there is no problem): The finance team changed the type of one field. The support service stopped working after the deploy. Now, for every small change, three teams must hold a meeting and deploy together. The phrase “active customer” in sales means “bought something in the last 3 months”, but in finance it means “has an open debt”. The reports of the two teams do not match. Now, why?
Why does this design slowly cause problems? (step by step)
- Each team needs only a small part of the fields. But all of them depend on the whole class.
- When one team changes a field, the code of the other teams also breaks. So teams cannot deploy separately.
- One word has a different meaning in each team. But in the code we have only one class and one field. So the meanings get mixed and cause logic bugs.
- The class keeps getting bigger. Nobody dares to remove a field, because nobody knows who uses it.
- A shared table means shared locks and shared slowness. One team’s heavy query slows down another team.
Solution: each context has its own model
// Sales context
public class Customer { public Guid Id; public string Name; public SalesStage Stage; }
// Support context
public class Requester { public Guid CustomerId; public SlaLevel Sla; }
// Billing context
public class Account { public Guid CustomerId; public TaxCode TaxCode; public Money CreditLimit; }
- Each context has its own model and database (or at least its own schema). Only the fields it needs. Even the class name can differ. In support it is called a “requester”, and in finance an “account”.
- Connection by ID. All contexts share one customer ID. But they do not read each other’s data directly.
- Connection by events. When sales creates a new customer, it publishes a “customer registered” event (for example, with Kafka). Finance and support receive it and build their own models.
- Anti-Corruption Layer. When a context must connect to an old system or another model, we put a translation layer in between. This layer converts the outside model into our own model. So a change in the outside model changes only this layer, not all our code.
- Context Map. We write down the relationship between every two contexts. Who is the supplier and who is the consumer? What contract is between them? For example, the event contract has versions and does not break without coordination.
- Make the change step by step. Everything does not need to change at once. First split one context (for example, finance), fill it with events, and then the others.
Also state the cost of this solution:
- Data is duplicated, and we have Eventual Consistency. It takes a few seconds for a new customer to show up in finance.
- Company-wide reports need a separate model, for example a data warehouse.
- For a small team with a simple product, this split may be too much.
Follow-up question: How do you find the boundaries of contexts?
Follow-up answer:
- From differences in language. When one word has two meanings in two teams, or there are two words for one thing, you have probably reached a boundary.
- With Event Storming. All business and technical people come together and place the domain events on a wall. For example, “order placed” and “invoice issued”. Where events form groups and the language changes, there is a context boundary.
- From the team structure. By Conway’s law, a system ends up like the communication structure of the teams. It is better for each context to belong to one team.
- From rules and data. Things that change together under one rule stay in one context.
- From the speed of change. Separate the part that changes every week from the stable part.
Red flag: Says “one shared model is better, because there is no duplication (DRY)”. Or the opposite: right away suggests one microservice for each table, without thinking about the business language and boundaries.
39The Mediator pattern and the MediatR libraryMedium to hardCQRS and MediatorDesign patterns
Question: In an ASP.NET Core API, the constructor of the orders controller looks like this:
public OrdersController(
IOrderService orders, ICustomerService customers,
IInventoryService inventory, IPaymentService payments,
IDiscountService discounts, INotificationService notifications,
IAuditService audit, IMapper mapper, ILogger<OrdersController> logger)
A teammate suggests moving all controllers to MediatR. This means each action is a Command or a Query with its own Handler. The controller takes only an IMediator. What do you think of this suggestion? What does the Mediator pattern give us, and what problem does it solve? When is it worth it, and when is it not?
Short answer: The Mediator pattern separates the sender of a request from the one that handles it. The controller only sends a message and does not know who runs it. Each use case has a small Handler. Shared work (validation, logging, transactions) is written in one place in the pipeline. But it also has a cost: an extra layer, many classes, and code that is harder to find. More important, 9 dependencies in a controller are a sign of a design problem. MediatR does not solve this problem, it only hides it.
Hint (if the candidate says there is no problem): We did the same thing in another project. After a few months, new developers said it was hard to find the code that runs. With “go to definition” they only reached the Send method. Some Handlers also called other Handlers with Send, and the chain had become long. Now, why?
Where is the main problem? (step by step)
- A controller with 9 dependencies means this class knows about many jobs. This goes against the single responsibility principle (SRP).
- Each action needs only two or three services. But to build the controller, all 9 services are created.
- Testing the controller is also hard. For each test, you must mock 9 dependencies.
- With MediatR, the controller takes only an IMediator. Each Handler has only its own dependencies. So the classes become small.
- But if the business logic is not split well, the same mess moves into the Handlers. So MediatR does not replace good design.
Benefits of Mediator:
- Sender separated from receiver: The controller only sends the message. The controller becomes thin.
- One Handler for each use case: It is similar to CQRS. Classes are small, have one clear job, and are separate from each other.
- Pipeline behaviors: Shared work like validation, logging, timing and transactions is written once and runs for all requests.
- Easier testing: You test each Handler separately, with few dependencies.
Costs and risks:
- An extra layer (indirection): With “go to definition” in the IDE, you do not reach the Handler. You must search for the Handler class.
- Many classes: For each simple job, you need a Command, a Handler and sometimes a Validator.
- Hiding the problem: Dependencies are hidden behind IMediator. You can no longer see from the constructor what the class depends on.
- Send chains: If Handlers call each other with Send, the code flow gets lost.
- Not needed for simple CRUD: In a small service, a few simple services are enough (YAGNI).
- License: The MediatR library has a commercial license since version 13 (in 2025). It is paid for larger companies, and only small companies get a free Community edition. You must check this before choosing it.
When is it worth it?
- When the project is large and has many use cases.
- When we have a lot of shared work that must run for all requests.
- When the team works with CQRS and Vertical Slice.
When is it not worth it?
- When the service is small or is mostly simple CRUD.
- When the only goal is fewer constructor parameters. In this case, it is better to split the controller into a few smaller controllers, or to split the services more correctly.
Simpler options: For this controller’s problem, these also work:
- Split the controller by job (for example, a separate controller for order payment).
- Group services that are always used together behind one business service.
- In Minimal API, each endpoint takes only its own dependencies.
- If we want the Mediator pattern but not the library, a simple internal dispatcher can be written in a few dozen lines.
Follow-up question: Suppose we chose MediatR. We want all Commands to be validated before they run. How do you do this without repeating code in every Handler?
Follow-up answer: With a Pipeline Behavior. This class runs around all Handlers. It gets all Validators for that request from DI and runs them. If there is an error, the Handler does not run at all:
public sealed class ValidationBehavior<TRequest, TResponse>(
IEnumerable<IValidator<TRequest>> validators)
: IPipelineBehavior<TRequest, TResponse> where TRequest : notnull
{
public async Task<TResponse> Handle(TRequest request,
RequestHandlerDelegate<TResponse> next, CancellationToken ct)
{
foreach (var validator in validators)
await validator.ValidateAndThrowAsync(request, ct);
return await next(ct);
}
}
Registering it in DI:
builder.Services.AddMediatR(cfg =>
{
cfg.RegisterServicesFromAssemblyContaining<Program>();
cfg.AddOpenBehavior(typeof(ValidationBehavior<,>));
});
builder.Services.AddValidatorsFromAssemblyContaining<Program>();
- In the same way, logging and transactions are added as separate Behaviors.
- The order of registering Behaviors matters. The first registered Behavior is the outermost layer.
Red flag: Says “MediatR is always needed and makes the architecture clean” and sees no cost in it. Or does not see that 9 dependencies in the controller are themselves a sign of a design problem.
40The Strategy pattern instead of a big switchMediumDesign patterns
Question: We have this code in the payment service:
public class PaymentService(BankAClient bankA, BankBClient bankB, WalletClient wallet)
{
public async Task<PaymentResult> PayAsync(Order order, string provider, CancellationToken ct)
{
switch (provider)
{
case "bank-a":
var token = await bankA.GetTokenAsync(ct);
return await bankA.PayAsync(token, order.Amount, ct);
case "bank-b":
return await bankB.PayAsync(order.Id, order.Amount * 10, ct);
case "wallet":
return await wallet.ChargeAsync(order.CustomerId, order.Amount, ct);
default:
throw new NotSupportedException(provider);
}
}
}
The product team said a few more payment gateways will be added in the coming months. You see this code in code review. What do you think? Does it have a problem? If you want to change it, what design do you suggest?
Short answer: This code works today. But with every new gateway, we must open this same class and change it. This goes against the Open/Closed principle. A better way is the Strategy pattern: one shared interface for all gateways, one class for each gateway, and choosing the gateway with DI. But if we have only 2 or 3 fixed cases, this switch is enough (YAGNI).
Hint (if the candidate says there is no problem): In one year, this class grew from 30 lines to 600 lines and got 15 dependencies. Once, a small change for bank B broke wallet payments. Two developers who were adding two new gateways at the same time had conflicts on this same file every time. To test one gateway, we also had to mock all the clients. Now, why?
Why does this code cause problems over time? (step by step)
- Each new gateway means a new case and a new dependency in the constructor. The class keeps getting bigger.
- All gateways are in one file. So a change in one may break another.
- Each change means testing the whole class again, not only the new gateway.
- To test one gateway, you must create or mock all the clients.
- Several people working on different gateways at the same time get conflicts on one file.
- Usually the same switch is also repeated in other places (refund, status check). So adding a gateway means changing several places in the code.
Solution: Strategy with DI
First, a shared interface:
public interface IPaymentProvider
{
string Key { get; }
Task<PaymentResult> PayAsync(Order order, CancellationToken ct);
}
Then, one class for each gateway:
public sealed class BankAProvider(BankAClient client) : IPaymentProvider
{
public string Key => "bank-a";
public async Task<PaymentResult> PayAsync(Order order, CancellationToken ct)
{
var token = await client.GetTokenAsync(ct);
return await client.PayAsync(token, order.Amount, ct);
}
}
And registering all of them in DI:
builder.Services.AddScoped<IPaymentProvider, BankAProvider>();
builder.Services.AddScoped<IPaymentProvider, BankBProvider>();
builder.Services.AddScoped<IPaymentProvider, WalletProvider>();
- Now, for a new gateway, you only write a new class and add one registration line. The old code is not touched.
- Each gateway has only its own dependency and is tested separately.
- The differences of each gateway (for example, converting toman to rial for bank B) stay inside its own class.
When is the switch better?
- When we have only 2 or 3 cases and they are not going to grow.
- When each case is only a line or two.
- In this case, building an interface and several classes is only extra complexity. The YAGNI principle says: do not build it until you need it.
Follow-up question: The user chooses the gateway at run time. How do you find the right gateway without writing a switch again?
Follow-up answer: Three common ways:
- Dictionary: Get all IPaymentProviders from DI and put them in a dictionary by Key:
public sealed class PaymentService(IEnumerable<IPaymentProvider> providers)
{
private readonly Dictionary<string, IPaymentProvider> _map =
providers.ToDictionary(p => p.Key);
public Task<PaymentResult> PayAsync(Order order, string key, CancellationToken ct) =>
_map.TryGetValue(key, out var provider)
? provider.PayAsync(order, ct)
: throw new NotSupportedException(key);
}
- Keyed Services: Since .NET 8, DI itself supports registering with a key:
builder.Services.AddKeyedScoped<IPaymentProvider, BankAProvider>("bank-a");
var provider = serviceProvider.GetRequiredKeyedService<IPaymentProvider>(key);
- A Factory class: A small class that takes the key and returns the gateway. Inside, it uses one of the two ways above. The rest of the code works only with this Factory.
- Note: when the key comes from user input, always handle the “unknown key” case properly and return a clear error.
Red flag: Sees no problem in the switch growing. Or the opposite: builds an interface, a Factory and several layers even for 2 simple cases, and cannot say when a switch is enough.
41The Decorator pattern and a Repository on EF CoreMedium to hardDesign patternsEntity Framework Core
Question: We have an IProductRepository implemented with EF Core. Now we want to add caching (with Redis) and logging around it. A teammate did it like this:
public async Task<Product?> GetByIdAsync(int id, CancellationToken ct)
{
_logger.LogInformation("Getting product {Id}", id);
var cached = await _cache.GetStringAsync($"product:{id}", ct);
if (cached is not null)
return JsonSerializer.Deserialize<Product>(cached);
var product = await _db.Products.FindAsync([id], ct);
await _cache.SetStringAsync($"product:{id}", JsonSerializer.Serialize(product), ct);
return product;
}
The same pattern is repeated in all methods of this repository. You see this code in code review. What do you think? If it were you, how would you add caching and logging? And is this repository on EF Core needed at all?
Short answer: This code mixes three jobs in one class: data access, caching and logging. A better way is the Decorator pattern. A new class implements the same interface, holds the original repository inside, and only adds caching around it. The logging class is separate too. These classes are stacked on each other in DI, and the original code is not touched. About the repository: the DbContext is itself a Unit of Work and a Repository. So a generic repository on top of it is often just an extra layer.
Hint (if the candidate says there is no problem): Once Redis went down, and all product read methods failed, even though the database was healthy. When we wanted to change the cache key, we had to change 12 methods one by one, and one was missed. The repository unit tests also did not run without Redis. In one place, a product that did not exist was saved in the cache as an empty value. Now, why?
Where is the problem in this code? (step by step)
- Each method has three responsibilities. This goes against the single responsibility principle (SRP).
- The cache logic is repeated in all methods. So any change in caching (key, expiration time, error handling) means changing all methods.
- If we want to turn the cache off somewhere (for example, in an internal job), there is no simple way.
- Testing the data access logic without the cache is not possible.
- It also has a small bug: when a product is not found, an empty value goes into the cache. The expiration time is not set either.
Solution: Decorator
The cache class implements the same interface and takes the original repository:
public sealed class CachedProductRepository(
IProductRepository inner, HybridCache cache) : IProductRepository
{
public async Task<Product?> GetByIdAsync(int id, CancellationToken ct) =>
await cache.GetOrCreateAsync($"product:{id}",
async token => await inner.GetByIdAsync(id, token),
cancellationToken: ct);
// Other methods: read from the cache, or just call inner and clear the cache
}
- Now the original repository works only with EF Core. The cache class knows only about caching.
- The logging class is a separate Decorator in the same way.
- The HybridCache class came in .NET 9 and later. It also prevents a rush of concurrent requests for one key (cache stampede).
Registering in DI:
With the Scrutor library, it is simple:
builder.Services.AddScoped<IProductRepository, EfProductRepository>();
builder.Services.Decorate<IProductRepository, CachedProductRepository>();
builder.Services.Decorate<IProductRepository, LoggingProductRepository>();
Without a library, you can also register it by hand:
builder.Services.AddScoped<EfProductRepository>();
builder.Services.AddScoped<IProductRepository>(sp =>
new LoggingProductRepository(
new CachedProductRepository(
sp.GetRequiredService<EfProductRepository>(),
sp.GetRequiredService<HybridCache>()),
sp.GetRequiredService<ILogger<LoggingProductRepository>>()));
- In Scrutor, the last Decorate is the outermost layer. So the run order is: logging, then cache, then EF Core.
Is a Repository on EF Core needed?
- The DbContext class is itself a Unit of Work. It collects changes and saves them all at once with SaveChanges.
- Each DbSet is itself similar to a Repository.
- So a generic repository (with Add, Update, GetAll methods for all tables) is often just an extra layer. Sometimes it also hides EF Core features like Include or projection.
- But a repository made for one aggregate (like this product repository) is useful in some places:
- When we want repeated queries to be in one place.
- When, in DDD, access must go only through the aggregate root.
- When we want one point for a Decorator (like this cache).
- In short: a generic repository on EF Core, no. A specific repository with a clear purpose, yes.
Follow-up question: Does the order of Decorators matter? If logging is outside the cache, or inside it, what is the difference?
Follow-up answer: Yes, it matters. Each Decorator only sees the requests that reach its layer.
- Logging outside, cache inside: All requests are logged, whether the answer comes from the cache or the database. So we measure the real time the user sees.
- Logging inside, cache outside: When the answer is in the cache, the request never reaches logging. So only requests that went to the database are logged. This is good for seeing the database load.
- Another example: If we have a Decorator for retry, it must be inside the cache. If it is outside, each retry goes to the cache again, which is useless.
- General rule: First decide what each layer must see, then set the order.
Red flag: Copies the cache into every method and sees no problem in it. Or builds a generic repository and a separate Unit of Work on EF Core, and does not know that the DbContext already does this.