Levelwise
English
Back to roadmap
Software Architect

Interview questions

Practice exam
1What does a software architect decide?EasyArchitecture stylesArchitecture Decision Records

Question: A team of 8 people works on an online shop. All the developers are senior. Now a new person joins with the title “software architect”. One developer asks: “We already write good code. What exactly does an architect do that we do not?” You are that architect. What do you answer?

Short answer: An architect works on decisions that are expensive to change later and affect the whole system. For example: the overall style (monolith or separate services), module boundaries, the type of database, and how the parts talk to each other. A senior developer mostly decides design inside one part: classes, functions, patterns. The architect weighs these decisions as trade-offs, writes them down, and explains them to everyone.

Architecture vs design (with an example):

  • A design decision: write the discount calculation with the Strategy pattern or with a few if statements. If it is wrong, fixing it takes about a day.
  • An architecture decision: keep orders and payments in one application, or split them into two services with two databases. If it is wrong, fixing it may take months.
  • The line between the two is not always sharp. A simple test: the cost of change and how many parts are affected.

What does an architect do? (step by step)

  1. Makes the quality needs clear. They ask: how many users? How fast? How available? How secure? What budget?
  2. Puts a few options side by side. For each one, writes down the benefits and the costs. No option has only benefits.
  3. Makes the decision with the team, not alone. The team must understand it and build it.
  4. Records the reason (for example, with an Architecture Decision Record), so nobody asks “why did we pick this?” six months later.
  5. Talks to non-technical people too. For example, tells the product manager: “If you want this feature, the server cost doubles.”
  6. After the decision, checks that the code really follows that path.

What an architect is not:

  • Not the person who makes every small decision. That makes the team slow and unmotivated.
  • Not the person who only draws diagrams and never reads code. An architect who does not know the code makes unrealistic decisions.
  • Not the person who always picks the newest technology. The simplest option that meets the needs is usually better.

Follow-up question: How do you know a decision is “architectural” and needs more time, and when should the team just decide quickly by itself?

Follow-up answer: A few simple questions help:

  1. If it is wrong, how much does it cost to undo? Days or months?
  2. How many teams or parts of the system does it touch?
  3. Does it affect an important quality need (speed, security, availability, cost)?
  4. If the answers are small, the team decides alone. If they are big, we compare options and record the decision.
  5. For reversible decisions, move fast. For hard-to-reverse decisions, go slow and be careful.

Red flag: Thinks an architect means “the person who picks the technology while everyone else just builds it”, or cannot name any cost of their own decisions.

Edit on GitHub

2Starting the architecture from a feature listEasyArchitecture styles

Question: A product manager defines a new food ordering system. They only give this list:

  • Users can order food from restaurants.
  • Users can pay online.
  • Users can track the status of their order.

You are asked to start the architecture for this system. What is the first thing you do?

Short answer: Before drawing any diagram, I ask about the quality attributes (also called non-functional requirements). The feature list says what the system does. Quality attributes say how well it must do it. Architecture is shaped mostly by quality attributes, not by the feature list. These attributes conflict with each other, so we must learn which ones matter most.

Hint (if the candidate says there is no problem): Imagine two teams build the same three features. The first team builds it for one restaurant with 100 orders a day. The second builds it for a big city with 50,000 orders an hour at lunch peak. Is the architecture the same?

Why is a feature list not enough? (step by step)

  1. You can build these three features with one simple application and one database.
  2. You can also build them with dozens of services, message queues and several databases.
  3. Both are correct for the features. They differ in load, speed, availability and cost.
  4. So without these numbers, any architecture we draw is only a guess.

What do I ask?

Quality attribute Example question
Load How many orders at peak hour? How much growth next year?
Latency How many milliseconds may the menu page take?
Availability If the system is down for 1 hour, what do we lose?
Security Do we store card data ourselves, or does the payment provider?
Cost What is the monthly server budget?
Team How many people? Which technologies do they know?
Time When must the first version go live?

Why do these attributes conflict?

  • Higher availability means more servers and more regions. So more cost and a more complex system.
  • More security (for example, extra checks) usually costs a little speed.
  • A short time to launch means a simpler path. So we may need to rewrite a part later for high load.
  • This is why I ask the product manager to rank them. Everything cannot be “very important”.

One note: A quality attribute must be measurable. “The system must be fast” does not help. “95% of menu requests under 300 ms” can be tested.

Follow-up question: The product manager says “Everything is important: it must be very fast, always available, and cheap.” What do you do?

Follow-up answer:

  1. For each attribute, I show its cost with numbers. For example, “more availability means twice the servers”.
  2. I ask which one we give up when they conflict.
  3. I ask per feature, not for the whole system. Maybe payment must always work, but the reviews page does not.
  4. I write down the result, so later it is clear why we chose this.

Red flag: Without asking anything about load, speed or the team, jumps straight to microservices, Kafka and Kubernetes.

Edit on GitHub

3A controller that skips the service layerEasyClean Architecture

Question: An online shop project has three layers: the API layer (controllers), the service layer that holds the business rules, and the data access layer. In code review you see this new handler. The author says “I went straight to the database to make it faster”:

@app.post("/orders/{order_id}/cancel")
def cancel_order(order_id: int):
    db.execute(
        "UPDATE orders SET status = 'cancelled' WHERE id = %s",
        (order_id,),
    )
    return {"ok": True}

The service layer already has a cancel method for orders. What do you think?

Short answer: This handler skips the business rules. The service layer is where rules are written once: for example, “a shipped order cannot be cancelled” or “after cancelling, return the stock and refund the money”. By going straight to the database, these rules do not run and the data becomes inconsistent. “Faster” is usually not true either: one extra function call costs almost nothing.

Hint (if the candidate says there is no problem): Support reports that some orders are “cancelled”, but their package was already shipped. Stock was also not returned for these orders. But when we cancel from the admin page, everything is correct.

Why does this cause problems? (step by step)

  1. The cancel method in the service layer does three things: checks the status, returns the stock, and sends a refund request.
  2. This handler only changes one column. So those three things do not happen.
  3. Now we have two ways to cancel an order, and they behave differently.
  4. If tomorrow a new rule is added (for example, “cancelling after 24 hours has a fee”), we must remember to change both places. Usually one is forgotten.

The dependency rule:

  • Each layer depends only on the layer below it. The API uses the service; the service uses the data layer.
  • Business rules live in one place. To understand or test them, we look only there.
  • If the top layer goes straight to the database, the table structure leaks into the API. Changing a table breaks the controllers too.

Is it always forbidden? No.

  • For simple reads (for example, a product list for display, with no rules), some teams have a direct read path on purpose. This is a simple form of the idea of separating reads and writes (CQRS).
  • The key condition: it must be a team decision that is written down, not one person’s choice in one PR.
  • For writes (status changes, money, stock) we always go through the service layer.
  • Consistency matters. If each person picks a different way each time, the code is not predictable.

The fix: The handler should just call the cancel method of the service layer. If speed is really a problem, measure first. Slowness usually comes from queries or the network, not from one extra layer.

Follow-up question: How do you stop this from happening again, without checking every PR yourself?

Follow-up answer:

  1. Check the rule with a tool. “Architecture test” tools or linters can say that the API module may not import the database module. This test runs in CI.
  2. Arrange the project so direct access is hard (for example, the data layer is only exposed to the services).
  3. Write the rule and its exceptions (like the simple read path) in an ADR.

Red flag: Says “it works, so there is no problem” and does not ask whether a rule in the service layer was skipped.

Edit on GitHub

4The debate that comes back every few monthsEasyArchitecture Decision Records

Question: A team of 12 people has worked on a system for three years. The system uses PostgreSQL and Kafka. Every few months, someone new or someone old asks: “Why PostgreSQL? MongoDB is easier” or “Why Kafka? RabbitMQ would be simpler”. A one-hour meeting happens. Nobody remembers the original reason exactly. The person who made the decision has left the company. You are the architect of this team. What do you do?

Short answer: The problem is not the decision itself. The problem is that the reason for the decision was never written down. The common fix is an Architecture Decision Record (ADR): a short file for each important decision, in the same repository as the code. It says what the situation was, what we decided, which options we rejected, and what the consequences are.

Why does this debate keep coming back? (step by step)

  1. The decision was made three years ago. The situation at that time (load, team, needs) lived only in a few people’s heads.
  2. Those people left or forgot.
  3. A new person sees only the result, not the reason. So it is natural for them to ask.
  4. Without a written record, every debate starts from zero. Team time is wasted, and decisions change based on the loudest voice.

What does an ADR contain?

# ADR-007: Use PostgreSQL as the main database

Status: Accepted (2023-04-10)

Context:
  Orders, payments and stock need transactions across tables.
  The team knows SQL well. Data is mostly relational.

Decision:
  Use PostgreSQL for all core services.

Alternatives considered:
  MongoDB: flexible schema, but multi-document transactions
  were not something the team had experience with.

Consequences:
  + Strong transactions and joins.
  - Schema changes need migrations.
  • Status: proposed, accepted, or superseded by another ADR.
  • Context: the problem and the limits at that time.
  • Decision: one or two clear sentences.
  • Consequences: the benefits and the costs. Writing the costs is important.

Simple rules for ADRs:

  • Keep it short. One page is enough. If nobody reads it, it is useless.
  • Keep it in the code repository, next to the code. Not in a separate wiki where it gets lost.
  • Never delete an old ADR. If the decision changes, write a new ADR and mark the old one “superseded”. So the history stays.
  • Only for important, expensive decisions. Do not write an ADR for every small thing.

What about old decisions? The people still on the team write down what they remember. If we do not know the reason, we honestly write “the original reason is unknown” and review today’s situation.

Follow-up question: Someone says: “Now that we have an ADR, does it mean we can never replace Kafka?”

Follow-up answer:

  1. No. An ADR does not lock a decision. It only makes the reason clear.
  2. Now the discussion gets better. The question becomes: “Is the context written in the ADR still true?”
  3. If the context changed (for example, load dropped a lot and Kafka costs too much), we write a new ADR.
  4. The new ADR also includes the cost of migrating.

Red flag: Their fix is “another meeting” or “a 40-page document in the wiki”, or they think recording a decision means it can never change.

Edit on GitHub

5One small change across 7 modulesEasyArchitecture styles

Question: A large online shop has 3 teams and about 20 modules. The product manager asks for a small change: “10% discount on the first purchase, but only if the cart total is above 50 dollars.” The team estimates:

  • 7 modules must change: cart, checkout page, orders, invoices, reports, mobile app, and emails.
  • Each of them calculates the discount on its own.
  • 3 teams must coordinate and deploy together.
  • Estimate: three weeks.

What does this situation tell you about the architecture?

Short answer: This is a sign of high coupling and low cohesion. The discount rule is one idea, but it is spread across 7 places. Things that change together do not live together. The fix is to move the discount behavior into one place (one module or service that owns discounts). Everyone else just asks it for the result.

Simple definitions:

  • Cohesion: things that change together live together. A module with high cohesion has one clear job.
  • Coupling: how much a change in one module forces other modules to change.
  • Goal: high cohesion inside a module, low coupling between modules.

Why is this bad? (step by step)

  1. The discount rule is written 7 times. Each copy may be slightly different.
  2. Each change means 7 changes. The chance of forgetting one is high.
  3. If one is forgotten, the cart shows one amount and the invoice shows another. Customers complain.
  4. Three teams must coordinate. So every team moves at the speed of the slowest one.
  5. These are real costs: three weeks for a one-line business change.

How do we fix it?

  1. Pick one owner for price and discount rules. For example, a “pricing” module.
  2. Move all discount calculations into that module.
  3. Everyone else asks: “What is the final price of this cart?” and uses the answer.
  4. Invoices, emails and reports do not calculate the discount. They show the amount saved on the order.
  5. Now the next change happens in one module and one team.

We do not do this all at once. We build this new rule in the new module. Then we connect the old modules to it one by one.

One trap: If we only put the shared code in a shared library that all 7 modules import, every change still means updating and redeploying 7 modules. The behavior needs an owner, not just shared code.

Follow-up question: How do you find high coupling before it is too late?

Follow-up answer:

  1. Look at git history. Files in different modules that always change in the same commit are a good sign.
  2. Count the changes that always need several teams.
  3. Draw the module dependency graph. Cycles (A depends on B, B depends on A) are a warning.
  4. Ask: “For this business change, how many places must change?” A good answer is “one”.

Red flag: Their only fix is “coordinate more” or “a weekly meeting between teams”, or they know coupling and cohesion only as textbook definitions and cannot point to them in this example.

Edit on GitHub

6Stock and likes during a network splitMediumConsistency and CAP

Question: An online shop runs in two data centers (A and B). Both data centers serve user requests and sync data between them. The product has two features:

  • Stock count during checkout: for example, only 3 units of a phone are left.
  • A “likes” counter under each product.

One day the network link between the two data centers is down for 20 minutes. Both data centers are healthy and users can still reach both, but the two data centers cannot talk to each other. How should each of these two features behave during these 20 minutes?

Short answer: The two features have different answers. By CAP, during a network split (partition) we must choose between consistency and availability. For stock, consistency matters more: it is better to reject checkout for a while, or let only one data center sell, so one item is not sold twice. For likes, availability matters more: both sides accept likes, and after the network comes back, the counts are added together.

Hint (if the candidate says there is no problem): During those 20 minutes, two users, one on data center A and one on data center B, buy the last phone. Each data center sees stock as 1 and accepts the sale. What happens now?

Why can we not have both? (step by step)

  1. The link between the data centers is down. So each one does not know what the other did.
  2. If both keep accepting writes (available), their data may differ (inconsistent).
  3. If we want the data to stay the same (consistent), one side must reject writes until the network is back.
  4. So during a split, we lose one of the two. This is a business choice, not only a technical one.

Stock: we choose consistency.

  • Selling one item to two people means a cancelled order, a refund, and an unhappy customer.
  • Option: each item’s stock has one “owner” data center, and only that one accepts sales. The other side shows “please try again in a few minutes”.
  • Another option: split the stock between the data centers ahead of time (for example, 2 here, 1 there). Each side sells only its own share.
  • Some businesses accept a little overselling on purpose and contact the customer later. That is also a choice, but it must be a conscious one.

Likes: we choose availability.

  • If the like count is a little wrong for a few minutes, nobody is hurt.
  • Rejecting a like makes the user experience worse, for no benefit.
  • Each data center keeps its own counter. After reconnecting, the counts are added. Because addition does not depend on order, merging is simple.

Beyond CAP: PACELC. Even when the network is fine, there is a choice. Wait until both data centers confirm a write (consistent but slower), or answer quickly (fast but temporarily inconsistent). For stock, we accept a bit of extra latency. For likes, we pick speed.

Follow-up question: The product manager says: “Pick one database for the whole system: either consistent or available.” What do you think?

Follow-up answer:

  1. The CAP choice is per type of data, not for the whole system.
  2. Money, stock and bookings usually need consistency.
  3. Likes, view counts, recommendations and shopping carts usually prefer availability.
  4. One system can behave differently per part, with one database tuned differently per use, or with different data stores.

Red flag: Says “we get all three”, or only memorized CAP as “pick two of three” and does not know that the choice only matters during a network split.

Edit on GitHub

7A 24-hour cache for product pagesMediumCaching patterns

Question: In an online shop, 95% of traffic is product page reads. The database is under load. The team proposes this plan:

  • Put Redis in front of the database.
  • When a product page is read, check the cache first. If it is not there, read from the database and put it in the cache.
  • The time to live (TTL) for each product in the cache is 24 hours.
  • When a price changes, only the database is updated.

There are about 200,000 products. The sales team changes prices several times a day, especially during sales. What do you think of this plan?

Short answer: A cache is the right idea for this load. The read pattern is cache-aside, which is fine. The problem is that when a price changes, the cache does not know. So for up to 24 hours, an old price can be shown. When a price changes, we must delete that product’s key from the cache (invalidation), and also make the TTL shorter, so that if the delete fails, the error does not last long.

Hint (if the candidate says there is no problem): A sale ends at 10 am and prices go back up. At 4 pm, some customers are still ordering at the sale price. Other customers complain that the product page shows one price and the checkout page shows another.

Why does this happen? (step by step)

  1. At 9 am, a product page is read. The sale price goes into the cache for 24 hours.
  2. At 10 am, the price changes in the database. The cache is not touched.
  3. Until 9 am tomorrow, everyone who opens the page gets the old price from the cache.
  4. If checkout reads the price from the database, we have two different prices. If it reads from the cache, we sell at the wrong price.

The fix:

  1. On price change: update the database first, then delete that product’s key from the cache. The next read loads the fresh value from the database.
  2. If several services can change prices, it is better to publish a “price changed” event and let one consumer clear the cache. So no path is forgotten.
  3. Make the TTL shorter (for example, a few minutes). This is a safety net: if a delete message is lost, the error only lasts a few minutes.
  4. For money, the database is the source of truth. Checkout always reads the final price from the database (or the pricing service), never from the cache.

Why delete and not update the cache? If two price changes happen at the same time, the cache writes can land in the wrong order and an older value can stay. Deleting is simpler: the next read always comes from the database.

Another thought: Maybe cache parts of the page separately. Descriptions and images rarely change and can have a long TTL. Price and stock get a short TTL.

Follow-up question: Now the TTL is 5 minutes. A best-selling sale product is read thousands of times per second. What happens when its key expires?

Follow-up answer:

  1. At that moment, thousands of requests see an empty cache and all go to the database together. This is called a cache stampede.
  2. Option one: only one request may load the value from the database (a short lock). The others wait a little or get the previous value.
  3. Option two: add a small random amount to the TTL so keys do not all expire at the same time.
  4. Option three: for very hot keys, refresh the value in the background before it expires.

Red flag: Sees the cache only as a speed tool and never asks “what happens when the data changes?”, or suggests reading the final checkout price from the cache too.

Edit on GitHub

8A five-service chain in checkoutMediumMicroservicesResilience with Polly

Question: An online shop gets about 20,000 orders a day. The team has prepared this plan for checkout:

Checkout API
  -> 1. User service       (check user)
  -> 2. Cart service       (get items)
  -> 3. Inventory service  (reserve stock)
  -> 4. Payment service    (charge card)
  -> 5. Email service      (send receipt)
  -> return "order placed"

All calls: HTTP, one after another
Timeout per call: 30 seconds
Retries: none

Each service alone is usually healthy 99.9% of the time and answers in about 100 ms. Review this plan. What do you think?

Short answer: In a sync chain, availability multiplies and latency adds up. So checkout as a whole is more fragile and slower than any single service. A 30-second timeout also means that if one service gets slow, requests and threads get stuck in every service above it, and the failure spreads (cascading failure). We should shorten the chain, make work that does not need to happen right now (like email) async, and set short, realistic timeouts.

Hint (if the candidate says there is no problem): One day the email service gets slow, and each request takes 25 seconds. The customer’s card is charged, but the checkout page keeps spinning. A few minutes later, the Checkout API stops answering completely, even for users who are only looking at their cart.

Why is this plan fragile? (step by step)

  1. Availability: each service is 99.9%. For success, all 5 must be healthy. 0.999 to the power of 5 is about 99.5%. That is roughly 5 times more failures.
  2. Latency: 5 calls one after another, about 100 ms each, means at least half a second. Any slow call adds to the total.
  3. In the worst case, each call waits up to 30 seconds. So one request can be stuck for two minutes.
  4. During that time, every stuck request holds a connection or a thread. New requests pile up behind them. The Checkout API runs out of resources and fails for everyone.
  5. Email does not need to be on the main path at all. But now slow email breaks checkout.

The fix:

  1. Make work the user does not wait for async. After the order is placed, publish an “order placed” event. The email service picks it up from a queue. If email is down for an hour, orders still go through.
  2. Reduce the number of hops. For example, send the cart contents with the request, or put the needed user data in the token.
  3. Give each call a short, realistic timeout (based on that service’s normal latency, not 30 seconds). Give the whole request a total time budget too.
  4. For temporary errors, add limited retries with backoff and some randomness (jitter). But only for operations that are safe to repeat (idempotent). Charging a card is not retried without a unique key.
  5. Put a circuit breaker in front of failing services so we fail fast instead of waiting.

Follow-up question: Why can simple retries (for example, 3 immediate retries) on every call make things worse?

Follow-up answer:

  1. If a service is overloaded, every retry adds more load. With 3 retries, load can grow up to 4 times.
  2. If every layer in the chain retries, the number of attempts multiplies.
  3. For payment, a retry without a unique key may charge the card twice.
  4. So: few retries, with growing delay and jitter, in one layer only, and only for idempotent operations.

Red flag: Only says “increase the timeout” or “add more servers”, and does not see that email should not be on the sync checkout path at all.

Edit on GitHub

9Two services sharing the same tablesMediumMicroservicesDomain-Driven Design

Question: A company has two teams: Orders and Billing. Each team has its own microservice and deploys on its own. To “save time”, both services connect to one database and read and write each other’s tables directly:

Orders service   --read/write-->  orders, order_items, invoices
Billing service  --read/write-->  orders, invoices, payments

For example, when a payment succeeds, the Billing service directly sets the status column in the orders table to “paid”. What do you think of this design?

Short answer: This design splits the two services for deployment, but glues them together through data. This is called a distributed monolith: we pay the costs of microservices (network, separate deploys, several teams), but we lose the main benefit (independent change and deploy). Each service must own its data. The other service reaches that data only through an API or through events.

Hint (if the candidate says there is no problem): The Orders team changes the status column from text to a number and runs the migration. That night, the Billing service starts failing and no invoices are created. The Billing team did not know about the change.

What is wrong with this design? (step by step)

  1. The table structure is a hidden contract between two teams. But nobody treats it as a contract.
  2. Any table change (column name, type, the meaning of a value) can break the other service.
  3. So before every migration, both teams must coordinate, and sometimes deploy together. They are no longer independent.
  4. Business rules are skipped. For example, Orders has a rule: “a cancelled order must not become paid”. Billing changes the column directly and never sees this rule.
  5. When data is wrong, it is not clear which service wrote it.
  6. Both services load the same database. You cannot scale one of them separately.

The fix: each service owns its data.

  1. The orders and order_items tables belong only to Orders. The invoices and payments tables belong only to Billing.
  2. If Billing needs order data, it has two options:
    • Call the Orders API, when fresh data is needed right now.
    • Listen to the “order placed” event and keep a copy of the fields it needs in its own database.
  3. When a payment succeeds, Billing publishes a “payment completed” event. Orders receives it and changes the status itself, with its own rules.
  4. Now the API and the events are the official contract. They can be versioned and tested.

Also name the cost: No more JOINs across the two services’ data. A copied value may be a few seconds behind. There is no shared transaction either, so multi-step work needs patterns like Outbox and Saga.

Migrate step by step: We do not split everything overnight. First, name the owner of each table. Then move writes into other teams’ tables behind an API. Split the reads last.

Follow-up question: If these two services need each other’s data so much, maybe they should not be separate at all. How do you decide?

Follow-up answer:

  1. If almost every change touches both, the boundary may be in the wrong place.
  2. A serious option is to merge them back into one service, with two separate modules inside (a modular monolith).
  3. If the two teams really are separate and the concepts are different (orders vs money), staying separate is right, with clear data ownership.

Red flag: Calls the shared database “simpler” and does not see that every table change locks the two teams together.

Edit on GitHub

10Renaming a field in a public APIMediumSchema evolution

Question: We have a public API used by Android and iOS mobile apps. There are about 500,000 active users. The current response of one endpoint is:

{
  "id": 42,
  "price": "129000",
  "user_name": "sara"
}

The backend team wants to make these changes in the next release:

{
  "id": 42,
  "price": 129000,
  "customerName": "sara"
}

So the type of the price field changes from text to a number, and user_name is renamed to customerName. The mobile team will update the app at the same time. What is your release plan?

Short answer: Both changes are breaking. You can change the server overnight, but not the mobile apps. Many users do not update the app for weeks or months. So old app versions still expect the old field with the old type, and they break. The right way: make the change additive (add the new field, keep the old one), have a deprecation period with a clear date, and measure how many users still run old versions. If the change is big, release a new API version.

Hint (if the candidate says there is no problem): On release day, the new app version works fine. But crash reports from phones with the previous app version go up. These users open a product page and the app closes.

Why does this happen? (step by step)

  1. A web app is changed for everyone with one deploy. A mobile app lives on the user’s phone, and we do not control it.
  2. The user must update the app. Many people turn off auto-update or update late.
  3. The old version looks for the user_name field. That field is gone. Depending on the app code, it shows an empty value or crashes.
  4. The old version expects price as text. Now a number comes. JSON parsing may fail.

The right plan:

  1. Only add, never remove or change. The new fields come next to the old ones:
{
  "id": 42,
  "price": "129000",
  "priceAmount": 129000,
  "user_name": "sara",
  "customerName": "sara"
}
  1. The new app version uses only the new fields.
  2. Mark the old fields as deprecated, and write the removal date in the docs.
  3. Using logs or an app version header, measure what percent of requests still come from old versions.
  4. When that number is very low, or when we have added a “please update the app” message (a forced minimum version) for very old versions, remove the old fields.

When do we need a new API version (for example, v2)? When the changes are many and structural, and adding fields is not enough. Then both versions run side by side for a while. The cost is maintaining two versions, so for one field it is usually not worth it.

Simple rule: Renaming a field, changing a type, removing a field, and making an input field required are all breaking. Adding an optional field is usually not breaking, as long as clients ignore unknown fields.

Follow-up question: How do you make sure a future change does not break the API without anyone noticing?

Follow-up answer:

  1. Write the API contract formally (for example, with OpenAPI) and keep it in the repository.
  2. In CI, compare the new contract with the previous one. There are tools that detect breaking changes in OpenAPI files.
  3. Write contract tests that check the server response still has what older client versions expect.
  4. Teach the mobile team to make clients ignore unknown fields.

Red flag: Thinks there is no problem because the mobile team updates the app at the same time, and does not see the old app versions still living on users’ phones.

Edit on GitHub

11A message consumer and duplicate messagesMediumIdempotency

Question: An online shop gets about 20,000 paid orders a day. After each payment, the payment service puts an “OrderPaid” message on a broker. The loyalty service reads these messages and gives the customer points. The broker is set to at-least-once delivery. This is the consumer design:

on message OrderPaid(orderId, customerId, amount):
    points = amount / 10
    UPDATE customers SET points = points + :points WHERE id = :customerId
    ack(message)

You see this design in an architecture review. What do you think? Is there a problem?

Short answer: With at-least-once delivery, a message can arrive more than once. This consumer adds points every time. So a duplicate message means double points. The fix is an idempotent consumer: processing the same message again does not change the result. This is usually done with a “processed messages” table and a natural key such as the order id.

Hint (if the candidate says there is no problem): Support says a few customers a week complain that their points are “too high”. The checks show it happens mostly on days when the loyalty service was redeployed or restarted.

Why do duplicates arrive? (step by step)

  1. The consumer gets the message and saves the points in the database.
  2. Just before the ack, the service restarts (deploy, crash, or a network break).
  3. For the broker, the message is still not confirmed.
  4. So the broker gives the same message to a consumer again.
  5. The consumer adds the points again. Now the customer got points twice.

The producer can also send duplicates. For example, if the broker’s reply does not reach it, it sends again. So do not expect “exactly once” from the broker. Handling duplicates is the consumer’s job.

The fix: an idempotent consumer

  1. Each message has a unique id. A natural key, like the order id, is better. If the producer rebuilds the message, the message id changes, but the order id does not.
  2. We create a table of processed messages with a unique constraint on this key.
  3. Writing to this table and adding the points happen in one transaction.
  4. If a duplicate arrives, the insert fails because of the unique constraint. So no points are added, and we just ack.
on message OrderPaid(orderId, customerId, amount):
    begin transaction
        INSERT INTO processed_orders(order_id) VALUES (:orderId)  -- unique
        UPDATE customers SET points = points + :points WHERE id = :customerId
    commit
    ack(message)
    -- on unique violation: rollback, then ack (already done)

Another option: instead of “add to the total”, store one row per order in a points ledger table (unique key on the order id). The balance is the sum of these rows. This design is idempotent by nature, and it also keeps history.

A few design notes:

  • The processed messages table grows. It needs a cleanup policy. Keep rows longer than the longest time a duplicate can still arrive.
  • “Read first, then write if missing” without a transaction and a constraint is not enough. Two consumer instances can see the same message at the same time.

Follow-up question: What if the consumer’s job is not a write to its own database, but a call to an external API (for example, sending an SMS)?

Follow-up answer:

  • Here we cannot have one shared transaction.
  • If the external API accepts an idempotency key, we send the order id as the key. Then the API itself ignores the duplicate.
  • If it does not, we record the state before the call (for example “sending”). But we must accept that a rare duplicate is still possible. So we decide which is worse: a duplicate SMS or a lost SMS.

Red flag: Thinks the broker guarantees “exactly once” so the consumer needs to do nothing, or relies only on “check first” with no transaction to stop duplicates.

Edit on GitHub

12A slow dependency and the product pageMediumResilience with Polly

Question: The product page of an online shop gets about 500 requests per second at peak time. To build each page, the product page service does these steps one after another:

  1. It gets product details from the catalog service.
  2. It gets price and stock from the inventory service.
  3. It gets a “recommended products” list from the recommendation service.
  4. It builds the page and returns it.

All three calls use HTTP with the HTTP client’s default settings. The product page service has a limited thread pool (or a limited number of workers). You see this plan in an architecture review. What do you think?

Short answer: This plan has no protection against a slow dependency. If the recommendation service gets slow, all workers wait for it and the whole product page goes down, even though recommendations are not critical. The fixes: a short timeout, a circuit breaker, a fallback (for example, the page without recommendations), and a bulkhead (separate resources for each dependency).

Hint (if the candidate says there is no problem): One day the recommendation service gets slow because of a heavy query. Each response takes about 5 seconds. It does not return errors; it is only slow. A few minutes later the whole site is down. Even pages that do not need recommendations stop responding.

Why did the whole site go down? (step by step)

  1. Each product page request holds one worker until it finishes.
  2. Before, each request took about 100 ms. Now it takes 5 seconds, because it waits for recommendations.
  3. With 500 requests per second and a 5 second wait, about 2,500 requests are in progress at once. That is far more than the number of workers.
  4. All workers are busy. New requests wait in a queue or get rejected.
  5. Users refresh the page. So the load goes up.
  6. If other pages run on the same service or the same pool, they go down too. This is called a cascading failure.

Key point: a slow service is more dangerous than a dead one. A dead service fails fast. A slow service holds on to our resources.

The fix:

  1. Timeouts: set a short, realistic timeout on every call. For example, if recommendations usually answer in under 200 ms, use a timeout around 300 ms. Client defaults are usually much too long.
  2. Circuit breaker: if the error or timeout rate goes over a limit, the circuit “opens”. For a while we do not call the recommendation service at all, and we return the fallback at once. After a few seconds we let a few test requests through (the half-open state). If they succeed, the circuit closes.
  3. Fallback: show the page without recommendations, or with a fixed list of best sellers from a cache. The user can still buy.
  4. Bulkhead: give each dependency its own concurrency limit. For example, at most 50 concurrent calls to recommendations. If it gets slow, only those 50 slots fill up, not all workers.
  5. Parallel calls: the three calls do not depend on each other. So they can run in parallel. The total time becomes the slowest call, not the sum.

Classify dependencies: the architect should ask which dependency is critical and which is optional. Without the catalog, the page makes no sense. Without recommendations, the page still works. Decide the failure behavior for each one in advance.

Follow-up question: The team wants to add 3 retries on every error. What do you think?

Follow-up answer:

  • For a service that is already overloaded, retries multiply the load. Three retries means up to four times the requests. This can knock the service down completely (a retry storm).
  • If retries are needed: only for idempotent operations, few of them, with exponential backoff and jitter (a random delay), and inside the request’s overall timeout.
  • Always together with a circuit breaker, so we do not retry while the service is clearly broken.

Red flag: Only thinks about “the service is down” and does not see “the service is slow”, or sees retries as the fix with no timeout and no limits.

Edit on GitHub

13"Checkout is sometimes slow"MediumDistributed tracing

Question: An online shop has 8 services (gateway, cart, pricing, discounts, inventory, payment, orders, notifications). Each service runs several instances. Each service only writes its own text logs to its own disk. The current dashboard shows only the average response time of each service, and all averages look fine. Users say “checkout is sometimes very slow”. As the architect, how do you find the cause? And what do you change in the system?

Short answer: With these tools you cannot follow one request across 8 services. You need three things: structured, central logs with a correlation id, metrics with percentiles (like p95 and p99) instead of averages, and distributed tracing (for example with OpenTelemetry) to see exactly where a slow request spent its time.

Why is finding the cause hard today? (step by step)

  1. One checkout request goes through several services and several instances.
  2. Each instance keeps its logs on its own machine. So we must visit many machines.
  3. There is no shared id in the logs. So we do not know which log line in inventory belongs to this request.
  4. “Sometimes” means the problem hits only a small share of requests. For example 2%.
  5. The average hides this small share. If 98% of requests take 200 ms and 2% take 8 seconds, the average is about 350 ms. That number looks “fine”.

The fix: the three pillars of observability

  1. Metrics: show response time as percentiles: p50, p95, p99. p99 means 99% of requests are faster than this number. This is where “sometimes” becomes visible. Add error rate and request count next to it.
  2. Logs: structured logs (for example JSON) sent to one central place. Every log line has a correlation id (or trace id). So one search finds all logs of one request across all services.
  3. Traces: each request is a trace. Each step in each service is a span with a start and end time. The trace id travels to the next service in headers. The result is a timeline that shows, for example, that 7 of the 8 seconds were spent in the discount service’s database call.

Why OpenTelemetry? It is an open, vendor-neutral standard for traces, metrics and logs. It has libraries for many languages and frameworks. So we can change the backend tool later without changing our code.

Practical steps for this problem:

  1. Add OpenTelemetry to all services. Start with the gateway and the checkout path.
  2. Tracing every request is expensive, so we sample. But it is better to always keep slow and failed traces (tail-based sampling), because those are the ones we need.
  3. Look at the checkout p99. Open a few slow traces and look for a pattern: always the same service? One instance? One kind of cart?
  4. Define an SLO for checkout. For example, “99% of checkouts finish under 2 seconds”. Alert on this number, not on the average.

Follow-up question: The trace shows the time is spent waiting to get a connection from the database connection pool, not in the query itself. What next?

Follow-up answer:

  • First look at pool metrics: how many connections are in use, how many requests are waiting.
  • Then ask who is holding the connections. For example, a transaction that calls an external API in the middle and keeps the connection open.
  • Just making the pool bigger usually does not fix the cause, and it can move the pressure to the database itself.

Red flag: Only says “we will write more logs”, trusts the average response time, or has no way to follow one request across services.

Edit on GitHub

14"We have no time for refactoring"MediumTechnical debt

Question: You are the architect of a product that has been in development for 5 years. Two teams of 6 people work on it. The product manager says: “We have no time for refactoring. Only features.” On the other side, the teams say they deliver fewer features every quarter and see more bugs after each release. A senior developer proposes: “Let’s stop features for six months and rewrite the whole system.” What do you do as the architect?

Short answer: Neither extreme is right. The architect’s job is to make technical debt visible, explain its cost in business terms (delivery time, bugs, risk), and propose small, continuous paydown that is tied to features. A big-bang full rewrite carries a very high risk.

Why does “only features” not work? (step by step)

  1. Every shortcut in the code makes the next change in that area harder. It is like interest on a loan.
  2. If nothing is ever paid back, the interest adds up.
  3. The result is what the teams see: each feature takes longer and more things break.
  4. So “we have no time” really means we will have less time every quarter.

Why is “a six month rewrite” also dangerous?

  • For six months the business gets nothing new.
  • The old system has many hidden rules that are in no document. Rewrites usually take longer than estimated.
  • The market does not wait. Soon there is pressure to add features to both systems.

The fix: a plan you can actually run

  1. Make it visible: build a technical debt list. For each item: where it is, what pain it causes, and roughly what it costs to fix.
  2. Measure: use numbers a manager understands. For example, the DORA metrics: time from commit to production (lead time), deploys per week, the share of deploys that caused problems, and time to recover. Add the number of bugs after each release. Show the trend of these numbers over several quarters.
  3. Prioritize by pain: not all debt matters. Debt in code that changes every week matters. Debt in code nobody has touched for years can wait.
  4. Tie it to features: when a feature touches an area, clean that area a little (the boy scout rule: leave the code cleaner than you found it). The feature estimate includes this work.
  5. A fixed share: keep a small, fixed share of each sprint’s capacity for debt. Agree on the exact number with the team and the manager, and show the result in the metrics.
  6. Stop new debt: automated tests, code review, and ADRs for important decisions. If debt is taken on purpose (for example, for an important deadline), record it and set a time to pay it back.

The right language with a manager: do not say “the code is dirty”. Say “payment features now take roughly twice as long as a year ago. If we fix these two modules over three sprints, we expect that time to go down. Then we will show it with numbers.”

Follow-up question: What if one part is really so bad that it must be replaced?

Follow-up answer:

  • Still no big-bang rewrite. Replace that part piece by piece with the Strangler Fig pattern (question 18).
  • Draw a clear boundary (an interface) around that part, write behavior tests, then change the implementation behind the boundary.
  • Each piece goes to production on its own. So the business gets value during the work too.

Red flag: Explains technical debt only in technical terms and cannot state its cost to a manager, or the only fix they offer is a full rewrite.

Edit on GitHub

15Build it ourselves or buy it?MediumArchitecture Decision Records

Question: A 40-person company sells online accounting software (SaaS) for small businesses. The tech team has 12 people. Now the team proposes to build its own authentication and user management system from scratch (login, passwords, two-factor login, SSO for big customers). Their reason: “We want full control and no dependency on any vendor.” What do you think?

Short answer: For this company, authentication is not the core domain. Customers pay for accounting, not for the login page. Building it brings high security risk and a high total cost of ownership (TCO). It is usually better to use a proven, ready solution (a managed service or a mature open source product) and control the dependency through standards (OpenID Connect, OAuth 2.0, SAML) and a thin layer of our own.

How I think about it (step by step)

  1. First I ask: does this part make us different from competitors? In DDD, the parts that make you different are the core domain. Parts everyone needs but that do not make you different are generic subdomains. For accounting software, authentication is generic.
  2. Then I look at risk: authentication leaves no room for small mistakes. Correct password storage, protection against password-guessing attacks, session handling, password reset, two-factor login, and correct use of the standards. One small mistake can mean a customer data leak.
  3. Then I count the total cost, not only the build cost: the first version is just the start. Then come maintenance, security updates, audits, night-time on-call, and new features customers ask for (for example SSO with their own company system).
  4. Then the opportunity cost: every month 3 people spend on the login page is a month they do not spend on accounting.
  5. Finally I make “control” precise: what control do we actually need? Usually the answer is: our user data, the look of the login page, and the option to change vendors. We can get these without building from scratch.

How do we control the dependency?

  • Connect to the identity system through open standards (OpenID Connect, SAML), not a vendor’s private API.
  • Our own code depends only on a small interface. So changing vendors later is possible, even if it takes work.
  • Check before choosing that we can export our user data.
  • Record the decision in an ADR: the options, why we chose this one, and the conditions that will make us review it.

When is building the right choice?

  • When the part really is core and makes us different.
  • When no ready solution meets a real need of ours (for example a special regulation or a strict performance limit).
  • When the team has the skills and the capacity to maintain it for years.
  • When the cost or scale of the ready solution is truly unacceptable, and we have shown this with numbers.

The same logic for building our own message queue: a message queue is usually generic too. Mature products (like Kafka or RabbitMQ) have spent years on hard problems such as durability, replication and message ordering. Building our own usually solves the same problems again, with lower quality.

Follow-up question: The finance manager says the monthly cost of the ready service grows as users grow. What do you answer?

Follow-up answer:

  • That is a fair point. We should calculate the cost for a few growth scenarios, and put the cost of building and running our own (salaries, infrastructure, risk) next to it.
  • There is also a third option: a mature open source product that we host ourselves. Lower license cost, but higher operating cost.
  • Define a review point. For example, “if the cost goes above a certain amount, we decide again”. Write it in the ADR.

Red flag: Only sees the cost of building the first version, not the maintenance cost and the security risk, or says “always buy” or “always build” with no criteria.

Edit on GitHub

16Saving an order and publishing an eventMedium to hardOutbox and InboxKafka

Question: The order service saves about 50,000 orders a day. After each order is saved, an “OrderCreated” event is published to Kafka. The inventory, email and reporting services read this event. This is the code that creates an order:

function createOrder(request):
    order = new Order(request)
    db.beginTransaction()
    db.insert(order)
    db.commit()
    kafka.publish("orders", OrderCreated(order.id, order.items))
    return order.id

You see this code in a code review. What do you think? Is there a problem?

Short answer: This code writes to two separate systems (the database and Kafka), and no transaction covers both. This is the dual write problem. If something fails between the two steps, the order is saved but the event is never published. The common fix is the Transactional Outbox: write the event in the same transaction, into an outbox table in the same database, and let a separate relay send it to Kafka.

Hint (if the candidate says there is no problem): A few times a month, an order exists in the system but inventory never reserved stock for it and no confirmation email was sent. These cases usually show up around deploys, or when Kafka was unavailable for a few seconds.

Why does this happen? (step by step)

  1. The database transaction commits. Now the order is surely saved.
  2. Before publish runs, the service crashes or is stopped for a deploy. Or Kafka is unavailable and publish fails.
  3. The order is in the database, but no event was sent. And nobody tries again.
  4. So inventory and email never hear about this order.

Changing the order does not help either:

  • If we publish first and then commit, the event may go out but the commit may fail. Now inventory reserves stock for an order that does not exist.
  • If we put publish inside the transaction, Kafka is still not part of the database transaction. If the commit fails after the publish, we have the same problem.

The fix: Transactional Outbox

  1. Create an outbox table in the same database as the orders.
  2. In the same transaction, save both the order and the event. So either both are saved, or neither.
  3. A separate process (the relay) reads unsent outbox rows, sends them to Kafka, and marks them as sent after Kafka confirms.
  4. Instead of polling the table, you can use CDC (Change Data Capture, for example with Debezium), which reads changes from the database log.
function createOrder(request):
    order = new Order(request)
    db.beginTransaction()
    db.insert(order)
    db.insert(outbox, { id: newId(), type: "OrderCreated", payload: order })
    db.commit()
    return order.id

// relay process, runs separately
loop:
    rows = db.select(outbox where sent = false order by id limit 100)
    for row in rows:
        kafka.publish("orders", row.payload, key = order id)
        db.update(outbox set sent = true where id = row.id)

An important result: at-least-once. If the relay crashes after publish but before marking the row, the same event is sent again. So the Outbox guarantees the event is not lost, but it may arrive more than once. Consumers must be idempotent (question 11). Each event has a unique id so duplicates can be detected.

One more note: if the order of events for one order matters, use the order id as the Kafka message key. Then all events of one order go to the same partition and keep their order.

Follow-up question: What about the consumer side? The inventory service reads the message, writes to its own database, and also publishes a new event.

Follow-up answer:

  • It is the same problem on the other side. So we do two things.
  • Inbox: store the message id in an inbox table, in the same transaction as the inventory changes. A duplicate message is ignored.
  • Outbox in the inventory service: the new event also goes out through the inventory service’s own outbox, not directly.

Red flag: Assumes “if the commit worked, the publish surely happens too”, or thinks a simple try/catch with a retry fixes it (that does not cover a process crash).

Edit on GitHub

17A transaction across order, payment and inventoryMedium to hardThe Saga pattern

Question: An online shop has three separate services: Order, Payment and Inventory. Each service has its own database and is owned by a separate team. At peak time about 200 orders per minute come in. To place an order, the order must be created, money must be taken from the customer’s card, and stock must be reserved. To make it “all or nothing”, the team proposes a distributed transaction with Two-Phase Commit (2PC) across the three databases. What do you think? What would you do?

Short answer: In microservices, 2PC is usually not a good choice. It locks the services together, it depends on all of them being available, and many databases, brokers and external APIs (like a payment gateway) do not support it at all. The common approach is a Saga: a sequence of local transactions. If one step fails, the earlier steps are undone with compensating actions. A Saga has two styles: orchestration and choreography.

Why is 2PC a problem here? (step by step)

  1. In 2PC a coordinator asks every participant “are you ready?”. Each one holds its locks and says “yes”.
  2. Then the coordinator says “commit”. Until that message arrives, the locks stay.
  3. If the coordinator fails between these two phases, participants wait with their locks held. This affects all other orders too.
  4. The whole operation is available only when all three services and the coordinator are healthy at the same time.
  5. The external payment gateway does not take part in our transaction. So 2PC cannot even cover all the work technically.
  6. The goal of microservices is independent teams and services. 2PC glues them back together.

The fix: a Saga

The steps and the compensating action for each step:

Step Action Compensating action
1 Create the order as “pending” Cancel the order
2 Reserve stock in inventory Release the reservation
3 Charge the card Refund the money
4 Confirm the order (last step)

If step 3 fails (the card is declined), the reservation is released and the order is cancelled.

The order of steps matters. Put work that is hard or costly to undo (like charging money) as late as you can. Work that cannot be undone at all (like shipping the goods) comes after every step that might fail.

The two Saga styles:

  • Orchestration: one coordinator (for example in the order service) calls the steps one by one and keeps the Saga state. The flow is visible in one place and tracing errors is easier. The downside: the coordinator can collect too much logic.
  • Choreography: there is no coordinator. Each service listens to events and publishes the next event. Simple for a short flow, with low coupling. The downside: with many steps, it becomes hard to see the whole flow.
  • Rule of thumb: for a short flow with two or three steps, choreography is enough. For a long or complex flow, orchestration is clearer.

Costs we must accept:

  • Eventual consistency: for a while the order is “pending”. The UI and reports must understand this state.
  • No isolation: the rest of the system can see in-between states. For example, stock that is reserved and later released.
  • Compensating actions can fail too. So they must be retried and must be idempotent. Steps talk through messages and an Outbox (question 16).

Follow-up question: The compensating action “refund the money” also fails after several retries. What do you do?

Follow-up answer:

  • Keep the Saga in a “needs attention” state and raise an alert.
  • Some failures must reach a human (for example the finance support team). The system needs tools to see and resolve these cases.
  • Every step and its result is recorded, so the accounts can be reconciled later.

Red flag: Proposes 2PC across microservices without seeing its cost, or explains a Saga with no compensating actions and no thought about in-between states.

Edit on GitHub

18A full rewrite of a 10-year-old monolithMedium to hardArchitecture stylesMicroservices

Question: An insurance company has a 10-year-old monolith. All of the company’s work runs on it: customers, policies, payments and claims. It has little documentation, and several of its original builders have left. Leadership has approved this plan: “A new team rewrites the whole system with a new architecture in 12 months. During that time the old system is frozen (no new features). Then, over one weekend, we switch everything to the new system.” They ask for your opinion as the architect. What do you think?

Short answer: A big-bang rewrite carries a very high risk: hidden rules in the old system, optimistic estimates, a frozen business, and one risky cut-over day. I would propose the Strangler Fig pattern: put a routing facade in front of the old system and move features to the new system slice by slice. Each slice goes to production on its own, until the old system slowly shrinks and can be switched off.

Why is a big-bang rewrite dangerous? (step by step)

  1. Ten years of bug fixes and special cases live in the old code. Many are in no document. The new system does not know about them.
  2. The 12 month estimate is based on what we know. What we do not know shows up later. So it usually takes longer.
  3. The business will not wait 12 months (or more). Pressure for new features starts. Now every feature must be built in two systems.
  4. Until cut-over day, the new system has never worked with real users and real data. So the first real feedback arrives at the riskiest moment.
  5. If a serious problem appears on cut-over day, going back is hard, because new data has already been written to the new system.

The fix: Strangler Fig

  1. Add a facade: put a routing layer (for example an API gateway or reverse proxy) in front of the old system. At first all requests still go to the old system. Users see no change.
  2. Pick the first slice: a feature with a clear boundary and low risk. For example, “view policy status”, which is read-only. Not “pay a claim” as the first step.
  3. Build and compare: build the feature in the new system. For a while you can send requests to both and compare the answers, but show users only the old system’s answer.
  4. Shift traffic gradually: the facade sends this feature to the new system step by step. First a small share of users. If there is a problem, one setting sends it back.
  5. Delete the old code: when all traffic for that feature goes to the new system, remove its old code.
  6. Repeat: take the next slice. The old system gets smaller each time.

The hardest part: data

  • Usually all features share one database. So moving data is harder than moving code.
  • During the transition, a feature may be written in one system and read in the other. So data must be synced between both sides. For example with CDC from the old database, or with events.
  • For each slice, decide who owns the data (the source of truth). At any moment only one side should own a given piece of data.
  • An Anti-Corruption Layer helps stop the old, messy model from leaking into the clean model of the new system.

The answer to leadership: new features are not frozen. New features are built in the new system. Value shows up in the first months, not after 12 months. And each step is small and can be rolled back.

Follow-up question: When is Strangler Fig not a good fit?

Follow-up answer:

  • When you cannot intercept and route requests in front of the old system. For example, a desktop app that connects straight to the database.
  • When the system is so small that a full rewrite takes a few weeks. Then the cost of a facade and data sync is not worth it.
  • When the old platform is no longer supported and there is a hard external deadline. A staged move is still better, but there may not be enough time for every step.

Red flag: Accepts a full rewrite with a frozen old system without seeing the risk, or names Strangler Fig but does not think about data and data ownership.

Edit on GitHub

19One Customer class for everyoneMedium to hardDomain-Driven Design

Question: An online retail company has four teams: Sales, Support, Billing and Shipping. All of them use one shared class called Customer and one customers table. The class has about 80 fields. Some of them:

Customer
  id, name, email, phone
  leadSource, salesRepId, discountTier        // sales
  openTickets, supportPlan, preferredLanguage // support
  taxId, billingAddress, paymentTerms, creditLimit // billing
  shippingAddresses, deliveryNotes, courierPreference // shipping
  ... (about 60 more fields)

The teams say almost every change to this class breaks another team’s work. For example, Sales changed the rule for “active customer” and the billing reports became wrong. Now every change needs a meeting with all four teams. What do you think? What would you do?

Short answer: The word “customer” means something different in each area. One model for everyone locks everyone together. In DDD, each area is a Bounded Context with its own model. Billing has a “billing account”, Shipping has a “recipient”, Sales has a “prospect or account”. Only a shared id passes between them. We show how the contexts relate in a Context Map, and each model has one owning team.

Why does one shared model cause trouble? (step by step)

  1. Each team has a different picture of “customer”. For Sales, it is someone who may buy. For Billing, someone who must pay. For Shipping, a name and an address on a parcel.
  2. When everyone shares one class, a word like “active” means a different thing to each team, but there is only one field for it.
  3. No team really owns the class. So each team makes its own change without knowing who else depends on that field.
  4. As the system grows, the class gets bigger and coordination gets more expensive. Every team slows down.

The fix: Bounded Contexts

  1. Find the boundaries: talk with the teams and the business experts (for example in an Event Storming workshop). Where a word starts to mean something different, there is usually a context boundary.
  2. One model per context: each context keeps only the fields it needs, with the names that team actually uses (the Ubiquitous Language).
Context Model Example fields
Sales Prospect / Account leadSource, salesRepId, discountTier
Support Requester supportPlan, preferredLanguage
Billing BillingAccount taxId, billingAddress, paymentTerms
Shipping Recipient shippingAddresses, deliveryNotes
  1. A shared id: all contexts share one customer id. Basic data (like name and email) has one clear owner. Other contexts keep a copy of what they need, updated by events (for example “CustomerRegistered” or “EmailChanged”).
  2. A context map: the Context Map shows which context depends on which, and the type of relationship. For example upstream/downstream, or an Anti-Corruption Layer when one side’s model must not leak into the other.
  3. Ownership: each context has one owning team. A change inside a context is that team’s decision. Only the contracts between contexts (events and APIs) need coordination.

Move gradually: do not change this in one step. First split out the context with the most pain (for example Billing): its own model, its own tables, and sync through events. Then the others.

Costs: data is copied in several places, with eventual consistency between them. But the copying is on purpose: each context keeps its own copy in its own shape, and there is still only one source of truth for each piece of data.

Follow-up question: So should every Bounded Context be its own microservice?

Follow-up answer:

  • Not necessarily. A Bounded Context is a boundary of model and language, not a deployment boundary.
  • Several contexts can live in one Modular Monolith, as long as the boundaries are kept in code and each module has its own tables.
  • But usually one microservice should not take pieces from several contexts. A good service boundary usually matches a context boundary.

Red flag: The fix is “clean up the class” or “group the fields” while keeping one model for everyone, or treats a Bounded Context as the same thing as a database table.

Edit on GitHub

20Microservices with layered teamsMedium to hardMicroservices

Question: A company has 30 developers in three teams: a web team (user interface), a backend team, and a database team. A year ago they decided to split the system into independent microservices based on business capabilities (orders, catalog, payments, shipping, and so on). Now they have 9 services. But almost every feature still needs all three teams: the web team builds the page, the backend team changes the service, and the database team changes the schema. Releases are still coordinated and go out together, and managers ask: “So where is the benefit of microservices?” What do you think is going on? What do you propose?

Short answer: This is Conway’s law: a system’s structure copies the communication structure of the organization that builds it. The organization is split into layers (web, backend, database), so in practice the work is split into layers too, even though the code is split into services. The fix is the Inverse Conway Maneuver: first shape the teams like the architecture we want. That means stream-aligned teams, each owning one or more business capabilities from the user interface down to the database.

Why did the microservices not become independent? (step by step)

  1. The goal of microservices is that one team can build and deploy its change without waiting for others.
  2. Here no team owns a whole capability. Every feature passes through three teams.
  3. Each team has its own work queue and priorities. So each feature waits on another team at least twice.
  4. Because the web, service and schema changes depend on each other, they must be released together.
  5. The result: the code is split into services, but the work and the decisions are still layered. We pay the cost of microservices and do not get their independence.
  6. Over time the code also drifts toward the team structure. For example, the database team keeps one shared schema for all services, because that is easiest for them.

The fix: shape the teams like the architecture

  1. Stream-aligned teams: each team owns one or a few related services (for example “orders and payments”). The team has web, backend and database skills inside it. So a normal feature is done start to finish inside one team.
  2. Team boundaries follow context boundaries: service and team boundaries come from Bounded Contexts, not from technical layers (question 19).
  3. A database per service: each service owns its data, and its team changes its own schema. Database experts are still needed, but as advisers and tool builders, not as gatekeepers for every change.
  4. A platform team: shared needs (CI/CD, monitoring, database infrastructure) are provided by a platform team as self-service, so stream-aligned teams do not wait for it.
  5. Contracts between teams: services talk through versioned APIs and events. So each team can release on its own.

These names (stream-aligned, platform) come from the book Team Topologies, which also describes two more types: enabling and complicated-subsystem.

A note for the architect: an architect cannot change the architecture with diagrams alone. If the team structure does not change, the architecture slowly drifts back to the shape of the organization. So this conversation must also happen with leadership and engineering managers.

Another option: if the organization will not or cannot change its teams, maybe microservices are the wrong choice. A Modular Monolith with one deploy has a lower coordination cost.

Follow-up question: How do you start changing the team structure without breaking everything at once?

Follow-up answer:

  • Start with one pilot team: one capability with a clear boundary (for example the catalog) and a few people from each of the three current teams.
  • Measure before and after: time from starting work to production, and how many teams each feature needs.
  • If the result is good, move to the next capability. At the same time, record clear ownership for every service.

Red flag: Sees the problem as purely technical (for example “we need better deploy tools”) and misses the link between team structure and architecture, or the fix is more coordination meetings.

Edit on GitHub

21When is Event Sourcing worth it?HardEvent Sourcing

Question: The company has an internal admin panel. Its job is simple: admins create, edit and delete product categories, banners and site settings. About 20 admins use it. The team has 3 people. One team member suggests building the panel with Event Sourcing, because it is “the modern way” and people talk about it a lot at conferences. What do you think?

Short answer: For this panel, Event Sourcing is not a good choice. A simple CRUD app with an audit log table is enough. Why:

  1. In Event Sourcing, we do not store the current state. We store only a list of events and build the state from them.
  2. This approach has real benefits, but it also has a high cost.
  3. This panel has none of the needs that Event Sourcing solves. So we only pay the cost.

Hint (if the candidate says there is no problem): Six months later, the team wants to add a new field to “banner”. They must decide what to do with old events that do not have this field. The banner list page also sometimes shows old data for a few seconds. And a user has asked for their personal data to be deleted. Now what?

Event Sourcing in one look:

  • Instead of “balance = 700”, we store: “account opened”, “1000 deposited”, “300 withdrawn”.
  • To know the current state, we replay the events in order.
  • Events are never changed or deleted. We only add new events.

Where Event Sourcing really helps:

  • When the full history has business value by itself. Examples: accounting, banking, insurance. We must be able to say exactly what happened and why.
  • When we have time-based questions: “What was the balance of this account on March 31?”
  • Complex domains where behavior matters more than data, and events are the shared language of the team and the business.
  • When we later want to build new reports or read models from the same events.

The costs of Event Sourcing (step by step):

  1. Event versioning: old events never change. So when the shape of an event changes, the code must still understand the old versions (for example, by converting old versions to the new one while reading).
  2. Read models (projections): to show lists and search, we cannot replay all events every time. So we build separate read tables and keep them updated from the events. This means more code and more places for bugs.
  3. Eventual consistency: if projections are updated separately, the user can see old data for a moment.
  4. Deleting personal data: laws like GDPR give people the right to have their data erased. But events are supposed to never change. There are workarounds, like keeping personal data outside the events, or encrypting it with a key that we delete later (crypto-shredding). All of them add complexity.
  5. Learning: the team must learn new patterns. For a team of 3, that is a lot of time.

What I suggest for this panel:

  • A simple CRUD app on a relational database.
  • If we want to know who changed what and when, an audit log table is enough: user, time, type of change, value before and after.
  • If one part ever really needs it, we can build only that part with Event Sourcing, not the whole system.

Follow-up question: What is the difference between an audit log and Event Sourcing? Why is an audit log enough for this panel?

Follow-up answer:

  • With an audit log, the source of truth is still the main table. The log is only a side copy. If the log is incomplete, the system still works correctly.
  • With Event Sourcing, the source of truth is the events themselves. The current state is only a result of them.
  • For the admin panel, the only important question is “who did what?”. An audit log answers that, without the cost of projections and versioning.

Red flag: Picks the pattern because it is “modern” and cannot name a single concrete cost of Event Sourcing.

Edit on GitHub

22A reporting dashboard on the main databaseHardCQRS and Mediator

Question: An online shop has one main relational database. All orders, payments and stock are written to it. The sales team has a reporting dashboard that refreshes every few seconds: today’s sales by city, category and campaign. The dashboard queries run directly on this database and JOIN several large tables. About 50 people in sales keep the dashboard open. Look at the design. What do you think?

Short answer: Heavy report reads and order writes compete on the same database. At peak hours, dashboard queries take CPU, memory and disk, and placing an order gets slow. We should move the reporting load away from the main database. There are two main options:

  1. A read replica for the dashboard.
  2. A separate read model (the CQRS idea) that is updated by events.

In both cases, the dashboard is a little behind the real data. For a sales dashboard, that is usually fine.

Hint (if the candidate says there is no problem): At night everything is fast. But from 8 to 10 PM, placing an order goes from 200 ms to several seconds. In database monitoring, the heaviest queries all belong to the dashboard. Sometimes order writes also wait for locks.

Why does this happen? (step by step)

  1. Each dashboard query JOINs and sums several large tables. This needs a lot of memory and disk.
  2. With 50 people and a refresh every few seconds, these queries run one after another.
  3. Orders need the same resources. So they wait.
  4. Depending on the database and isolation level, long reads may also have lock conflicts with writes.
  5. Result: the work that brings money (orders) suffers because of work that can wait a few seconds (reports).

Option 1: read replica

  • The database has a read-only copy that receives changes from the main one. The dashboard connects only to the copy.
  • Benefit: simple. The report code barely changes.
  • Cost: the copy is a bit behind (replication lag). The data model is still designed for writes, so the JOINs are still heavy. They just run somewhere else.

Option 2: separate read model (CQRS)

  • The idea of CQRS is to separate the write model from the read model.
  • When an order is placed, an event is published (ideally with the Outbox pattern, so the event is not lost).
  • A consumer takes the event and updates a summary table. For example: “today’s sales, city, category, total amount”.
  • This table can live in another database, even one built for reporting.
  • Benefit: the dashboard query has no JOINs and is very fast. The main database is completely free.
  • Cost: more code, a message queue, and we must handle duplicate or missed events (the consumer must be idempotent).

Which one do I pick?

Situation Suggestion
The problem is only load, and reports are few Read replica, the simplest step
Reports are slow even on the replica Separate read model with summary tables

Follow-up question: The sales manager says, “The dashboard must show the exact real-time number.” What do you answer?

Follow-up answer:

  • First ask what decision is made with this number. Usually no decision depends on a few seconds.
  • Then state the cost clearly: a real-time number means reading from the main database again, with the same slow orders.
  • The dashboard can show the time of the last update. Then users know how fresh the number is.

Red flag: Only says “let’s buy a bigger database server” and does not see that the real problem is reads and writes competing for the same resource.

Edit on GitHub

23Data design for a multi-tenant SaaSHardSystem designSecurity

Question: We are building a B2B SaaS product. We have about 5,000 small customers, each with a few users and little data. We also have 3 very large customers, each with as much data and traffic as hundreds of small ones. Two of these large customers have it in their contract that their data must be isolated from everyone else’s. How do you design the customer data?

Short answer: No single model is best for everyone. I suggest a hybrid model:

  1. For the 5,000 small customers: one shared database and one shared schema. Every table has a tenant_id column.
  2. For the large customers that require isolation: a separate database for each.
  3. A central catalog says where each customer’s data lives, and the app connects to the right database based on it.

The three main models:

Model Benefit Cost
Shared schema with tenant_id Cheap, simple for thousands of tenants, one migration Risk of data leaks, noisy neighbors
Schema per tenant Better isolation inside one database Migrations across thousands of schemas are hard
Database per tenant Strongest isolation, separate backup and restore Expensive, managing thousands of databases is hard

Problem 1: the noisy neighbor

  1. In the shared model, all customers use the CPU, memory and disk of one database together.
  2. A large customer runs a heavy report.
  3. The database gets busy, and 5,000 small customers get slow too, even though they did nothing.
  4. That is why very large customers should have their own place, even if they do not ask for isolation.

Problem 2: data leaking between tenants

  1. In the shared model, every query must filter by tenant_id.
  2. It only takes one developer forgetting this filter once.
  3. Then customer A sees customer B’s data. For B2B, that is a serious security and contract incident.

Ways to reduce the leak risk:

  • Put the tenant_id filter in one central layer (for example, a global filter in the ORM), not by hand in every query.
  • Use Row Level Security in the database (for example, in PostgreSQL). The database itself hides other tenants’ rows, even if the code forgets the filter.
  • Take the tenant id from the user’s token, not from input the user sends.
  • Write automated tests with two tenants to make sure one tenant’s data never reaches the other.

A small Row Level Security example in PostgreSQL:

ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;

CREATE POLICY tenant_isolation ON invoices
  USING (tenant_id = current_setting('app.tenant_id')::uuid);

Why a hybrid?

  • For small customers, a database each is very expensive and not worth it.
  • For large customers, a separate database satisfies the contract and also solves the noisy neighbor problem. We usually include its cost in the contract price.
  • Because all databases use the same schema, the app code is the same. Only the connection string changes.

Follow-up question: A small customer has grown and is now a noisy neighbor itself. How do you move it to its own database?

Follow-up answer:

  1. Create a new database with the same schema.
  2. Copy that customer’s data using the tenant_id filter, and keep following new changes so the copy stays current.
  3. Pause writes for that customer for a few minutes and move the last changes.
  4. Change that customer’s location in the central catalog.
  5. After checking everything, delete the old data from the shared database.

Red flag: Does not see the risk of a forgotten tenant_id filter, or suggests a database per tenant for all 5,000 small customers without thinking about cost.

Edit on GitHub

24Checking the user's token in every serviceHardSecurityAuthentication and authorization

Question: A system has 15 microservices. There is also one central Auth Service. After login, the user gets a random token. On every request, each service sends the token to the Auth Service and asks, “Is this token valid? Who is the user?”. One user request usually passes through 3 or 4 services. When services call each other, they forward the same user token. Look at the design. What do you think?

Short answer: This design has two main problems:

  1. The Auth Service is a single point of failure. If it gets slow or goes down, the whole system stops.
  2. Each user request goes to the Auth Service several times. That adds a lot of latency and load.

The common fix: a signed token (like a JWT) with a short lifetime. Each service checks the signature itself with a public key and no longer calls the Auth Service. For service-to-service calls, each service also has its own identity.

Hint (if the candidate says there is no problem): On a busy day, the Auth Service has the highest load in the whole system, more than the order service. Once, the Auth Service restarted for 5 minutes. During those 5 minutes no service worked, not even pages that only list products.

Why is this design a problem? (step by step)

  1. A user request passes through 4 services. So the Auth Service is called 4 times.
  2. Each call is a network round trip. These delays add up.
  3. The load on the Auth Service is several times the total user traffic.
  4. If the Auth Service is not available, no service can accept any request.
  5. Forwarding the user token from one service to another also means the target service does not know which service is really calling it.

The fix:

  1. Signed token: after login, the Auth Service creates a JWT and signs it with its private key. The token contains the user id, roles and expiry time.
  2. Local check: each service fetches the public key once (for example, from a JWKS URL) and keeps it in memory. Then it checks the signature and expiry itself. No network call.
  3. Short lifetime: the access token is valid for, say, a few minutes. To get a new one, the client uses a refresh token, and only then is the Auth Service called.
  4. Service identity: when service A calls service B, A has its own identity too. Two common ways: mTLS (each service has a certificate) or OAuth2 Client Credentials (each service gets its own token). This way B knows both who the user is and which service is calling.

The cost of this approach: revocation

  • With signed tokens, services do not ask the Auth Service. So if we block a user, their token still works until it expires.
  • That is why the lifetime is short: shorter means less risk, but more refreshes.
  • For very sensitive actions (like changing a password or a large payment), we can still ask the Auth Service, or share a small list of revoked tokens with the services.

Follow-up question: The signing key has leaked, or we must rotate it regularly. How do you change the key without downtime?

Follow-up answer:

  1. Create a new key and publish its public key next to the old one at the JWKS URL. Each token has a key id (kid) in its header.
  2. Services reload the keys and now accept both.
  3. The Auth Service starts signing new tokens with the new key.
  4. After all old tokens have expired, remove the old key.
  5. If the key has leaked, remove the old key right away and accept that users must log in again.

Red flag: Does not see the single point of failure, or suggests JWT without knowing that revoking it is hard.

Edit on GitHub

25Capacity estimate for a photo appHardSystem designEstimation and planning

Question: We have a photo sharing app. It has 10 million daily active users (DAU). On average, each user uploads 2 photos per day. An average photo is 2 MB. Each user views 50 photos per day. With a quick back-of-the-envelope estimate, tell me: how many requests per second do we have? How many at peak? How much storage do we need per year? What is the bandwidth? Say the steps out loud.

Short answer: The goal is not an exact number. The goal is the method: state the assumptions, round the numbers, and find out which part of the system is the heaviest. The rough result:

Item Average Peak (about 3x)
Uploads about 200 per second about 600 per second
Photo views about 5,000 per second about 15,000 per second
Photo storage per year about 15 PB (no copies)

The main point: the system is read-heavy (25 views per upload), and the main cost is storage and the bandwidth for viewing photos, not CPU.

Step 1: assumptions and rounding

  • One day is 86,400 seconds. To keep the math simple, we use 100,000 seconds. The error is about 15%, which is fine for an estimate.
  • Traffic is not even through the day. We assume the peak hour is about 2 to 3 times the average. Here we use 3 to be safe.

Step 2: requests

  1. Uploads per day: 10 million × 2 = 20 million.
  2. Uploads per second: 20 million ÷ 100,000 = about 200. At peak, about 600.
  3. Views per day: 10 million × 50 = 500 million.
  4. Views per second: 500 million ÷ 100,000 = about 5,000. At peak, about 15,000.

Step 3: storage

  1. Per day: 20 million × 2 MB = 40 million MB = about 40 TB.
  2. Per year: 40 × 365 ≈ 40 × 400 = 16,000 TB. So about 15 PB (the more exact number is about 14.6).
  3. If each file has 3 copies (for durability), it becomes about 45 PB.
  4. Metadata (user, time, caption) is maybe about 1 KB per photo: 20 million × 1 KB = about 20 GB per day. That is tiny compared to the photos.

Step 4: bandwidth

  1. Incoming (uploads): 200 × 2 MB = about 400 MB per second on average.
  2. Outgoing (views), if we send the original 2 MB photo: 5,000 × 2 MB = about 10 GB per second. At peak, about 30 GB per second. That is a lot.
  3. So in practice we also store smaller sizes of each photo (for example, thumbnails for lists). If the photo we show is about 200 KB, outgoing traffic drops by 10 times.

What decisions do these numbers give us?

  • Photos do not go into the database. They go into object storage, and the database keeps only the address and metadata.
  • Photo views must come from a CDN, because outgoing traffic is much bigger than incoming, and photos do not change after upload. So they cache very well.
  • Uploads can go directly from the user to object storage (for example, with a temporary signed URL), so our servers do not carry the file bytes.
  • Creating smaller sizes is heavy work. So we do it asynchronously with a queue.

Follow-up question: After 5 years, the storage cost is very high. What do you do?

Follow-up answer:

  • Old photos are viewed very rarely. So we move them to a cheaper, slower storage tier (cold storage).
  • We keep the original in a more compressed format, if the quality stays acceptable.
  • We really delete photos that users have deleted.
  • Before doing anything, we measure which group of data costs the most.

Red flag: Starts calculating without stating assumptions, gets lost in the math because they do not round, or ignores the difference between average and peak traffic.

Edit on GitHub

26A plan for going multi-regionHardSystem designConsistency and CAP

Question: We run an online booking service. The whole system (app servers and one relational database) runs in one cloud region in the US. Now 40% of users are in Europe and 20% are in Asia. These users say the site is slow. The product manager asks: “Should we go multi-region?”. What is your plan?

Short answer: Going multi-region is a big and expensive decision, especially for data. I would go step by step:

  1. First, measure where the slowness comes from. Often a CDN and fewer round trips solve a big part of it.
  2. Then move reads close to users (app servers and a read-only database copy in Europe and Asia).
  3. Only if needed, also accept writes in several regions (active-active). This is the hardest step, because data conflicts appear.

Next to all of this, we must talk to the legal team about data laws (like GDPR in Europe).

Why does distance cause slowness?

  1. Light in optical fiber has a limited speed. No software can change that.
  2. One round trip from Europe or Asia to the US takes tens to hundreds of milliseconds.
  3. A page usually needs several round trips: connection, TLS, several API calls.
  4. So these delays multiply, and the user waits for seconds.

Step 1: cheap wins

  • Serve static files (images, JS, CSS) from a CDN. They are cached close to the user.
  • Terminate TLS at the network edge close to the user (many CDNs do this).
  • Reduce the number of API calls per page.

Step 2: active-passive, or local reads

  • Each region has app servers and a read-only copy of the database.
  • Reads (search, list views) are local and fast.
  • Writes (making a booking) still go to the main database in the US.
  • If the main region fails, we can promote another region. So we are also better prepared for outages.
  • Cost: copies are a little behind. A user may not see their own booking right after making it. A simple fix: after a write, that user’s reads go to the main database for a few seconds.

Step 3: active-active (writes in several regions)

  • Every region can write. So writes are fast too.
  • The main problem is conflicts: two people in two regions book the last seat at the same moment.
  • Options:
    • Each piece of data has a “home region”. For example, each user is written only in their own region. This is the simplest way.
    • For shared resources (like a seat), writes happen in only one place, even if it is slower.
    • A conflict rule like “last write wins” is dangerous for bookings, because data gets lost.

Data residency laws:

  • GDPR has strict rules about transferring personal data of people in the EU to places outside Europe.
  • Some countries also have laws that data must stay inside the country.
  • So we may need to keep European users’ data in Europe. This is not a technical decision. It must be made with the legal team. The “home region” model makes it easier.

Cost: each region means servers, databases, monitoring, data transfer between regions, and a team to run it all. So we must know how much business this slowness actually costs us.

Follow-up question: You moved to local reads. How do you know it worked?

Follow-up answer:

  • Before and after, measure response time per region and with percentiles (like p95), not only one global average.
  • Also monitor database replication lag in each region.
  • Look at a business number too: for example, the booking completion rate in Europe and Asia.

Red flag: Jumps straight to active-active without thinking about data conflicts, cost and data laws, or decides before measuring.

Edit on GitHub

27Design a URL shortenerDesignWeight ×2System designCaching patterns

Question: Design a URL shortener. A user gives a long URL and gets a short one, like this:

https://sho.rt/aZ3kQ9x

Anyone who opens the short URL is redirected to the original. Requirements:

  • 100 million new links per month.
  • Reads are much more common than writes: about 100 opens for each link.
  • Links are kept for 5 years.
  • The link owner wants to see click statistics.

Short answer: This is a read-heavy system. The heart of the design is three things:

  1. Generating short, unique keys without collisions.
  2. A simple key-to-URL table, with a cache in front of it for fast redirects.
  3. Recording clicks asynchronously, so redirects do not slow down.

Step 1: requirements and numbers

  1. Writes: 100 million per month ≈ 3.3 million per day. Using 100,000 seconds for a day, about 40 writes per second.
  2. Reads: 100 times more, about 4,000 per second. At peak maybe 3 times that, about 12,000.
  3. Total links in 5 years: 100 million × 60 months = 6 billion.
  4. If each record is about 500 bytes, all data is about 3 TB. Not much.

Step 2: key length

  • With lowercase, uppercase and digits (base62), each character has 62 options.
  • 6 characters give about 56 billion options, and 7 characters about 3.5 trillion.
  • For 6 billion links, 7 characters leave plenty of room.

Step 3: generating keys (the hard part)

Method Benefit Cost
Hash the URL and take the first 7 characters Simple Collisions can happen; must check and retry
Global counter converted to base62 No collisions The central counter is a bottleneck; keys are guessable
Pre-reserved ranges No collisions, no round trip per link A bit more complex; if a server dies, part of a range is wasted
  • I suggest ranges: each server takes, say, 10,000 numbers at once from a central service and uses them itself.
  • If we do not want keys to be sequential and guessable, we shuffle the number with a reversible function before converting it.

Step 4: data model

  • Main table: short key (primary key), long URL, owner, created time, expiry time.
  • Access is only by key, with no JOINs. So a key-value store or a relational database partitioned by key both work.

Step 5: redirect and cache

  1. The request reaches a server. It checks the cache (like Redis) first.
  2. On a miss, it reads from the database and puts the result in the cache.
  3. Links almost never change after creation. So caching works very well. Usually a small share of links gets most of the clicks.

301 or 302?

  • With 301 (permanent), the browser caches the answer. Next time it does not even reach our server. Less load, but we do not count later clicks.
  • With 302 (temporary), every click reaches our server. More load, but more accurate statistics, and we can change or disable the link later.
  • Because click statistics are a requirement, I choose 302.

Step 6: click statistics

  • On the redirect path, the server only drops a click event into a queue (like Kafka) and answers right away.
  • A consumer aggregates clicks (for example, count per hour per link) and writes them to a separate reporting database.
  • If the statistics system gets slow, redirects do not.

Step 7: scaling

  • App servers are stateless and scale out behind a load balancer.
  • The database is partitioned by the short key. Load spreads evenly.
  • The cache has several nodes.

Follow-up question: Someone uses the service to spread phishing links. What do you add to the design?

Follow-up answer:

  • Rate limits on creating links, per user and per IP.
  • Check the target URL against lists of known malicious URLs when the link is created.
  • A way to disable a link. This is one more reason for 302, because a 301 is cached in the user’s browser and is no longer under our control.

Red flag: Says nothing about the numbers, does not see key collisions, or does not know the difference between 301 and 302 for statistics.

Edit on GitHub

28Design a notification system (email, SMS, push)DesignWeight ×2System designKafkaIdempotency

Question: Design a central notification system that other services in the company use. Requirements:

  • Three channels: email, SMS and mobile push.
  • 50 million messages per day. A large part of them are marketing campaigns sent all at once.
  • Some messages are urgent, like one-time login codes (OTP).
  • Users choose which messages they get on which channel.
  • If sending fails, try again.
  • Avoid sending duplicates where possible.
  • External providers (email and SMS senders) have rate limits.

Short answer: We do not send messages inline. An API takes the message and puts it in a queue. Then:

  1. There is one queue per channel, so a slow channel does not affect the others.
  2. Urgent and marketing messages have separate queues, so an OTP does not wait behind a big campaign.
  3. Each message has an idempotency key, so duplicates are dropped.
  4. Each channel’s workers send at the speed the provider allows, and retry failures with growing delays.

Step 1: numbers

  1. 50 million per day ÷ 100,000 seconds = about 500 messages per second on average.
  2. But campaigns are bursty. Several million messages may arrive in a few minutes.
  3. So the system must absorb bursts in a queue and send at a steady pace. This is the main reason for the queue.

Step 2: main parts

  1. API service: takes the request (user, message type, data, idempotency key). It only validates and puts it in the input queue. It answers quickly with 202.
  2. Router: reads the user’s preferences. For example, the user does not want marketing SMS. Then it fills in the template and puts one message in the queue of each needed channel.
  3. Queues: for example with Kafka or RabbitMQ. For each channel, one urgent queue and one normal queue.
  4. Channel workers: read from the queue and send to the external provider.
  5. Status table: for each message: queued, sent, failed, delivered (if the provider reports it).

Step 3: data model

  • User preferences: user, message type, channel, on or off, quiet hours.
  • Message: id, idempotency key (unique), user, channel, priority, status, attempt count, timestamps.
  • Devices: the push token of each of the user’s devices.

Step 4: the hard parts

A) Duplicates:

  1. Queues usually deliver “at least once”. So a message can reach a worker twice.
  2. Before sending, the worker checks the idempotency key in the status table. If it is already “sent”, it does nothing.
  3. One hard case remains: we sent it, the provider delivered it, but its answer never reached us (timeout). We do not know if the message went out.
  4. If the provider accepts an idempotency key, we send the same key again. If not, we must choose: resend (maybe a duplicate) or not (maybe lost). For an OTP, resending is usually better. That is why we say “where possible”.

B) Provider rate limits:

  • Workers use a shared token bucket (for example in Redis) to keep the send rate below the provider’s limit.
  • If the provider returns a “too many requests” code (like 429), the worker waits a bit and slows down.

C) Retries:

  • Temporary errors (timeout, 500, 429) are retried with exponential backoff and some randomness (jitter).
  • Permanent errors (invalid number, expired push token) are not retried. We delete the expired token.
  • After a few attempts, the message goes to a dead letter queue (DLQ) for review.

D) Priority: the urgent queue has its own workers and always has free capacity. A 5-million-message campaign must not delay an OTP by minutes.

Step 5: scaling

  • Each part is stateless and scales on its own. Is the email queue busy? Add email workers.
  • Have two providers per channel. If one goes down, switch to the other.

Follow-up question: The marketing team wants to send a campaign to 10 million users “right now”. What happens, and how do you handle it?

Follow-up answer:

  • If the provider’s limit is, say, a few hundred messages per second, 10 million messages take several hours. This is math, and code cannot change it. Say this clearly to the marketing team.
  • The campaign goes to the normal queue and is sent at a steady pace. The urgent queue is not touched.
  • Respect users’ quiet hours. Messages for users where it is night now are scheduled for the morning.

Red flag: Puts all channels and all priorities in one queue, or claims that “exactly once” with an external provider is always possible.

Edit on GitHub

29Design a simple chat appDesignWeight ×2System design

Question: Design a simple chat app (like a simple WhatsApp). Requirements:

  • 1-to-1 chats and small groups (up to about 100 people).
  • 20 million daily active users.
  • Online status (online / last seen).
  • Messages in each conversation must show in the right order.
  • Delivery and read receipts (like the check marks).
  • If the recipient is offline, the message reaches them later.

Short answer:

  1. Each online user keeps one open WebSocket connection to one of the connection servers (gateways).
  2. A registry (for example in Redis) says which gateway each user is connected to right now.
  3. A message is stored first, then sent to the recipient. So being offline is not a problem.
  4. The server decides the order with a sequence number per conversation, not the phone’s clock.

Step 1: numbers (with assumptions)

  1. Assumption: each user sends 40 messages per day. So 20 million × 40 = 800 million messages per day.
  2. Divided by 100,000 seconds: about 8,000 messages per second. At peak maybe 3 times that.
  3. Assumption: at peak, about a quarter of users are online. That is about 5 million concurrent connections.
  4. If each gateway holds about 50,000 connections (we must measure this with a load test), we need about 100 connection servers.

Step 2: main parts

  1. Connection servers (gateways): hold the WebSocket connections. No business logic.
  2. Connection registry: which gateway each user is on.
  3. Chat service: receives a message, assigns a sequence number, stores it and hands it off for delivery.
  4. Message store: keeps messages grouped by conversation.
  5. Push service: sends a push notification to users who are offline.

Step 3: data model

  • Message: conversation id, sequence number, message id (created by the sender’s phone), sender, text, time.
  • The partition key is the conversation id, and inside it messages are sorted by sequence number. This fits a wide-column database (like Cassandra), because we almost always read “the latest messages of one conversation”.
  • Membership: conversation, user, last delivered sequence, last read sequence.

Step 4: the flow of one message (step by step)

  1. Ali’s phone sends the message with a unique id to the gateway.
  2. The chat service gives it the next sequence number of that conversation and stores it.
  3. It answers Ali with “stored” (one check mark).
  4. It asks the registry where Sara is connected and sends the message to her gateway.
  5. Sara’s phone confirms receipt. The server updates the last delivered sequence and tells Ali (two check marks).
  6. When Sara opens the chat, a “read” receipt comes back the same way.

Step 5: the hard parts

A) Message order: phone clocks cannot be trusted. The server assigns the sequence number per conversation. If all messages of one conversation are handled by one partition or one service instance, generating the number is simple. We do not need order across different conversations.

B) Duplicates: if the connection drops, the phone sends the message again. Because the phone created the message id, the server detects the duplicate and does not store it twice.

C) Offline messages: after reconnecting, the phone says “for each conversation, this is the last sequence I have”. The server sends everything after that. So no separate offline queue is needed. The message store itself is the source of truth.

D) Group fan-out: in a group of 100, a message is stored once and sent to 100 members. For small groups this is fine. Channels with millions of members need a different design, which is outside these requirements.

E) Online status: the phone sends a heartbeat every few seconds. The server keeps “last seen” in Redis with an expiry time. We do not broadcast status changes to all contacts, because that creates millions of useless messages. Only when someone has a chat screen open do they get the other person’s status.

Step 6: scaling

  • Connection servers scale out. If one dies, phones reconnect to another server and fetch missed messages using the sequence number.
  • The message store is partitioned by conversation id.

Follow-up question: A user is logged in on two devices (phone and laptop). What changes in the design?

Follow-up answer:

  • Instead of “user to one gateway”, the registry must keep “user to several devices, each on a gateway”.
  • A message is sent to all of the recipient’s devices, and also to the sender’s other devices, so all stay in sync.
  • Each device has its own last sequence number. So sync works per device.

Red flag: Uses the phone’s clock for ordering, or sends a message only when the recipient is online and does not store it.

Edit on GitHub

30Design the payment part of an online shopDesignWeight ×2System designIdempotencyThe Saga pattern

Question: Design the payment part of an online shop. The payment itself is done by an external payment provider. We call its API, and it also tells us the result with a webhook. Requirements:

  • About 100,000 orders per day.
  • Never charge a customer twice.
  • The provider is sometimes slow or times out.
  • Webhooks may arrive duplicated, late, or in the wrong order.
  • At the end of each day, our money must match the provider’s report.

Short answer: The traffic is small. The hard part of this system is correctness, not scale. Five main tools:

  1. An idempotency key for each payment attempt, also sent to the provider.
  2. A clear state machine for the payment, with an “unknown” state.
  3. An idempotent webhook handler that checks the signature.
  4. The Outbox pattern to tell the rest of the system.
  5. A daily reconciliation job and a simple ledger.

Step 1: the payment state machine

State Meaning
CREATED The payment record exists, not yet sent to the provider
PENDING The request was sent, waiting for the result
UNKNOWN No answer came (timeout); we do not know if money was taken
SUCCEEDED The money was taken
FAILED The money was not taken
  • Only forward moves are allowed. For example, you cannot go from SUCCEEDED back to PENDING.
  • This rule also solves out-of-order webhooks: an older webhook cannot move the state backward.

Step 2: preventing double charges (step by step)

  1. When the customer clicks Pay, we create a payment record with a unique idempotency key. The order has a unique constraint so it can have only one active payment.
  2. We send the request to the provider with this key. Many providers accept an idempotency key and do not run a repeated request with the same key again.
  3. If it times out, we set the state to UNKNOWN. We do not immediately retry with a new key. That is exactly how you charge twice.
  4. For UNKNOWN: we either resend with the same key, ask the provider’s API for the status, or wait for the webhook.
  5. We tell the customer honestly “your payment is being checked”, not “failed”, so they do not pay again.

Step 3: webhooks

  1. First, verify the webhook signature. Anyone can send a request to our URL.
  2. Store the event id in a table. If we have seen it before, just answer success and do nothing.
  3. Move the payment state forward according to the state machine.
  4. Answer 200 quickly. Do heavy work later, asynchronously. Otherwise the provider thinks it failed and sends it again.

Step 4: the Outbox pattern

  1. When a payment succeeds, the order must become “paid”, and stock and shipping must be told.
  2. If we update the database first and then send a message, the service may die in between, and the message is never sent.
  3. So in one transaction, we change the payment state and also write a row to an outbox table.
  4. A separate process reads the outbox rows and publishes them. Receivers must be idempotent, because a message may arrive twice.

Step 5: the ledger

  • Every money movement is a new row. No row is ever edited or deleted. A refund is a new row, not the deletion of the old one.
  • In double-entry bookkeeping, each transaction has at least two rows that sum to zero. For example, the “customer” account is debited and the “sales” account is credited.
  • Store amounts as integers in the smallest currency unit (for example, cents), together with the currency, not as floating point numbers, to avoid rounding errors.

Step 6: daily reconciliation

  1. Every day, fetch the transaction report from the provider.
  2. Compare it row by row with our own ledger.
  3. Every mismatch (money taken but FAILED on our side, or the other way around) creates a case to review.
  4. Old UNKNOWN payments are also resolved here.
  5. This is the last safety net: even if a webhook is lost, this job finds the mistake.

Follow-up question: Reconciliation shows that a customer was charged, but their order is “failed” in our system, and the customer paid again. What do you do?

Follow-up answer:

  • First, the customer: refund the extra payment, with a new row in the ledger.
  • Then find the cause: probably a timeout was wrongly recorded as FAILED instead of UNKNOWN, and the customer paid again with a new key.
  • Fix the code so a timeout always goes to UNKNOWN, and add an alert for old UNKNOWN payments.

Red flag: Retries the payment with a new request after a timeout, accepts webhooks without checking the signature or removing duplicates, or stores money as a floating point number.

Edit on GitHub