Kubernetes for .NET Developers
In Kubernetes you say "I want three copies of the order service", and it runs them and keeps them healthy. For this to work well, the app must ask for the right resources, report its health and shut down safely.
Author: bezzad
The problem: we have a container, but who looks after it?
Our shop’s order service now has an Image. On busy days, three copies of it must run on several servers. There are many questions:
- If one copy crashes, who starts it again?
- If a server goes down, where do the copies on it go?
- How do we deploy a new version so users see no errors?
- At busy hours, who creates more copies?
You cannot do all this by hand. Kubernetes exists for this.
The idea: write the desired state
In Kubernetes we do not give orders like “do this”. We only write the desired state in a YAML file. For example, “three copies of the order Image, version 1.4”. Then Kubernetes keeps comparing the real state with the desired state and fixes the difference.
A few main ideas:
- Pod. The smallest unit that runs. It usually has one app container. Every Pod is temporary. It may die at any moment, and a new Pod with a new address takes its place.
- Deployment. It says how many copies of a Pod we want, and how to replace them with a new version.
- Service. It gives us a fixed address and name. It sends traffic only to Pods that are ready.
Resources: request and limit
For each container we write two numbers:
- The request is a reservation. The scheduler places the Pod only on a server that has this amount free.
- The limit is a ceiling. The container cannot use more than this.
But CPU and memory behave very differently at the ceiling:
- When CPU hits the ceiling. The container is not killed. It only gets slow. This is called throttling. Response time goes up, but you see no error in the logs.
- When memory goes over the ceiling. The container is killed at once, and its status is OOMKilled. Then it starts again.
Two points special to .NET:
- The number of cores comes from the CPU limit. If the CPU limit is one core, the app thinks it has only one processor. The thread pool and the GC are also tuned based on this.
- The GC sees the memory limit. By default it uses about 75% of the container’s memory limit for the heap. The rest stays for other things, like stacks and native memory.
Health: three kinds of Probe
How does Kubernetes know a Pod is healthy? From three kinds of check (Probe):
- The readiness probe. “Can you take traffic right now?” If the answer is no, the Service sends no traffic to this Pod. The Pod is not killed.
- The liveness probe. “Are you alive, or are you stuck?” If the answer is no, Kubernetes restarts the container.
- The startup probe. “Have you finished starting up?” Until the answer is yes, the other two probes do not start. It is useful for an app that starts slowly.
Autoscaling with HPA
The HPA (Horizontal Pod Autoscaler) moves the number of Pods up and down based on a metric. For example, “if the average CPU goes above 70%, add a Pod”.
- The CPU percent is measured against the request. Not against the limit. If the request is half a core, 70% means about 0.35 of a core.
- The metric must show the real bottleneck. If the service waits for the database, CPU stays low and the HPA does nothing. More Pods would only put more load on the database.
- A queue consumer needs a different metric. For a service that reads from Kafka, the number of unread messages (lag) is a better metric. The KEDA tool does this.
- The useful maximum of Pods for a Kafka consumer is limited. In a consumer group, each partition goes to only one consumer. So if the topic has six partitions, the seventh Pod stays idle.
- Autoscaling reacts late. Creating a Pod and warming up the app takes time. For busy times you can predict, raise the minimum number of Pods in advance.
Safe shutdown
Every deploy means the old Pods shut down. If this is not done right, users see a 502 error. Let’s see, step by step, how Kubernetes shuts down a Pod:
- Two things start at the same time. Removing the Pod from the Service’s list of targets, and sending a shutdown signal to the app.
- Removing it from the list takes a few seconds. This change must reach all servers and Load Balancers.
- The SIGTERM signal arrives at once. With this signal, an ASP.NET Core app stops accepting new requests.
- So we have a dangerous gap. The app is closed, but the Load Balancer still sends it requests for a few seconds.
- At the end, SIGKILL comes. If the app is still running after the grace period (30 seconds by default), it is killed at once, even in the middle of work.
The fix is to wait a few seconds before SIGTERM (with preStop). This way, traffic stops first, and then the app closes.
Live example
Choose a mode and see which requests get errors:
Code
The Deployment file
apiVersion: apps/v1
kind: Deployment
metadata:
name: orders-api
spec:
replicas: 3
selector:
matchLabels: { app: orders-api }
strategy:
rollingUpdate:
maxUnavailable: 0 # never go below 3 ready pods
maxSurge: 1 # add one new pod at a time
template:
metadata:
labels: { app: orders-api }
spec:
terminationGracePeriodSeconds: 45
containers:
- name: api
image: registry.example.com/shop/orders-api:1.4.0
ports: [{ containerPort: 8080 }]
resources:
requests: { cpu: 500m, memory: 512Mi }
limits: { memory: 512Mi }
readinessProbe:
httpGet: { path: /health/ready, port: 8080 }
periodSeconds: 5
livenessProbe:
httpGet: { path: /health/live, port: 8080 }
periodSeconds: 10
lifecycle:
preStop:
sleep: { seconds: 5 }
The timing math in this file is simple: five seconds of preStop plus thirty seconds to finish work is less than the 45-second total grace period.
The HPA file
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: orders-api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: orders-api
minReplicas: 3
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70 # percent of the CPU request
The .NET app side
using Microsoft.AspNetCore.Diagnostics.HealthChecks;
using Microsoft.Extensions.Diagnostics.HealthChecks;
var builder = WebApplication.CreateBuilder(args);
// How long to wait for running requests after SIGTERM.
builder.Services.Configure<HostOptions>(o => o.ShutdownTimeout = TimeSpan.FromSeconds(30));
builder.Services.AddHealthChecks()
// Liveness: only "is the process alive?"
.AddCheck("self", () => HealthCheckResult.Healthy(), tags: ["live"])
// Readiness: can we serve traffic? (needs the EF Core health checks package)
.AddDbContextCheck<OrdersDbContext>(tags: ["ready"]);
var app = builder.Build();
app.MapHealthChecks("/health/live", new HealthCheckOptions
{
Predicate = check => check.Tags.Contains("live")
});
app.MapHealthChecks("/health/ready", new HealthCheckOptions
{
Predicate = check => check.Tags.Contains("ready")
});
app.Run();
In background work (BackgroundService), also pass the stopping token to every method. This way no new work starts after SIGTERM. A Kafka consumer must also finish the current message first, commit the offset, and then close.
Key rules
- The app must be stateless. Any Pod may die at any moment.
- Write a request for every container. Base it on real usage, not a guess.
- Set the memory request equal to the memory limit. The behavior is more predictable. Teams do not agree about a CPU limit, because a CPU limit causes throttling.
- The liveness probe checks only the app itself. Dependencies go in readiness.
- Design the shutdown. The preStop step, the ShutdownTimeout in the app and the Pod’s total grace period must fit together.
- Message processing must be idempotent. Even with all this, a Pod sometimes dies suddenly. For example, a server breaks. So a message may be processed twice.
Common mistakes
| Mistake | Result | Right way |
|---|---|---|
| Keeping the session or shopping cart in the Pod’s memory | On every deploy or scale, user data is lost. | Store it in Redis or a database. |
| Checking the database in liveness | A slow database restarts all Pods together. | The liveness probe checks only the app itself. |
| No preStop | On every deploy, some requests get a 502 error. | Wait a few seconds with preStop. |
| Total grace period shorter than needed | Messages and requests are cut by SIGKILL in the middle of work. | preStop plus ShutdownTimeout less than the total grace period. |
| Scaling a Kafka consumer by CPU | Messages fall behind, but the HPA sees nothing. | Scale by lag, for example with KEDA. |
| Maximum Pods more than the number of partitions | The extra Pods are idle and only cost money. | Pod ceiling equal to the number of partitions. |
When to use Kubernetes?
Good fit
- Several services with several copies that must stay healthy and scale automatically.
- A team or company that has someone to run the cluster, or uses a managed cloud service.
- Several deploys a day with no downtime.
Poor fit
- A small app with one copy. One server or a simple cloud service is enough.
- A small team with no Kubernetes experience and no time to learn it.
Summary in seven lines
- In Kubernetes you write the desired state, and it keeps the real state correct.
- Every Pod is temporary, so the app must be stateless.
- The request is a reservation and the limit is a ceiling. The CPU ceiling slows the app; the memory ceiling kills it.
- The readiness probe stops traffic; the liveness probe restarts. Do not check the database in liveness.
- The HPA percent is measured against the request, and its metric must be the real bottleneck.
- For a safe shutdown: preStop, then SIGTERM and finishing the work, all inside the total grace period.
- Still, make message processing idempotent, because a Pod sometimes dies suddenly.