Levelwise
English
Infrastructure and DevOps

Kubernetes for .NET Developers

In Kubernetes you say "I want three copies of the order service", and it runs them and keeps them healthy. For this to work well, the app must ask for the right resources, report its health and shut down safely.

Not reviewedWritten with AI helpReading time: 17 minOnline shop order service exampleYAML and C# code with .NET 10

Author: bezzad

The problem: we have a container, but who looks after it?

Our shop’s order service now has an Image. On busy days, three copies of it must run on several servers. There are many questions:

  1. If one copy crashes, who starts it again?
  2. If a server goes down, where do the copies on it go?
  3. How do we deploy a new version so users see no errors?
  4. At busy hours, who creates more copies?

You cannot do all this by hand. Kubernetes exists for this.

The idea: write the desired state

In Kubernetes we do not give orders like “do this”. We only write the desired state in a YAML file. For example, “three copies of the order Image, version 1.4”. Then Kubernetes keeps comparing the real state with the desired state and fixes the difference.

A few main ideas:

  • Pod. The smallest unit that runs. It usually has one app container. Every Pod is temporary. It may die at any moment, and a new Pod with a new address takes its place.
  • Deployment. It says how many copies of a Pod we want, and how to replace them with a new version.
  • Service. It gives us a fixed address and name. It sends traffic only to Pods that are ready.
Shop usersService: orders-apiOne fixed address for all copiesPod 1ReadyPod 2ReadyPod 3Not ready yetDeployment: replicas = 3If a copy dies, it creates a new one
The Service sends traffic only to Pods that are ready. The Deployment keeps the number of Pods fixed.
What this means for your code: Because a Pod is temporary, the app must be stateless. Do not keep the customer’s shopping cart in the Pod’s memory. Put it in Redis or a database. Otherwise the customer’s cart is lost on every deploy or scale.

Resources: request and limit

For each container we write two numbers:

  1. The request is a reservation. The scheduler places the Pod only on a server that has this amount free.
  2. The limit is a ceiling. The container cannot use more than this.

But CPU and memory behave very differently at the ceiling:

CPUMemoryReservedrequestlimitCan use more, up to limitHit the ceiling: gets slowReserved = limitlimitrequestOver the ceiling: killed
Hitting the CPU ceiling makes the app slow. Going over the memory ceiling kills the Pod.
  • When CPU hits the ceiling. The container is not killed. It only gets slow. This is called throttling. Response time goes up, but you see no error in the logs.
  • When memory goes over the ceiling. The container is killed at once, and its status is OOMKilled. Then it starts again.

Two points special to .NET:

  1. The number of cores comes from the CPU limit. If the CPU limit is one core, the app thinks it has only one processor. The thread pool and the GC are also tuned based on this.
  2. The GC sees the memory limit. By default it uses about 75% of the container’s memory limit for the heap. The rest stays for other things, like stacks and native memory.
If you do not set a request: The scheduler does not know how many resources the Pod needs. Many Pods may pile up on one server. At busy times they all get slow together, or get evicted.

Health: three kinds of Probe

How does Kubernetes know a Pod is healthy? From three kinds of check (Probe):

  1. The readiness probe. “Can you take traffic right now?” If the answer is no, the Service sends no traffic to this Pod. The Pod is not killed.
  2. The liveness probe. “Are you alive, or are you stuck?” If the answer is no, Kubernetes restarts the container.
  3. The startup probe. “Have you finished starting up?” Until the answer is yes, the other two probes do not start. It is useful for an app that starts slowly.
A dangerous mistake: The liveness probe must not check the database. Imagine the database gets slow for a few seconds. All Pods fail liveness together and all restart together. A small database problem turns into a full outage of the service. Check dependencies only in readiness, and even then with care.

Autoscaling with HPA

The HPA (Horizontal Pod Autoscaler) moves the number of Pods up and down based on a metric. For example, “if the average CPU goes above 70%, add a Pod”.

  1. The CPU percent is measured against the request. Not against the limit. If the request is half a core, 70% means about 0.35 of a core.
  2. The metric must show the real bottleneck. If the service waits for the database, CPU stays low and the HPA does nothing. More Pods would only put more load on the database.
  3. A queue consumer needs a different metric. For a service that reads from Kafka, the number of unread messages (lag) is a better metric. The KEDA tool does this.
  4. The useful maximum of Pods for a Kafka consumer is limited. In a consumer group, each partition goes to only one consumer. So if the topic has six partitions, the seventh Pod stays idle.
  5. Autoscaling reacts late. Creating a Pod and warming up the app takes time. For busy times you can predict, raise the minimum number of Pods in advance.

Safe shutdown

Every deploy means the old Pods shut down. If this is not done right, users see a 502 error. Let’s see, step by step, how Kubernetes shuts down a Pod:

  1. Two things start at the same time. Removing the Pod from the Service’s list of targets, and sending a shutdown signal to the app.
  2. Removing it from the list takes a few seconds. This change must reach all servers and Load Balancers.
  3. The SIGTERM signal arrives at once. With this signal, an ASP.NET Core app stops accepting new requests.
  4. So we have a dangerous gap. The app is closed, but the Load Balancer still sends it requests for a few seconds.
  5. At the end, SIGKILL comes. If the app is still running after the grace period (30 seconds by default), it is killed at once, even in the middle of work.

The fix is to wait a few seconds before SIGTERM (with preStop). This way, traffic stops first, and then the app closes.

Time from the shutdown command0s5s45spreStop sleeponly waitsRemoved from targetstakes a few secondsSIGTERMFinish current requests and messagesSIGKILLterminationGracePeriodSeconds = 45The total grace period starts at the first momentWait time plus time to finish work must be less than the total grace period
The preStop step only waits until removal from the Service is complete. Then SIGTERM arrives.

Live example

Choose a mode and see which requests get errors:

A Pod shutting down during a deploy
In Service list: yes
App: takes requests
Second 0
Choose a mode. Each square is one user request to this Pod.

    Code

    The Deployment file

    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: orders-api
    spec:
      replicas: 3
      selector:
        matchLabels: { app: orders-api }
      strategy:
        rollingUpdate:
          maxUnavailable: 0   # never go below 3 ready pods
          maxSurge: 1         # add one new pod at a time
      template:
        metadata:
          labels: { app: orders-api }
        spec:
          terminationGracePeriodSeconds: 45
          containers:
            - name: api
              image: registry.example.com/shop/orders-api:1.4.0
              ports: [{ containerPort: 8080 }]
              resources:
                requests: { cpu: 500m, memory: 512Mi }
                limits: { memory: 512Mi }
              readinessProbe:
                httpGet: { path: /health/ready, port: 8080 }
                periodSeconds: 5
              livenessProbe:
                httpGet: { path: /health/live, port: 8080 }
                periodSeconds: 10
              lifecycle:
                preStop:
                  sleep: { seconds: 5 }

    The timing math in this file is simple: five seconds of preStop plus thirty seconds to finish work is less than the 45-second total grace period.

    About preStop: The built-in sleep action exists in newer versions of Kubernetes. In older versions you had to run the sleep program with an exec command. But chiseled Images have no shell and no sleep program at all. Check your own Kubernetes version.

    The HPA file

    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    metadata:
      name: orders-api
    spec:
      scaleTargetRef:
        apiVersion: apps/v1
        kind: Deployment
        name: orders-api
      minReplicas: 3
      maxReplicas: 10
      metrics:
        - type: Resource
          resource:
            name: cpu
            target:
              type: Utilization
              averageUtilization: 70   # percent of the CPU request

    The .NET app side

    using Microsoft.AspNetCore.Diagnostics.HealthChecks;
    using Microsoft.Extensions.Diagnostics.HealthChecks;
    
    var builder = WebApplication.CreateBuilder(args);
    
    // How long to wait for running requests after SIGTERM.
    builder.Services.Configure<HostOptions>(o => o.ShutdownTimeout = TimeSpan.FromSeconds(30));
    
    builder.Services.AddHealthChecks()
        // Liveness: only "is the process alive?"
        .AddCheck("self", () => HealthCheckResult.Healthy(), tags: ["live"])
        // Readiness: can we serve traffic? (needs the EF Core health checks package)
        .AddDbContextCheck<OrdersDbContext>(tags: ["ready"]);
    
    var app = builder.Build();
    
    app.MapHealthChecks("/health/live", new HealthCheckOptions
    {
        Predicate = check => check.Tags.Contains("live")
    });
    app.MapHealthChecks("/health/ready", new HealthCheckOptions
    {
        Predicate = check => check.Tags.Contains("ready")
    });
    
    app.Run();

    In background work (BackgroundService), also pass the stopping token to every method. This way no new work starts after SIGTERM. A Kafka consumer must also finish the current message first, commit the offset, and then close.

    Key rules

    1. The app must be stateless. Any Pod may die at any moment.
    2. Write a request for every container. Base it on real usage, not a guess.
    3. Set the memory request equal to the memory limit. The behavior is more predictable. Teams do not agree about a CPU limit, because a CPU limit causes throttling.
    4. The liveness probe checks only the app itself. Dependencies go in readiness.
    5. Design the shutdown. The preStop step, the ShutdownTimeout in the app and the Pod’s total grace period must fit together.
    6. Message processing must be idempotent. Even with all this, a Pod sometimes dies suddenly. For example, a server breaks. So a message may be processed twice.

    Common mistakes

    Mistake Result Right way
    Keeping the session or shopping cart in the Pod’s memory On every deploy or scale, user data is lost. Store it in Redis or a database.
    Checking the database in liveness A slow database restarts all Pods together. The liveness probe checks only the app itself.
    No preStop On every deploy, some requests get a 502 error. Wait a few seconds with preStop.
    Total grace period shorter than needed Messages and requests are cut by SIGKILL in the middle of work. preStop plus ShutdownTimeout less than the total grace period.
    Scaling a Kafka consumer by CPU Messages fall behind, but the HPA sees nothing. Scale by lag, for example with KEDA.
    Maximum Pods more than the number of partitions The extra Pods are idle and only cost money. Pod ceiling equal to the number of partitions.

    When to use Kubernetes?

    Good fit

    • Several services with several copies that must stay healthy and scale automatically.
    • A team or company that has someone to run the cluster, or uses a managed cloud service.
    • Several deploys a day with no downtime.

    Poor fit

    • A small app with one copy. One server or a simple cloud service is enough.
    • A small team with no Kubernetes experience and no time to learn it.

    Summary in seven lines

    1. In Kubernetes you write the desired state, and it keeps the real state correct.
    2. Every Pod is temporary, so the app must be stateless.
    3. The request is a reservation and the limit is a ceiling. The CPU ceiling slows the app; the memory ceiling kills it.
    4. The readiness probe stops traffic; the liveness probe restarts. Do not check the database in liveness.
    5. The HPA percent is measured against the request, and its metric must be the real bottleneck.
    6. For a safe shutdown: preStop, then SIGTERM and finishing the work, all inside the total grace period.
    7. Still, make message processing idempotent, because a Pod sometimes dies suddenly.