Levelwise
English
Observability

Metrics and alerts

A metric is a number that we measure over time, like orders per minute or response time. Prometheus collects these numbers and Grafana shows them. A good alert is set on user pain, waits a little, and only wakes someone when there is really something to do.

Not reviewedWritten with AI helpReading time: 15 minOnline shop payment exampleC# code with OpenTelemetry and Prometheus

Author: bezzad

The problem: the customer finds out first, then us

It is Friday night and the shop is busy. The bank is slow, and half of the payments fail. Our team finds out one hour later, and only from angry customer messages.

We had logs and traces. So what was the problem?

  1. Logs and traces are for the details of one event. To answer “how is the whole system doing right now?”, we must count millions of lines. This is slow and expensive.
  2. Nobody was looking at the dashboard. On Friday night, nobody sits at the system.

So we need two things: summary numbers about the health of the system, and an alert that tells us by itself.

What is a metric?

A metric is a number that we measure over time. For example, orders per minute, the percent of failed payments, or API response time. Each metric can also have a few labels, like the payment method or the HTTP status code.

Log Trace Metric
Which question does it answer? What exactly happened? Where was the time spent? How is the whole system doing?
Unit One event One request One summary number per time window
Cost when traffic grows Grows Grows, unless you use sampling Stays almost the same
Good for alerts Low Low Very good

The cost of a metric stays the same, because the app does not store each request on its own. It only adds one to a counter. That is why metrics are the best base for dashboards and alerts.

Three main types of metric

CounterOnly goes upOrder count, error countGaugeGoes up and downQueue length, memory useHistogramSpread of valuesResponse time, cart sizeEach metric is a number over time, with a few labels
  1. Counter. It only goes up. Like the total number of orders. The number itself is not important. How fast it changes is important, for example orders per second.
  2. Gauge. A value at one moment that goes up and down. Like the number of messages in a queue, or memory use.
  3. Histogram. It counts values in a few buckets. For example, how many requests were under 100 milliseconds, how many under 500 milliseconds, and so on. With a histogram you can calculate percentiles.

The average lies

Imagine 100 payment requests. 95 of them are fast and 5 take a few seconds. The average response time looks good. But five customers out of every hundred waited a few seconds.

Response time of 100 payment requestsAverage: looks goodp99: some customers wait seconds95 fast requests5 slow requestsThe average hides the pain of these 5 customers
For response time, look at percentiles, not the average.

The 99th percentile (p99) means 99 percent of requests were faster than this number. For response time we usually look at the 50th, 95th and 99th percentiles. In a high-traffic system, one percent means thousands of customers a day.

Code: metrics in .NET

ASP.NET Core and HttpClient have ready-made metrics from .NET 8 onward. For example, the http.server.request.duration metric is a histogram of the response time of all requests. We only need to collect them and expose them.

For business numbers, we build our own metric with the Meter class. The IMeterFactory interface gives it to us from DI:

public sealed class ShopMetrics
{
    public const string MeterName = "Shop.Orders";

    private readonly Counter<long> _ordersPlaced;
    private readonly Histogram<double> _paymentDuration;

    public ShopMetrics(IMeterFactory meterFactory)
    {
        var meter = meterFactory.Create(MeterName);
        _ordersPlaced = meter.CreateCounter<long>("shop.orders.placed");
        _paymentDuration = meter.CreateHistogram<double>("shop.payment.duration", unit: "s");
    }

    public void OrderPlaced(string paymentMethod) =>
        _ordersPlaced.Add(1, new KeyValuePair<string, object?>("payment.method", paymentMethod));

    public void PaymentFinished(TimeSpan elapsed) =>
        _paymentDuration.Record(elapsed.TotalSeconds);
}

Now, with OpenTelemetry, we expose all metrics at the metrics address for Prometheus:

builder.Services.AddSingleton<ShopMetrics>();

builder.Services.AddOpenTelemetry()
    .ConfigureResource(r => r.AddService("orders-api"))
    .WithMetrics(metrics => metrics
        .AddAspNetCoreInstrumentation()
        .AddHttpClientInstrumentation()
        .AddMeter(ShopMetrics.MeterName)
        .AddPrometheusExporter());

var app = builder.Build();
app.MapPrometheusScrapingEndpoint();   // exposes /metrics
Pre-release version: The OpenTelemetry.Exporter.Prometheus.AspNetCore package has no stable version yet. It is published as a beta. Another way is to send the metrics with OTLP to the OpenTelemetry Collector, and let the Collector give them to Prometheus.

Data flow: Prometheus and Grafana

Order service/metrics3 instancesPrometheusStores numbers over timeChecks alert rulesGrafanaDashboards and chartsTo see trendsAlertmanagerGroups and sends alertsOn-call personSMS, call, team chatScrapeevery 15 sQueryPromQL
The app only keeps the numbers ready. Prometheus comes to get them by itself.
  1. The app shows its current numbers at the metrics address.
  2. Prometheus reads this address from all instances at fixed intervals (Scrape). It stores the numbers with their time.
  3. Grafana asks Prometheus questions in the PromQL language and draws charts.
  4. Prometheus checks the alert rules all the time. If a rule becomes true, it gives it to Alertmanager. Alertmanager groups similar alerts and sends them to the right person.

Here are two common questions in PromQL. Metric names change a little in Prometheus: a dot becomes an underscore, and the unit is added to the end of the name.

# Requests per second that ended with a 5xx status, over the last 5 minutes
sum(rate(http_server_request_duration_seconds_count{http_response_status_code=~"5.."}[5m]))

# p99 response time of the checkout endpoint
histogram_quantile(0.99,
  sum by (le) (rate(http_server_request_duration_seconds_bucket{http_route="/checkout"}[5m])))

What should we measure?

For each service, take three main numbers. This method is called RED:

  1. Rate. How many requests come in each second?
  2. Errors. What percent of them fail?
  3. Duration. What are the response time percentiles?

For resources, like CPU, memory, queues and the Connection Pool, how full they are (Saturation) also matters. For example, if the database Connection Pool is always full, requests will soon start to wait.

Next to these, do not forget business numbers: orders per minute and the percent of successful payments. Sometimes all technical numbers are healthy, but no order is placed.

Be careful with labels. Each new combination of label values creates a new time series in Prometheus. If you use the user ID or the order number as a label, millions of series are created, and Prometheus gets slow or stops working. A label must have a few limited values, like the payment method. IDs belong in logs and traces.

A good alert

An alert means waking up a person. So each alert must be worth it:

  1. On user pain, not on the cause. “The payment error rate is high” is user pain. “CPU is at eighty percent” may cause no pain at all.
  2. Wait a little. A spike of a few seconds is not worth waking someone up. In Prometheus, the for section says the condition must stay true for some time without a break.
  3. Make it actionable. The person who gets the alert must know what to do. Put a link to a guide for fixing the problem (Runbook) next to the alert.
  4. Give it a level. Only an urgent problem should wake a person. The rest can be a ticket or a message in the team chat.

See what the for section does in practice:

Live example: payment error rate alert

Rule: if the error rate is above 5% for five minutes in a row, tell the on-call person. Each step of the chart is one minute.

5%20%0%

    The same rule looks like this in Prometheus:

    groups:
      - name: payments
        rules:
          - alert: PaymentErrorRateHigh
            expr: |
              sum(rate(http_server_request_duration_seconds_count{http_route="/pay", http_response_status_code=~"5.."}[5m]))
                /
              sum(rate(http_server_request_duration_seconds_count{http_route="/pay"}[5m]))
                > 0.05
            for: 5m
            labels:
              severity: page
            annotations:
              summary: "More than 5% of payments fail"
              runbook_url: "https://wiki.example.com/runbooks/payment-errors"
    Service level objective (SLO): More mature teams first write a goal. For example: “99.5 percent of payments in a month must succeed.” Then they set the alert on this question: with the current error speed, will we miss the goal for the month? This method cuts unneeded alerts a lot.

    Common mistakes

    Mistake Result Right way
    Looking at the average response time Slowness for some customers stays hidden. The 95th and 99th percentiles with a histogram.
    User or order ID as a label The number of series explodes and Prometheus stops working. Labels with limited values.
    Alerts on CPU and memory Many alerts with no real pain. Alerts on errors and slowness that the user feels.
    Alert with no waiting time Every small spike wakes someone up. A for section of a few minutes.
    Many alerts that nobody acts on The team gets used to alerts and ignores the real one too. Remove or fix every alert with no action.
    Only technical metrics The system looks healthy, but no order is placed. Have business metrics too.

    Summary in six lines

    1. A metric is a number over time, and its cost does not grow with traffic.
    2. There are three main types: counter, gauge and histogram.
    3. For response time, look at percentiles, not the average.
    4. For each service, take the request rate, the error percent and the response time, next to business numbers.
    5. Labels must have limited values. IDs belong in logs and traces.
    6. An alert must be on user pain, wait a little, and always have a clear action.