Metrics and alerts
A metric is a number that we measure over time, like orders per minute or response time. Prometheus collects these numbers and Grafana shows them. A good alert is set on user pain, waits a little, and only wakes someone when there is really something to do.
Author: bezzad
The problem: the customer finds out first, then us
It is Friday night and the shop is busy. The bank is slow, and half of the payments fail. Our team finds out one hour later, and only from angry customer messages.
We had logs and traces. So what was the problem?
- Logs and traces are for the details of one event. To answer “how is the whole system doing right now?”, we must count millions of lines. This is slow and expensive.
- Nobody was looking at the dashboard. On Friday night, nobody sits at the system.
So we need two things: summary numbers about the health of the system, and an alert that tells us by itself.
What is a metric?
A metric is a number that we measure over time. For example, orders per minute, the percent of failed payments, or API response time. Each metric can also have a few labels, like the payment method or the HTTP status code.
| Log | Trace | Metric | |
|---|---|---|---|
| Which question does it answer? | What exactly happened? | Where was the time spent? | How is the whole system doing? |
| Unit | One event | One request | One summary number per time window |
| Cost when traffic grows | Grows | Grows, unless you use sampling | Stays almost the same |
| Good for alerts | Low | Low | Very good |
The cost of a metric stays the same, because the app does not store each request on its own. It only adds one to a counter. That is why metrics are the best base for dashboards and alerts.
Three main types of metric
- Counter. It only goes up. Like the total number of orders. The number itself is not important. How fast it changes is important, for example orders per second.
- Gauge. A value at one moment that goes up and down. Like the number of messages in a queue, or memory use.
- Histogram. It counts values in a few buckets. For example, how many requests were under 100 milliseconds, how many under 500 milliseconds, and so on. With a histogram you can calculate percentiles.
The average lies
Imagine 100 payment requests. 95 of them are fast and 5 take a few seconds. The average response time looks good. But five customers out of every hundred waited a few seconds.
The 99th percentile (p99) means 99 percent of requests were faster than this number. For response time we usually look at the 50th, 95th and 99th percentiles. In a high-traffic system, one percent means thousands of customers a day.
Code: metrics in .NET
ASP.NET Core and HttpClient have ready-made metrics from .NET 8 onward. For example, the http.server.request.duration metric is a histogram of the response time of all requests. We only need to collect them and expose them.
For business numbers, we build our own metric with the Meter class. The IMeterFactory interface gives it to us from DI:
public sealed class ShopMetrics
{
public const string MeterName = "Shop.Orders";
private readonly Counter<long> _ordersPlaced;
private readonly Histogram<double> _paymentDuration;
public ShopMetrics(IMeterFactory meterFactory)
{
var meter = meterFactory.Create(MeterName);
_ordersPlaced = meter.CreateCounter<long>("shop.orders.placed");
_paymentDuration = meter.CreateHistogram<double>("shop.payment.duration", unit: "s");
}
public void OrderPlaced(string paymentMethod) =>
_ordersPlaced.Add(1, new KeyValuePair<string, object?>("payment.method", paymentMethod));
public void PaymentFinished(TimeSpan elapsed) =>
_paymentDuration.Record(elapsed.TotalSeconds);
}
Now, with OpenTelemetry, we expose all metrics at the metrics address for Prometheus:
builder.Services.AddSingleton<ShopMetrics>();
builder.Services.AddOpenTelemetry()
.ConfigureResource(r => r.AddService("orders-api"))
.WithMetrics(metrics => metrics
.AddAspNetCoreInstrumentation()
.AddHttpClientInstrumentation()
.AddMeter(ShopMetrics.MeterName)
.AddPrometheusExporter());
var app = builder.Build();
app.MapPrometheusScrapingEndpoint(); // exposes /metrics
Data flow: Prometheus and Grafana
- The app shows its current numbers at the metrics address.
- Prometheus reads this address from all instances at fixed intervals (Scrape). It stores the numbers with their time.
- Grafana asks Prometheus questions in the PromQL language and draws charts.
- Prometheus checks the alert rules all the time. If a rule becomes true, it gives it to Alertmanager. Alertmanager groups similar alerts and sends them to the right person.
Here are two common questions in PromQL. Metric names change a little in Prometheus: a dot becomes an underscore, and the unit is added to the end of the name.
# Requests per second that ended with a 5xx status, over the last 5 minutes
sum(rate(http_server_request_duration_seconds_count{http_response_status_code=~"5.."}[5m]))
# p99 response time of the checkout endpoint
histogram_quantile(0.99,
sum by (le) (rate(http_server_request_duration_seconds_bucket{http_route="/checkout"}[5m])))
What should we measure?
For each service, take three main numbers. This method is called RED:
- Rate. How many requests come in each second?
- Errors. What percent of them fail?
- Duration. What are the response time percentiles?
For resources, like CPU, memory, queues and the Connection Pool, how full they are (Saturation) also matters. For example, if the database Connection Pool is always full, requests will soon start to wait.
Next to these, do not forget business numbers: orders per minute and the percent of successful payments. Sometimes all technical numbers are healthy, but no order is placed.
A good alert
An alert means waking up a person. So each alert must be worth it:
- On user pain, not on the cause. “The payment error rate is high” is user pain. “CPU is at eighty percent” may cause no pain at all.
- Wait a little. A spike of a few seconds is not worth waking someone up. In Prometheus, the for section says the condition must stay true for some time without a break.
- Make it actionable. The person who gets the alert must know what to do. Put a link to a guide for fixing the problem (Runbook) next to the alert.
- Give it a level. Only an urgent problem should wake a person. The rest can be a ticket or a message in the team chat.
See what the for section does in practice:
Rule: if the error rate is above 5% for five minutes in a row, tell the on-call person. Each step of the chart is one minute.
The same rule looks like this in Prometheus:
groups:
- name: payments
rules:
- alert: PaymentErrorRateHigh
expr: |
sum(rate(http_server_request_duration_seconds_count{http_route="/pay", http_response_status_code=~"5.."}[5m]))
/
sum(rate(http_server_request_duration_seconds_count{http_route="/pay"}[5m]))
> 0.05
for: 5m
labels:
severity: page
annotations:
summary: "More than 5% of payments fail"
runbook_url: "https://wiki.example.com/runbooks/payment-errors"
Common mistakes
| Mistake | Result | Right way |
|---|---|---|
| Looking at the average response time | Slowness for some customers stays hidden. | The 95th and 99th percentiles with a histogram. |
| User or order ID as a label | The number of series explodes and Prometheus stops working. | Labels with limited values. |
| Alerts on CPU and memory | Many alerts with no real pain. | Alerts on errors and slowness that the user feels. |
| Alert with no waiting time | Every small spike wakes someone up. | A for section of a few minutes. |
| Many alerts that nobody acts on | The team gets used to alerts and ignores the real one too. | Remove or fix every alert with no action. |
| Only technical metrics | The system looks healthy, but no order is placed. | Have business metrics too. |
Summary in six lines
- A metric is a number over time, and its cost does not grow with traffic.
- There are three main types: counter, gauge and histogram.
- For response time, look at percentiles, not the average.
- For each service, take the request rate, the error percent and the response time, next to business numbers.
- Labels must have limited values. IDs belong in logs and traces.
- An alert must be on user pain, wait a little, and always have a clear action.