Clouds
Observability·4 min read

Cardinality: Why Your Metrics Bill Exploded

A single label added to a counter can multiply your time series by ten thousand, and the invoice arrives weeks later pointing at code that looks perfectly reasonable.


The monitoring bill tripled between two releases. Nothing in the changelog mentioned monitoring. The release added one label (user_id) to a request counter, so the dashboard could show per-user traffic for a customer who had asked for it once.

One label. The series count went from thirty thousand to eighteen million.

What a label actually costs

A time series is identified by metric name plus every label value. Each distinct combination is a separate sequence the database stores, indexes, and samples on every scrape.

text
http_requests_total{method="GET", status="200", route="/api/orders"}   → 1 series
http_requests_total{method="GET", status="200", route="/api/orders",   → 1 series
                    user_id="8f2a..."}

The second line is not a variation of the first. It is a new row, a new index entry, and a new value written at every scrape interval for the lifetime of the deployment, multiplied by every route, every status, and every user who ever made a request.

The arithmetic that catches teams is that labels multiply rather than add. Three labels with 10, 5, and 1,000,000 values produce fifty million series. The cardinality of the highest-cardinality label dominates, and it is almost always an identifier.

The categories worth separating

Not all labels are equal, and the distinction is whether the value is bounded and known in advance.

Bounded labels are safe: an HTTP method, a status class, a region, a service name, a build version. You can enumerate them before the code ships, and the count grows only when you add a new route or deploy to a new region.

Unbounded labels are the problem: a user ID, an order ID, a request ID, a URL path with an ID in it, a dynamic hostname, a free-text error string. These grow with traffic, which is what you did not predict.

The clearest signal is whether you would ever write a WHERE clause against it to investigate an incident. If nobody would, it does not belong in a metric; it belongs in a log line or a trace, where the storage model is completely different.

Moving the detail where it belongs

The fix is not to lose the information. It is to stop paying per-sample cost for detail you will query one user at a time.

text
metric  →  "how many requests failed, and for whom at the aggregate level"
log     →  "which specific request, with which user, and what the error said"
trace   →  "which calls downstream contributed to this one slow request"

Metrics answer counts and rates over bounded groups. Logs answer questions about a specific event. Traces answer questions about a specific causal path. Putting a user ID in a metric answers a log question at metric prices, and it answers it poorly: you get a count per user rather than the actual failure.

Concretely: keep http_requests_total{route, method, status} for the alert, and attach user_id to the structured log emitted on a non-2xx response. The dashboard stays cheap, the investigation still has the field, and the cost moves to storage that scales with events rather than with time.

Catching it before the invoice

The check belongs in the build, not the postmortem. Prometheus exposes the damage:

promql
# series per metric, descending: run this after every deploy
count by (__name__) ({__name__=~".+"})

If http_requests_total appears near the top, the label set is growing fast. A stricter version is a CI assertion on the scrape endpoint: compare series count against the previous build and fail on a jump larger than some threshold.

bash
curl -s localhost:9090/metrics \
  | grep -c '^http_requests_total{'
# deploy, then run again and diff

The version most teams need warns rather than blocks: a unit test over the metric registry that flags any label whose value is set from a request-derived identifier. Static analysis catches this during review: far cheaper than a billing console, and it is the same gate observability that answers questions recommends for signal quality.

The alert that should exist

The most useful alert is not on your application. It is on your metric cardinality itself:

text
alert: MetricCardinalityExplosion
expr:   count by (__name__)({__name__=~".+"}) > 1000000
for:    10m

It fires before the database does, while the release is still fresh enough to attribute, and staging is where catching it is cheap.

Summary

Series cost multiplies across label values, so one identifier-shaped label can dominate the entire bill. Keep metric labels bounded and enumerable, push per-entity detail into logs and traces where it is queried one at a time, and assert on series count in CI. The invoice that triples overnight is almost always one line of code that looked entirely reasonable in review.

#observability#metrics#prometheus#monitoring#cost

SDP Clouds Team

DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations: every article is based on real incidents and real pipelines, not docs-page rewrites.

More about us →

Related articles