not accepting clients
← back to blog

prometheus cardinality

every unique label combination is its own series in prometheus, and one label holding an id or a raw path can outweigh every other metric combined.

prometheus has been running for a year at a steady few gigabytes of memory. a service ships a release, and over the next two hours the prometheus pod climbs past its limit and is oom-killed. it restarts, replays its write-ahead log, climbs again, and is killed again. nothing in the release looked like a monitoring change. somebody added a label to a histogram so a dashboard could break latency down by customer.

prometheus does not store metrics, it stores series, and a series is a metric name plus one exact set of label values. a label is not a column that gets filled in; every new value it takes creates a new series, with its own index entry and its own chunks in memory. the cost of a metric is how many distinct label combinations it ever produces.

the short version

one histogram with three ordinary labels on a modest fleet:

method   (4) ─┐
route   (25) ─┼─▶ 600 label sets ─▶ ×14 series each ─▶ 8,400 per pod ─▶ ×30 pods ─▶ 252,000
code     (6) ─┘                     12 buckets, _sum, _count

every factor multiplies. the histogram turns each label set into fourteen series because the default go client buckets are eleven boundaries plus +Inf, and _sum and _count come along with them. the pod count multiplies again because each target carries its own instance and pod labels. a label with 25 values is fine. replace route with the raw request path, and the 25 becomes the number of distinct urls the service has ever seen - one per order id, one per user id - and the right-hand side stops being a number anyone chose.

why memory follows series

new samples land in the head block, which holds the most recent two to three hours of data in memory before it is compacted into a block on disk. every series that received a sample in that window is in the head: its labels in the inverted index, its open chunk, its postings. a commonly quoted rule of thumb is a few kilobytes per active series, and it is the series count, not the sample rate, that moves the number.

that is why the release above took two hours to kill prometheus rather than two minutes. each new customer that sent a request added a set of series, and none of them left until the head was truncated. the current head size is one query away:

prometheus_tsdb_head_series

graph it next to the container’s memory and they move together.

the usual offenders

ids. user ids, tenant ids, request ids, trace ids, order numbers. any label whose value comes from a database key grows with the business. per-request detail belongs in logs and traces; if the point is to jump from a slow bucket to an example request, that is what exemplars are for - a trace id attached to an observation, not a label on the series.

raw paths. /orders/8f14e45f and /orders/c9f0f895 are the same endpoint and two series. the label has to hold the route template the router matched, /orders/{id}, not the path the client sent. raw paths also let anyone on the internet create series by requesting urls that do not exist.

error strings. err.Error() as a label value carries addresses, ports, timestamps and ids inside the message. a closed set of error kinds - timeout, refused, invalid - says the same thing in a handful of series.

pod names on high-churn workloads. this one is subtler, and is covered under churn below.

finding the culprit

the tsdb status page in the prometheus ui, and the api behind it, report the head block’s heaviest metrics and labels without running a query:

curl -s 'http://prometheus:9090/api/v1/status/tsdb?limit=10' | jq '.data'

the response has headStats with the total series count, seriesCountByMetricName, labelValueCountByLabelName, memoryInBytesByLabelName, and seriesCountByLabelValuePair. a label name with tens of thousands of values in labelValueCountByLabelName is almost always the answer, and seriesCountByMetricName says which metric carries it.

the same question in promql works on any prometheus, including one behind a query layer, but it touches every series in the database and is expensive on exactly the instance that is already struggling:

topk(10, count by (__name__)({__name__=~".+"}))

once the metric is known, count the distinct values of a suspect label:

count(count by (route) (http_request_duration_seconds_count))

and to find which target is producing it, prometheus records the post-relabeling sample count of every scrape as a synthetic series:

topk(10, scrape_samples_post_metric_relabeling)

for a block already on disk, promtool reads the tsdb directly and reports label pair cardinality and churn without touching the running server:

promtool tsdb analyze /prometheus --limit 20

fixing it at the source

the right fix is in the code that registers the metric. with client_golang, the route label is filled in once per handler at registration time from the pattern it is registered under, so no request can influence it:

var requestDuration = promauto.NewHistogramVec(
	prometheus.HistogramOpts{
		Name:    "http_request_duration_seconds",
		Help:    "time spent serving http requests.",
		Buckets: prometheus.DefBuckets,
	},
	[]string{"route", "method", "code"},
)

func instrument(route string, h http.Handler) http.Handler {
	return promhttp.InstrumentHandlerDuration(
		requestDuration.MustCurryWith(prometheus.Labels{"route": route}),
		h,
	)
}

func routes() http.Handler {
	mux := http.NewServeMux()
	mux.Handle("GET /orders/{id}", instrument("/orders/{id}", http.HandlerFunc(getOrder)))
	mux.Handle("POST /orders", instrument("/orders", http.HandlerFunc(createOrder)))
	return mux
}

the version this replaces calls requestDuration.WithLabelValues(r.URL.Path, ...) from a middleware, and is correct in every test that uses one order id. routers that resolve the pattern at request time expose it too - chi’s chi.RouteContext(r.Context()).RoutePattern() returns the matched template once routing has finished - and the rule is the same: the label value comes from the route table, never from the request.

fixing it at scrape time

when the code cannot change quickly - a third-party exporter, a release that is already out - the series can be dropped before they are stored, with metric_relabel_configs on the scrape config:

scrape_configs:
  - job_name: orders
    metric_relabel_configs:
      # drop a whole metric nobody reads
      - source_labels: [__name__]
        regex: go_memstats_.*
        action: drop
      # strip the label that exploded
      - regex: customer_id
        action: labeldrop

the order matters. relabel_configs runs before the scrape, against the target’s labels from service discovery - it decides what to scrape and what the target is called, and it never sees a metric name. metric_relabel_configs runs after the scrape, against every sample, before ingestion. a drop there keeps the series out of the tsdb, but the target still renders them, the network still carries them, and prometheus still parses them. it is a tourniquet, and the fix still belongs in the code.

guardrails

the scrape config can also refuse a target that misbehaves:

scrape_configs:
  - job_name: orders
    sample_limit: 20000
    label_limit: 30
    label_value_length_limit: 200

each limit is checked after metric relabeling, and exceeding any of them fails the entire scrape, not just the excess samples. the target is marked down, up goes to 0, and every series from it goes stale until the target comes back under the limit. that is the point: a hard, visible failure on the one target instead of a slow climb that takes prometheus down for everyone. it does mean the limit has to sit well above normal, and prometheus_target_scrapes_exceeded_sample_limit_total needs an alert of its own.

with the prometheus operator, the same limits are sampleLimit, labelLimit and labelValueLengthLimit on a ServiceMonitor or PodMonitor, and enforcedSampleLimit on the Prometheus resource sets a ceiling that individual teams cannot raise.

churn is a different problem

cardinality is how many series exist at once. churn is how fast they are replaced. a deployment with 30 replicas has 30 values of pod at any moment, which is fine. roll it out ten times a day and the head sees 300 values, because the series from the old pods are still in memory until the head is truncated. batch jobs and cronjobs that create a pod per run, and autoscalers that add and remove replicas all day, do the same thing with no deploy involved.

rate(prometheus_tsdb_head_series_created_total[1h])

a steady head series count with a high creation rate is churn. the fixes are different from cardinality: keep pod-level labels off metrics that are only ever aggregated per service, and keep kube-state-metrics from copying every kubernetes label onto its series - --metric-labels-allowlist is empty by default for a reason.

native histograms

the ×14 in the diagram is the classic histogram’s cost: one series per bucket, plus _sum and _count. a native histogram stores the whole distribution in a single series per label set, with buckets chosen exponentially at write time instead of fixed in code. they are stable as of prometheus 3.8, enabled per scrape config with scrape_native_histograms: true, and in client_golang a histogram exposes them when NativeHistogramBucketFactor is set in its options. they do not fix a bad label - 600 label sets are still 600 series - but they take the multiplier off.

what to watch out for

labeldrop can merge series. if the dropped label was the only thing distinguishing two series, they become the same series with two samples at one timestamp, and prometheus rejects one of them. drop labels that are redundant or that the metric can lose entirely; aggregate the rest in the code.

the limits fail the whole target. a sample_limit set at today’s count plus ten percent turns the next legitimate new metric into a full outage of that target’s monitoring, including its alerts. set it as a ceiling for runaway growth, not a quota.

the discovery query is part of the load. count by (__name__) over every series is the heaviest query most people ever run. use the tsdb status endpoint first, and run the promql against a replica if there is one.

remote write multiplies the bill. a hosted backend that prices by active series charges for every one of them, and the label that only made prometheus slower locally becomes a line on an invoice. the same fixes apply, and they apply before the data leaves the cluster.

the metric is usually fine and the label is not. removing a metric to save memory loses a signal. removing the one label that made it unbounded usually keeps everything anyone was using.

references

[1] prometheus documentation. “configuration: scrape_config.”
prometheus.io/docs/prometheus/latest/configuration/configuration

[2] prometheus documentation. “http api: tsdb stats.”
prometheus.io/docs/prometheus/latest/querying/api

[3] prometheus documentation. “promtool: tsdb analyze.”
prometheus.io/docs/prometheus/latest/command-line/promtool

[4] prometheus documentation. “native histograms.”
prometheus.io/docs/specs/native_histograms

[5] prometheus documentation. “instrumentation: use labels.”
prometheus.io/docs/practices/instrumentation

[6] prometheus operator documentation. “api reference.”
prometheus-operator.dev/docs/api-reference/api

# ask the author

a question
about this
post?

direct line