Metric Cardinality: Keep Labels Within a Budget

Control metric cardinality with bounded labels, route templates and per-metric budgets, while moving user-specific detail into appropriately protected logs.

In this article

Metric cardinality is the number of distinct time series created by label combinations. A label containing user IDs, complete URLs or request IDs can create a new series for nearly every event. Keep metrics labels bounded and reserve individual-event detail for logs or traces with suitable access controls.

The goal is not to remove useful context. It is to choose context that supports aggregation without turning a monitoring system into an expensive record of every request.

Count combinations, not just label names

Suppose a metric has 10 routes, 4 methods and 5 status groups. That permits up to 200 combinations before adding other labels. Add 20 instances and the potential grows to 4,000. A customer identifier with thousands of values can expand that number far beyond the original intention.

These figures are arithmetic examples, not measurements of a particular monitoring backend. Actual series counts depend on observed combinations and the system's behavior. Still, estimating the upper bound during review catches obvious problems before deployment.

json
{"metric":"http_requests_total",
 "labels":{"route":"/orders/:id","method":"GET","status_group":"2xx"},
 "excluded_labels":["customer_id","request_id","raw_url"]}

Inspect the proposed schema with JSON Viewer. Compare releases with JSON Diff so an accidental new unbounded label is visible in review.

Use route templates rather than raw paths

For a request to /orders/12345, the metric can use a route template such as /orders/:id. Otherwise every order creates a distinct label value. Query strings make raw URLs even less suitable because tokens, search terms and identifiers can vary continuously.

Prefer a route name supplied by the framework's router rather than a regular expression assembled after the fact. If you must normalize paths, test the normalization on representative routes and unknown paths. A catch-all fallback should not expose complete private URLs.

The Prometheus naming guidance discusses metric and label design. The operational lesson is that label values need an intended, reviewable domain rather than an open-ended string supplied by a request.

Separate dimensions from event details

Good metric dimensions answer questions such as which service, operation or status class is affected. Event details answer which exact request or customer experienced the problem. Both matter, but they do not need to live in the same storage model.

Keep detailed records in appropriately protected logs or traces, with retention and redaction rules. Link aggregate incidents to those records through supported observability features where useful. Do not paste raw customer identifiers into labels simply because dashboards make filtering convenient.

Hashing an identifier does not solve cardinality. A unique input usually remains a unique hash. It may also remain linkable to the same person across systems. Removing a readable name is not the same as creating a bounded metric dimension.

Budget each important metric

List labels, expected value counts and acceptable series growth. Give the metric an owner and a review trigger. A growing product can legitimately add services or routes, but an unexpected explosion should prompt investigation.

Include histogram behavior in the estimate if the monitoring system represents buckets as additional series. Do not count only the metric's business labels and forget the representation overhead. Use your backend's documented accounting when estimating storage or ingestion costs.

The Prometheus instrumentation guide offers instrumentation principles. Apply them to your workload instead of inventing a universal safe label count; reasonable budgets vary with traffic and available infrastructure.

Diagnose a sudden increase

Identify which metric and label are growing. Compare the time of growth with deployments, new endpoints and instrumentation changes. Inspect a small sample of label values to determine whether a field became unbounded.

Do not export all values if they may contain personal data or secrets. Raw query strings in metrics can create a privacy incident as well as a resource problem. Restrict access and preserve the relevant evidence according to your operational process.

If a deployment caused the growth, remove the problematic label or roll back the instrumentation safely. Existing series may remain until retention expires; a code fix can stop new growth without immediately shrinking storage. Explain that distinction to the team.

Test the label policy

Send requests with different object IDs, search terms and unknown paths in staging. Confirm the number of route-label values remains bounded. Test that no credential-like query parameter reaches telemetry labels.

Run concurrent traffic and inspect exported metrics, not just application configuration. A wrapper may attach labels after your reviewed instrumentation. Keep a regression test for the actual output schema when a high-traffic service depends on it.

Review labels added by automatic instrumentation as well as labels your team wrote. Default request or exception attributes may have broader value ranges than expected. Check a deployed sample after framework upgrades, and keep a documented fallback label for unknown operations so unexpected traffic cannot expand the schema without a limit.

Keep metrics aggregatable

Use bounded labels that explain system behavior, estimate their combinations and monitor growth after deployment. Move individual-event detail to the right evidence channel. A well-designed metric should remain useful as traffic increases, without creating a new series for every user action.

Advertisement
Metric Cardinality: Keep Labels Within a Budget | Duck Cloud