Suraj Nath

@electron0zero

Metrics from Traces at Scale: Surviving the Cardinality Explosion

Submitted Sep 29, 2026

Session title

Metrics from Traces at Scale: Surviving the Cardinality Explosion

One-line summary

Generating metrics from traces gives you RED metrics and service graphs with zero instrumentation, until one span attribute like http.url creates a new series per request and takes down your metrics backend.

This is the production story of building and operating cardinality protection in Grafana Tempo’s metrics-generator

What problem are you addressing? What real engineering challenge, question, or experience is this session about?

Trace-derived metrics are the worst-behaved metric source by construction. Labels come from span attributes that application developers never designed for metrics, and a remote-write producer bypasses every scrape-time protection like sample_limit, label_limit, relabeling rules, etc.

Thousands of app teams can’t be relied on to keep attributes disciplined, so the producer is the last controllable point.

The challenge was capping the high cardinilatiy labels without dropping every other useful dimension, while keeping the cost low at scale.

Who is the intended audience?

SRE, Platform Engineering, Infrastructure, Developers working on observability pipelines

Level

Intermediate

List one or two practical takeaways

  • How to estimate per-label cardinality demand cheaply with HyperLogLog sketches, and cap only the offending label without losing other dimensions

What will you share?

Architecture and design decisions, code and implementation details, benchmarks, failure modes, operational challenges, open-source tooling

Everything shown is shipped in Grafana Tempo 3.0 and applies to any traces-to-metrics pipeline, including OpenTelemetry Collector connectors.

Code is open source: https://github.com/grafana/tempo

What is your experience with this problem?

Production system, open-source project, real incident or failure, hard-earned engineering lesson

What approaches failed, disappointed, or created unexpected problems?

OTel semantic-convention metrics (target_info, host_info) that had to be exempted from the limiter and sanitizer,

What will you do differently today?

Not a whole lot, this is working fine for us and customers like the feature, and it’s running in production without major downsides.

What trade-offs did you consider?

  • Producer-side vs consumer-side enforcement (TSDB limits, scrape limits, sidecars).
  • Probabilistic sketches (about 5KB per tenant, roughly 3% error) vs exact counting (about 16 bytes per series).
  • Per-label overflow vs dropping whole series.
  • Fixing at instrumentation time vs at the producer.

How can this help other practitioners?

A pattern to adopt: producer-side per-label limiting.
A set of operational practices: dry run, alerts, rollout checklist, etc
A mistake to avoid: limits without observability into the limit itself.

Current state

Production experience

#observability #sre #databases #scalability #failurestory #demo #casestudy

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy