Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Suraj Nath
@electron0zero
Submitted Sep 29, 2026
Session title
Metrics from Traces at Scale: Surviving the Cardinality Explosion
One-line summary
Generating metrics from traces gives you RED metrics and service graphs with zero instrumentation, until one span attribute like http.url creates a new series per request and takes down your metrics backend.
This is the production story of building and operating cardinality protection in Grafana Tempo’s metrics-generator
What problem are you addressing? What real engineering challenge, question, or experience is this session about?
Trace-derived metrics are the worst-behaved metric source by construction. Labels come from span attributes that application developers never designed for metrics, and a remote-write producer bypasses every scrape-time protection like sample_limit, label_limit, relabeling rules, etc.
Thousands of app teams can’t be relied on to keep attributes disciplined, so the producer is the last controllable point.
The challenge was capping the high cardinilatiy labels without dropping every other useful dimension, while keeping the cost low at scale.
Who is the intended audience?
SRE, Platform Engineering, Infrastructure, Developers working on observability pipelines
Level
Intermediate
List one or two practical takeaways
What will you share?
Architecture and design decisions, code and implementation details, benchmarks, failure modes, operational challenges, open-source tooling
Everything shown is shipped in Grafana Tempo 3.0 and applies to any traces-to-metrics pipeline, including OpenTelemetry Collector connectors.
Code is open source: https://github.com/grafana/tempo
What is your experience with this problem?
Production system, open-source project, real incident or failure, hard-earned engineering lesson
What approaches failed, disappointed, or created unexpected problems?
OTel semantic-convention metrics (target_info, host_info) that had to be exempted from the limiter and sanitizer,
What will you do differently today?
Not a whole lot, this is working fine for us and customers like the feature, and it’s running in production without major downsides.
What trade-offs did you consider?
How can this help other practitioners?
A pattern to adopt: producer-side per-label limiting.
A set of operational practices: dry run, alerts, rollout checklist, etc
A mistake to avoid: limits without observability into the limit itself.
Current state
Production experience
#observability #sre #databases #scalability #failurestory #demo #casestudy
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}