saurabh hirani

saurabh hirani

@saurabh_hirani

Observing AI apps: know your instrumentation

Submitted Oct 3, 2026

Summary

As we helped customers instrument their AI applications, the tooling matured alongside the work, and each engagement pushed us to cover more surface area and get more visibility: from vanilla OpenTelemetry on RAG apps, to OpenLLMetry, the Bifrost AI gateway, and OpenLIT for agents, and now Langfuse. We rebuilt each stage on an open-source reference setup to show what each approach sees, what it misses, and why we moved to the next one.

What problem are you addressing?

Teams building AI features need answers to questions that get harder over time. We set out to answer questions like is it up, what does it cost, are the retrieved documents relevant, which model is the spend going to, etc., one layer of instrumentation at a time.

What made it hard:

  • AI failures don’t show up in HTTP status codes. Agents loop, retrievals come back empty, and a request can return 200 with a useless answer.

To compare approaches fairly, we built one reference setup to showcase our learnings:

  • A RAG app for the first few experiments.
  • An incident-triage agent for the later ones.
  • An OTel collector gateway with swappable backends (Prometheus, Tempo, Loki, Grafana), and an optional AI gateway.

Who is the intended audience?

  • Platform Engineering, SRE, and DevOps engineers who run or are about to run LLM-backed services.
  • Engineering leaders who need the right business metrics from their AI applications and insight into how their agents perform.

Level

Intermediate. Assumes OpenTelemetry and Prometheus basics; no ML background needed.

Practical takeaways

  • A comparison of the main ways to instrument AI applications (manual OpenTelemetry, auto-instrumentation libraries, a hybrid of both, and an AI gateway): the pros, cons, and limitations of each, and what each one lets you see.

What will you share?

  • Architecture or design decisions

    • How we decided between manual OpenTelemetry, auto-instrumentation, and a hybrid of both, based on what the application needed to show at that stage.
    • What we moved into an AI gateway and what stayed in application instrumentation.
    • Why domain-specific signals, such as retrieval relevance, stayed in application code regardless of the library used.
  • Failure modes and debugging

    • For each approach, alongside the demo code, we have documented which failure modes it catches, and which it can’t.
  • Live demo

    • The open source repo has locally runnable docker setups accompanied by demo screenshots which can be used during the talk.
  • Open-source tooling

What is your experience with this problem?

  • Research / investigation
  • Hard earned engineering lesson

What approaches failed, disappointed, or created unexpected problems?

  • Auto-instrumentation removed visibility we already had.
    • OpenLLMetry added token and model data with one import.
    • It did not see our RAG pipeline steps (embedding, vector search, generation), so we had to add manual spans back.
  • Agent traces without agent metrics.
    • OpenLLMetry traced our Agents SDK agent but did not emit workflow or tool duration metrics.
    • The dashboard could not explain where most of the request time went.
  • Not understanding histogram fundamentals led us to misread dashboards.
    • Latency percentiles we read as measurements were bucket boundaries.
    • Percentiles for retrieval similarity scores were meaningless until we fixed the buckets.

What will you do differently today?

  • Instead of going to either extreme, rolling OpenTelemetry by hand or adopting a managed solution wholesale, evaluate what makes sense for the application’s current stage and choose the instrumentation method accordingly.

What trade-offs did you consider?

  • Manual vs auto-instrumentation for agents.
    • The hand-written loop gave full visibility but meant writing and maintaining both the loop and its instrumentation.
    • The Agents SDK cut the agent code to a few lines but left visibility up to the instrumentation library.
  • Gateway vs in-app instrumentation.
    • A gateway gives per-app and per-model cost, latency, and spend limits without touching app code.
    • It sees only the model calls, not what the app does around them.
  • What we gave up.
    • We dropped OpenLLMetry for the agent experiments and moved to OpenLIT, which emitted the workflow, tool, and cost metrics we needed.

How can this help other practitioners?

  • A way to evaluate competing approaches.
    • Keep the app the same, change the instrumentation layer, and compare what you can see.

Tags

#observability #agents #finops #sre #platformengineering #opentelemetry

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy