As we helped customers instrument their AI applications, the tooling matured alongside the work, and each engagement pushed us to cover more surface area and get more visibility: from vanilla OpenTelemetry on RAG apps, to OpenLLMetry, the Bifrost AI gateway, and OpenLIT for agents, and now Langfuse. We rebuilt each stage on an open-source reference setup to show what each approach sees, what it misses, and why we moved to the next one.
Teams building AI features need answers to questions that get harder over time. We set out to answer questions like is it up, what does it cost, are the retrieved documents relevant, which model is the spend going to, etc., one layer of instrumentation at a time.
What made it hard:
- AI failures don’t show up in HTTP status codes. Agents loop, retrievals come back empty, and a request can return 200 with a useless answer.
To compare approaches fairly, we built one reference setup to showcase our learnings:
- A RAG app for the first few experiments.
- An incident-triage agent for the later ones.
- An OTel collector gateway with swappable backends (Prometheus, Tempo, Loki, Grafana), and an optional AI gateway.
- Platform Engineering, SRE, and DevOps engineers who run or are about to run LLM-backed services.
- Engineering leaders who need the right business metrics from their AI applications and insight into how their agents perform.
Intermediate. Assumes OpenTelemetry and Prometheus basics; no ML background needed.
- A comparison of the main ways to instrument AI applications (manual OpenTelemetry, auto-instrumentation libraries, a hybrid of both, and an AI gateway): the pros, cons, and limitations of each, and what each one lets you see.
- Research / investigation
- Hard earned engineering lesson
- Auto-instrumentation removed visibility we already had.
- OpenLLMetry added token and model data with one import.
- It did not see our RAG pipeline steps (embedding, vector search, generation), so we had to add manual spans back.
- Agent traces without agent metrics.
- OpenLLMetry traced our Agents SDK agent but did not emit workflow or tool duration metrics.
- The dashboard could not explain where most of the request time went.
- Not understanding histogram fundamentals led us to misread dashboards.
- Latency percentiles we read as measurements were bucket boundaries.
- Percentiles for retrieval similarity scores were meaningless until we fixed the buckets.
- Instead of going to either extreme, rolling OpenTelemetry by hand or adopting a managed solution wholesale, evaluate what makes sense for the application’s current stage and choose the instrumentation method accordingly.
- Manual vs auto-instrumentation for agents.
- The hand-written loop gave full visibility but meant writing and maintaining both the loop and its instrumentation.
- The Agents SDK cut the agent code to a few lines but left visibility up to the instrumentation library.
- Gateway vs in-app instrumentation.
- A gateway gives per-app and per-model cost, latency, and spend limits without touching app code.
- It sees only the model calls, not what the app does around them.
- What we gave up.
- We dropped OpenLLMetry for the agent experiments and moved to OpenLIT, which emitted the workflow, tool, and cost metrics we needed.
- A way to evaluate competing approaches.
- Keep the app the same, change the instrumentation layer, and compare what you can see.
#observability #agents #finops #sre #platformengineering #opentelemetry
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}