Agent systems work through loops: plan, call tools, inspect results, recover, and decide when to stop. As these systems move beyond demos, the harder question is not whether they can complete a task, but whether the path was reliable, efficient, grounded, and worth its cost and what should change when it was not.
This Birds of a Feather session will examine how to evaluate and improve these agent quality loops. When an agent gets stuck, retries a tool, wastes tokens, calls the wrong interface, uses weak evidence, or returns a plausible but incomplete answer, how do we determine whether the problem lies in the model, prompt, tool contract, context, orchestration, or evaluation itself?
The panel will establish a shared frame for traces, logs, dashboards, and evals: what each can prove and what level of rigor is useful. The panel will then drive the discussion through concrete failure patterns and turn-level signals, while participants share their own experiences, approaches, and open questions.
Together, the room will examine how to distinguish useful work from wasted turns and token costs, which signals lead to meaningful engineering decisions, and how findings should flow through the quality loop:
Observe the turns → evaluate the path and outcome → classify the failure → change the system → measure again
The goal is to build a shared map of what agent quality loops should contain, what teams are building today, and which patterns should become common practice for production agent systems.
- What should a useful agent quality view show at the turn level?
- What is the difference between agent traces, logs, and evals?
- How should teams measure wasted tokens, unnecessary turns, retries, and avoidable latency?
- How do we distinguish model failures from tool, context, schema, orchestration, or evaluation failures?
- Which path-level (agent trajectory) signals reveal real quality problems, and which are merely interesting metrics?
- How should findings feed back into tooling, benchmarks, evaluations, and release decisions?
- What level of rigor is appropriate before an agent moves from experimentation to a production workflow?
- A practical model for evaluating both an agent’s answer and the path it took.
- A clearer way to turn trace, cost, and evaluation evidence into engineering improvements.
- AI and platform engineers building production agent systems
- Technical leads designing and operating agent quality loops
- Leaders responsible for agent evaluation, observability, cost, or reliability
Prateek Mandloi, Principal Engineer at Atlassian works on agentic AI systems, enterprise work graphs, CLI and tooling interfaces, and benchmark-driven quality loops for AI agents. Prateek has architected large scale distributed systems and is applying alot of learnings into engineer AI native tools and infrastructure at scale. He previously hosted a Birds of a Feather session at The Fifth Elephant 2024 on enterprise data lifecycle and AI analytics.
Sunny R Gupta, Head of Engineering and India Site Lead at Metaforms, an AI platform building agentic workflows for market research agencies. Based in Bengaluru, he has architected and operated large-scale systems handling billions of events per day across data infrastructure, streaming, and consumer products. Sunny also runs TeamShiksha and Impromptu Meetups, mentoring engineers and hosting community learning sessions on distributed systems, data platforms, and developer productivity. He brings operability practices from real-world systems, with a bias toward first-principles thinking under constraints.
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}