The Fifth Elephant 2026 Annual Conference
Built for humans. Now rebuilding for agents.
Jul 2026
20 Mon
21 Tue
22 Wed
23 Thu
24 Fri
25 Sat
26 Sun
Jul 2026
27 Mon
28 Tue
29 Wed
30 Thu
31 Fri 08:45 AM – 06:00 PM IST
1 Sat
2 Sun
Shashank Rao
@shashankpr Facilitator
Submitted Jul 1, 2026
Track: Track 2 – Building & Implementing AI Tools & Agents in Production
Format: Birds of a Feather (BoF) Session
How do we evaluate AI agents for domain-specific, real-world workflows?
AI agents deployed in real-world products do more than generate a response. They interpret user goals, make decisions, retrieve information, use tools, follow policies, take actions, recover from failures, and determine when to involve a human.
Evaluating such systems is challenging because there is often no public benchmark or universally correct answer for the workflow being automated. What constitutes a successful agent depends heavily on the product, domain, and intended user outcome.
For example, an agent handling a customer-support request may need to resolve the underlying problem, follow organisational policies, minimise user effort, avoid unsafe actions, and escalate appropriately. Similar challenges arise in incident response, IT operations, data quality, finance operations, procurement, compliance, and other domain-specific workflows.
In this BoF session, we will discuss how AI practitioners, product teams, and domain experts can jointly translate the intended outcome of an agentic feature into meaningful evaluation criteria.
This includes identifying domain-specific metrics, determining whether to evaluate only the final outcome or also the agent’s decisions and actions, and defining acceptable trade-offs between quality, user experience, reliability, latency, and cost.
We will also explore how teams can build representative evaluation datasets when no ready-made benchmark exists. Potential sources include production traces and logs, historical cases, domain-expert-authored scenarios, observed failures, and synthetically generated examples.
An important part of the discussion will be how offline evaluation results can be validated against online product, user, operational, and business metrics after deployment.
The goal is to exchange practical approaches for answering a central question:
How do we know that an agentic product feature is genuinely effective, reliable, and valuable in its intended domain?
The discussion will focus on evaluation design, domain reasoning, datasets, metrics, and product outcomes rather than any particular framework or implementation stack.
Defining the outcome
What user, operational, or business problem is the agent expected to solve? What would constitute success, partial success, or failure in that domain?
Defining domain-specific metrics
How do teams derive metrics for:
Evaluating outcomes and agent behaviour
Is evaluating the final result sufficient, or should teams also examine:
Building representative evaluation datasets
How can teams combine:
How can this be done without losing realism, introducing dataset leakage, or overfitting to previously observed cases?
Connecting offline and online evaluation
How do we determine whether improvements in offline evaluation actually correlate with:
Choosing evaluation and judging methods
Participants should leave with:
This session is intended for cross-functional teams building and evaluating agentic systems, particularly:
Participants do not need to work in a particular domain. Examples may include customer support, incident response, IT operations, CRM, compliance, procurement, finance operations, data workflows, developer tools, and other products in which agents perform multi-step tasks.
A Survey on Evaluation of LLM-based Agents
A broad overview of agent evaluation covering planning, tool use, interactive environments, application-specific benchmarks, safety, robustness, and evaluation using humans or language models (https://arxiv.org/abs/2503.16416)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Introduces fine-grained progress metrics for evaluating what an agent accomplishes during multi-step execution, rather than relying only on final task success (https://arxiv.org/abs/2401.13178)
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Evaluates agents that interact with users, use domain tools, modify environment state, and comply with policies in realistic workflows (https://openreview.net/forum?id=roNSXZpUDN)
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Demonstrates how a domain-specific evaluation environment can be built around realistic enterprise knowledge-work tasks (https://arxiv.org/abs/2403.07718)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Explores realistic evaluation of agents across sales, service, and business workflows, including multi-turn interactions, confidentiality requirements, and expert-validated tasks (https://arxiv.org/abs/2505.18878)
Demystifying Evals for AI Agents
A practical guide to designing agent evaluation tasks, datasets, graders, and evaluation processes during product development (https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)
Hosted by
Supported by
Platinum Sponsor
Platinum Sponsor
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}