Track: Track 2 – Building & Implementing AI Tools & Agents in Production
Format: Birds of a Feather (BoF) Session
How do we evaluate AI agents for domain-specific, real-world workflows?
AI agents deployed in real-world products do more than generate a response. They interpret user goals, make decisions, retrieve information, use tools, follow policies, take actions, recover from failures, and determine when to involve a human.
Evaluating such systems is challenging because there is often no public benchmark or universally correct answer for the workflow being automated. What constitutes a successful agent depends heavily on the product, domain, and intended user outcome.
For example, an agent handling a customer-support request may need to resolve the underlying problem, follow organisational policies, minimise user effort, avoid unsafe actions, and escalate appropriately. Similar challenges arise in incident response, IT operations, data quality, finance operations, procurement, compliance, and other domain-specific workflows.
In this BoF session, we will discuss how AI practitioners, product teams, and domain experts can jointly translate the intended outcome of an agentic feature into meaningful evaluation criteria.
This includes identifying domain-specific metrics, determining whether to evaluate only the final outcome or also the agent’s decisions and actions, and defining acceptable trade-offs between quality, user experience, reliability, latency, and cost.
We will also explore how teams can build representative evaluation datasets when no ready-made benchmark exists. Potential sources include production traces and logs, historical cases, domain-expert-authored scenarios, observed failures, and synthetically generated examples.
An important part of the discussion will be how offline evaluation results can be validated against online product, user, operational, and business metrics after deployment.
The goal is to exchange practical approaches for answering a central question:
How do we know that an agentic product feature is genuinely effective, reliable, and valuable in its intended domain?
The discussion will focus on evaluation design, domain reasoning, datasets, metrics, and product outcomes rather than any particular framework or implementation stack.
Defining the outcome
What user, operational, or business problem is the agent expected to solve? What would constitute success, partial success, or failure in that domain?
Defining domain-specific metrics
How do teams derive metrics for:
- Resolution quality
- Task completion
- Policy adherence
- Human handoff quality
- Safety & Latency
- Other product or domain-specific requirements
Evaluating outcomes and agent behaviour
Is evaluating the final result sufficient, or should teams also examine:
- The agent’s decisions and intermediate actions
- Tool selection and tool usage
- Adherence to constraints and policies
- Recovery from errors or incomplete information
- Escalation and human-handoff behaviour
- Whether the chosen path was efficient and appropriate
Building representative evaluation datasets
How can teams combine:
- Production traces and logs
- Historical cases
- Domain-expert-authored scenarios
- Observed failure modes
- Synthetic data
How can this be done without losing realism, introducing dataset leakage, or overfitting to previously observed cases?
Connecting offline and online evaluation
How do we determine whether improvements in offline evaluation actually correlate with:
- Better user experience and Improved business outcomes
- Higher task-resolution rates
- Safer agent behaviour
- Reduced latency or cost
Choosing evaluation and judging methods
- Which criteria can be evaluated deterministically?
- Which require human or domain-expert review?
- Where can LLM-based judges be useful, and how should they be calibrated or validated before being trusted?
Participants should leave with:
- A clearer approach for translating product and domain goals into measurable evaluation criteria.
- Ideas for evaluating both the outcome of an agentic workflow and the decisions or actions taken to reach it.
- Practical strategies for creating evaluation datasets when no public benchmark or labelled dataset exists.
- A better understanding of how production traces, domain expertise, synthetic scenarios, edge cases, and failure analysis can complement one another.
- Approaches for connecting offline evaluation scores with online product, user, operational, and business metrics.
- A framework for deciding which aspects of evaluation should use deterministic checks, human review, domain experts, or LLM-based judges.
- A shared understanding of which parts of agent evaluation may be reusable across domains and which must remain specific to the product and workflow.
This session is intended for cross-functional teams building and evaluating agentic systems, particularly:
- Applied AI, ML, data science, and evaluation practitioners designing metrics, datasets, rubrics, graders, and evaluation processes.
- Product managers and product leaders defining agentic features and determining whether those features deliver meaningful user and business value.
- Engineering and platform teams operating production agents and investigating reliability, latency, cost, safety, observability, and failure modes.
- Domain experts and operations teams whose knowledge is required to define correct behaviour, important edge cases, organisational policies, and acceptable outcomes.
- Open-source contributors and researchers building agent-evaluation tools, datasets, test environments, observability systems, or reusable evaluation methodologies.
Participants do not need to work in a particular domain. Examples may include customer support, incident response, IT operations, CRM, compliance, procurement, finance operations, data workflows, developer tools, and other products in which agents perform multi-step tasks.
-
A Survey on Evaluation of LLM-based Agents
A broad overview of agent evaluation covering planning, tool use, interactive environments, application-specific benchmarks, safety, robustness, and evaluation using humans or language models (https://arxiv.org/abs/2503.16416)
-
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Introduces fine-grained progress metrics for evaluating what an agent accomplishes during multi-step execution, rather than relying only on final task success (https://arxiv.org/abs/2401.13178)
-
τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Evaluates agents that interact with users, use domain tools, modify environment state, and comply with policies in realistic workflows (https://openreview.net/forum?id=roNSXZpUDN)
-
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
Demonstrates how a domain-specific evaluation environment can be built around realistic enterprise knowledge-work tasks (https://arxiv.org/abs/2403.07718)
-
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
Explores realistic evaluation of agents across sales, service, and business workflows, including multi-turn interactions, confidentiality requirements, and expert-validated tasks (https://arxiv.org/abs/2505.18878)
-
Demystifying Evals for AI Agents
A practical guide to designing agent evaluation tasks, datasets, graders, and evaluation processes during product development (https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}