The Fifth Elephant 2026 Annual Conference

The Fifth Elephant 2026 Annual Conference

Built for humans. Now rebuilding for agents.

Tickets

Loading…

Shashank Rao

@shashankpr Facilitator

Beyond SWE-bench: How do we evaluate AI agents for real-world workflows?

Submitted Jul 1, 2026

Track: Track 2 – Building & Implementing AI Tools & Agents in Production
Format: Birds of a Feather (BoF) Session

Title

How do we evaluate AI agents for domain-specific, real-world workflows?

Description

AI agents deployed in real-world products do more than generate a response. They interpret user goals, make decisions, retrieve information, use tools, follow policies, take actions, recover from failures, and determine when to involve a human.

Evaluating such systems is challenging because there is often no public benchmark or universally correct answer for the workflow being automated. What constitutes a successful agent depends heavily on the product, domain, and intended user outcome.

For example, an agent handling a customer-support request may need to resolve the underlying problem, follow organisational policies, minimise user effort, avoid unsafe actions, and escalate appropriately. Similar challenges arise in incident response, IT operations, data quality, finance operations, procurement, compliance, and other domain-specific workflows.

In this BoF session, we will discuss how AI practitioners, product teams, and domain experts can jointly translate the intended outcome of an agentic feature into meaningful evaluation criteria.

This includes identifying domain-specific metrics, determining whether to evaluate only the final outcome or also the agent’s decisions and actions, and defining acceptable trade-offs between quality, user experience, reliability, latency, and cost.

We will also explore how teams can build representative evaluation datasets when no ready-made benchmark exists. Potential sources include production traces and logs, historical cases, domain-expert-authored scenarios, observed failures, and synthetically generated examples.

An important part of the discussion will be how offline evaluation results can be validated against online product, user, operational, and business metrics after deployment.

The goal is to exchange practical approaches for answering a central question:

How do we know that an agentic product feature is genuinely effective, reliable, and valuable in its intended domain?

The discussion will focus on evaluation design, domain reasoning, datasets, metrics, and product outcomes rather than any particular framework or implementation stack.

Discussion Outline

Defining the outcome
What user, operational, or business problem is the agent expected to solve? What would constitute success, partial success, or failure in that domain?

Defining domain-specific metrics
How do teams derive metrics for:

  • Resolution quality
  • Task completion
  • Policy adherence
  • Human handoff quality
  • Safety & Latency
  • Other product or domain-specific requirements

Evaluating outcomes and agent behaviour
Is evaluating the final result sufficient, or should teams also examine:

  • The agent’s decisions and intermediate actions
  • Tool selection and tool usage
  • Adherence to constraints and policies
  • Recovery from errors or incomplete information
  • Escalation and human-handoff behaviour
  • Whether the chosen path was efficient and appropriate

Building representative evaluation datasets
How can teams combine:

  • Production traces and logs
  • Historical cases
  • Domain-expert-authored scenarios
  • Observed failure modes
  • Synthetic data

How can this be done without losing realism, introducing dataset leakage, or overfitting to previously observed cases?

Connecting offline and online evaluation
How do we determine whether improvements in offline evaluation actually correlate with:

  • Better user experience and Improved business outcomes
  • Higher task-resolution rates
  • Safer agent behaviour
  • Reduced latency or cost

Choosing evaluation and judging methods

  • Which criteria can be evaluated deterministically?
  • Which require human or domain-expert review?
  • Where can LLM-based judges be useful, and how should they be calibrated or validated before being trusted?

Key Takeaways

Participants should leave with:

  • A clearer approach for translating product and domain goals into measurable evaluation criteria.
  • Ideas for evaluating both the outcome of an agentic workflow and the decisions or actions taken to reach it.
  • Practical strategies for creating evaluation datasets when no public benchmark or labelled dataset exists.
  • A better understanding of how production traces, domain expertise, synthetic scenarios, edge cases, and failure analysis can complement one another.
  • Approaches for connecting offline evaluation scores with online product, user, operational, and business metrics.
  • A framework for deciding which aspects of evaluation should use deterministic checks, human review, domain experts, or LLM-based judges.
  • A shared understanding of which parts of agent evaluation may be reusable across domains and which must remain specific to the product and workflow.

Who Should Attend?

This session is intended for cross-functional teams building and evaluating agentic systems, particularly:

  • Applied AI, ML, data science, and evaluation practitioners designing metrics, datasets, rubrics, graders, and evaluation processes.
  • Product managers and product leaders defining agentic features and determining whether those features deliver meaningful user and business value.
  • Engineering and platform teams operating production agents and investigating reliability, latency, cost, safety, observability, and failure modes.
  • Domain experts and operations teams whose knowledge is required to define correct behaviour, important edge cases, organisational policies, and acceptable outcomes.
  • Open-source contributors and researchers building agent-evaluation tools, datasets, test environments, observability systems, or reusable evaluation methodologies.

Participants do not need to work in a particular domain. Examples may include customer support, incident response, IT operations, CRM, compliance, procurement, finance operations, data workflows, developer tools, and other products in which agents perform multi-step tasks.

Suggested Reading and References

  1. A Survey on Evaluation of LLM-based Agents
    A broad overview of agent evaluation covering planning, tool use, interactive environments, application-specific benchmarks, safety, robustness, and evaluation using humans or language models (https://arxiv.org/abs/2503.16416)

  2. AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
    Introduces fine-grained progress metrics for evaluating what an agent accomplishes during multi-step execution, rather than relying only on final task success (https://arxiv.org/abs/2401.13178)

  3. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
    Evaluates agents that interact with users, use domain tools, modify environment state, and comply with policies in realistic workflows (https://openreview.net/forum?id=roNSXZpUDN)

  4. WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
    Demonstrates how a domain-specific evaluation environment can be built around realistic enterprise knowledge-work tasks (https://arxiv.org/abs/2403.07718)

  5. CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
    Explores realistic evaluation of agents across sales, service, and business workflows, including multi-turn interactions, confidentiality requirements, and expert-validated tasks (https://arxiv.org/abs/2505.18878)

  6. Demystifying Evals for AI Agents
    A practical guide to designing agent evaluation tasks, datasets, graders, and evaluation processes during product development (https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Get your hybrid access ticket

Hosted by

Jumpstart better data engineering and AI futures

Supported by

Platinum Sponsor

Atlassian unleashes the potential of every team. Our agile & DevOps, IT service management and work management software helps teams organize, discuss, and compl

Platinum Sponsor

Sahaj is an artisanal technology services company crafting purpose-built AI and data-led solutions for businesses.

Gold Sponsor

Skyflow secures the flow of data across datastores, models, and agents. Enterprises turn to Skyflow as their runtime AI data control layer to protect sensitive

Bronze Sponsor

Internet infrastructure APIs for IP geolocation and more

Bronze Sponsor

Open Source Analytical Database for the AI era.

Community sponsor

Real-time Observability & Governance layer for AI agents