Unavailable

This livestream is restricted

Already a member? Login with your membership email address

The Fifth Elephant 2026 Annual Conference

Built for humans. Now rebuilding for agents.

Tickets

Loading…

Priyanga P Kini

Priyanga P Kini

@PriyangaPKini

Designing Reliable Agentic Loops

Submitted Jul 23, 2026

Overview

“Loops” are becoming a popular mental model for how we build with agents. Agents reach their goals by planning, calling tools, inspecting results, retrying failed steps, using memory, and sometimes calling other agents. As these loops become more autonomous, the question is no longer just whether the agent produced a good final answer, but whether each part of the loop is working reliably.

A bad agent outcome may come from many places: weak prompt instructions, poor retrieval, wrong tool selection, bad tool arguments, ignored observations, stale memory, unnecessary retries, or final generation that is not grounded in evidence. Evaluating only the final answer hides these failure modes.

This BOF will discuss component-wise evaluation of agent loops: which parts to evaluate separately, what metrics are useful for each component, and how error analysis can turn failed runs into regression tests. We’ll also discuss the role of annotation tools in collecting better human feedback, building better eval datasets, and improving prompts, retrieval, tool use, and agent behavior over time.

Discussion questions

  • Which components of an agent loop should be evaluated separately?
  • Which metrics are useful for prompt, retrieval, planning, tool use, memory, final generation, cost, and latency?
  • How can error analysis turn failed agent runs into regression tests?
  • What makes an annotation tool easy enough for experts to use regularly?

Takeaways

Participants should leave with:

  • A practical checklist for component-wise evaluation of agent loops.
  • A clearer understanding of how error analysis, regression tests, and annotation tools improve agent reliability over time.

Who is this BOF for?

The session is intended for participants with practical interest in building, testing, operating, or evaluating agentic workflows.

  • AI engineers who are building agents or agentic workflows.
  • Evals engineers testing non-deterministic AI output.
  • Platform / observability engineers supporting AI systems.
  • Engineering / product leads deciding how much autonomy to give AI systems.

Bio

Priyanga is a Software Engineer at nilenso. She recently co-built Megasthenes, a library that lets an AI agent answer questions about any codebase with sourced, evidence-backed answers. Lately she’s been rebuilding her own engineering workflows around AI agents.

Reach out to her on LinkedIn.

Recent Blogs

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Get your hybrid access ticket

Hosted by

Jumpstart better data engineering and AI futures

Supported by

Platinum Sponsor

Atlassian unleashes the potential of every team. Our agile & DevOps, IT service management and work management software helps teams organize, discuss, and compl

Platinum Sponsor

Sahaj is an artisanal technology services company crafting purpose-built AI and data-led solutions for businesses.

Gold Sponsor

Skyflow secures the flow of data across datastores, models, and agents. Enterprises turn to Skyflow as their runtime AI data control layer to protect sensitive

Gold Sponsor

Bronze Sponsor

Internet infrastructure APIs for IP geolocation and more

Bronze Sponsor

Open Source Analytical Database for the AI era.

Community sponsor

Real-time Observability & Governance layer for AI agents