Enterprise AI in Production: Mumbai Call for Proposals

Share what it takes to run AI inside an enterprise: architectures, trade-offs, and lessons from production.

Swetha A

Swetha A

@swetha_03

Understanding Your Agents From Traces to Trust

Submitted Oct 2, 2026

Abstract

While building and operating LLM-powered systems in production, one lesson keeps repeating itself which is that the hardest failures are often invisible.
Unlike traditional software, agentic systems can fail in countless ways while every dashboard remains green. An agent may misunderstand user intent, choose the wrong tool, retrieve stale information, drift from its objective, or silently regress after a model upgrade. In many cases, users discover the problem before the engineering team does.

In this talk, we will explore what it takes to truly understand, evaluate, and operate AI agents in production. We will examine why testing and monitoring agents differ fundamentally from traditional software and even classical ML systems, and introduce practical frameworks for identifying failures across comprehension, specification, and generalization.

Drawing from experiences building production agentic systems, we will cover the full evaluation lifecycle: component-level evaluations, workflow and end-to-end testing, synthetic data generation, observability and tracing, safety and alignment checks, human review processes, LLM-as-a-Judge systems, Agent-as-a-Judge architectures, and feedback loops that continuously improve agent behavior.

Along the way, we will discuss the strengths, limitations, and trade-offs of each approach, including how to handle subjective evaluations and measure quality at scale.

Key Takeaways

  • Design effective component-level, workflow, and end-to-end evaluations for production agents.
  • Use synthetic data and scenario generation to test edge cases and failure modes at scale.
  • Build effective observability and tracing to understand what an agent actually did and why.
  • Incorporate human-in-the-loop review, LLM-as-a-Judge, and Agent-as-a-Judge approaches for evaluating agent behaviour.
  • Establish safety, alignment, and regression checks as part of the production lifecycle.
  • Build feedback loops that turn production failures and user feedback into continuous improvements.
  • Develop a practical evaluation and observability strategy that helps teams detect failures before users do.

Target Audience

  • AI/ML Engineers building production-grade agentic systems
  • Software Engineers and Developers working with LLM-based applications
  • Solution Architects and Technical Leads designing multi-agent architectures
  • Engineers working on conversational AI, enterprise assistants, and workflow automation

Anyone moving agentic systems from prototype to production and dealing with real-world failures

Speaker Bio

I’m a Solution Consultant at Sahaj Software. My work spans multi agentic systems and the practical applications of GenAI in engineering.
I enjoy exploring how AI technologies can augment human creativity and decision-making. My research has been presented at the International Conference on Data Analytics and Management, and I’ve spoken at multiple DevDays events and Fifth Elephant conferences, sharing insights on AI agents, AI-assisted software development, and emerging agentic architectures.

Linked In:

https://www.linkedin.com/in/swetha0302

Past talks & workshops:

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

Jumpstart better data engineering and AI futures