Manas Chaturvedi

@manaschaturvedi

From AI Agents to AI Platform: Lessons From Building an In-House Multi-Agent Platform

Submitted Oct 4, 2026

From AI Agents to AI Platform: Lessons From Building an In-House Multi-Agent Platform

One-line Summary

Building AI agents is relatively easy; operating them reliably in production is not. This talk shares how we evolved from standalone AI agents to a production-grade multi-agent platform that emphasizes deterministic execution, observability, evaluations, cost control, and continuous improvement.


What problem are you addressing? What real engineering challenge, question, or experience is this session about?

As AI agents became integral to our production workflows, we quickly discovered that the agents themselves were only a small part of the engineering challenge. The real complexity lay in orchestrating long-running workflows, managing failures, maintaining state, controlling costs, evaluating quality, and ensuring predictable behavior across hundreds of executions.

We evaluated several popular agent frameworks and SDKs, but found that our production requirements - deterministic execution, strict SLAs, resumability, execution transparency, operational observability, and platform-level governance - required capabilities beyond what generic frameworks could provide. This led us to build an in-house multi-agent platform that provides orchestration, standardized execution semantics, guardrails, model abstraction, and shared infrastructure for AI workloads. The platform is intentionally orchestrated rather than fully autonomous because predictability matters more than autonomy in production.

More recently, we extended the platform with a feedback loop that continuously evaluates agent performance, generates improved skill definitions, validates them through A/B testing, and safely promotes better versions - moving from simply running agents to enabling continuous improvement of agent capabilities.

This session focuses on the engineering lessons learned while building and operating this platform, rather than on prompt engineering or LLM fundamentals.


Who is the intended audience?

  • Platform Engineers
  • Infrastructure Engineers
  • DevOps Engineers
  • Site Reliability Engineers (SREs)
  • Backend Engineers
  • AI Platform Engineers
  • Engineering Managers and Tech Leads responsible for production AI systems

Level

Intermediate to Advanced


Practical Takeaways

  • Understand the anatomy of a production-grade multi-agent platform.
  • Learn how to design a production-ready multi-agent platform that prioritizes determinism, reliability, scalability, and operational simplicity over autonomous agent behavior.
  • Learn how evaluations, observability, cost controls, and self-improving feedback loops can be integrated into an AI platform to continuously improve agent performance while maintaining production safety.

What will you share?

This talk will cover:

  • The evolution from standalone AI agents to a reusable internal AI platform.
  • Why we chose to build our own multi-agent platform instead of adopting existing agent frameworks.
  • The anatomy of a production-grade multi-agent platform, including orchestration, execution, state management, guardrails, tool integration, and model abstraction.
  • Design decisions around deterministic execution, checkpointing, retries, failure recovery, queues, DAGs, and long-running workflows.
  • Building evaluation pipelines to measure agent quality instead of relying on intuition.
  • Cost optimization techniques for operating AI agents at scale.
  • Operational lessons around debugging, observability, governance, and production reliability.
  • Designing a self-improving agent feedback loop where execution traces are evaluated, candidate skills are generated, validated through A/B testing, and promoted only after demonstrating measurable improvements.

What is your experience with this problem?

This platform powers production AI workflows inside our Risk Platform and has evolved from independent AI agents into a shared execution platform supporting multiple AI-driven use cases.


What approaches failed, disappointed, or created unexpected problems?

Our initial assumption was that building AI agents was the primary challenge. In reality, the operational complexity emerged only after deploying them into production.

Some of the challenges we encountered included:

  • Agent implementations diverged across teams, making them difficult to standardize and maintain.
  • Generic agent frameworks provided useful abstractions for building agents but did not fully address our operational requirements around deterministic execution, resumability, observability, governance, and execution semantics.
  • Long-running workflows required checkpointing, state persistence, and resumability to avoid restarting entire workflows after failures.
  • Debugging AI workflows without structured execution traces and evaluation pipelines became increasingly difficult.
  • Prompt changes and skill updates often required manual experimentation without an objective mechanism to measure whether agents had actually improved.
  • As adoption increased, cost management, retry policies, and execution guardrails became platform-level concerns rather than application-level concerns.

These challenges fundamentally changed how we viewed AI systems. Instead of treating agents as isolated applications, we began treating them as workloads running on a production platform - requiring the same engineering discipline around reliability, scalability, observability, governance, and continuous improvement that we apply to any other production infrastructure.

Speaker’s Bio

Manas Chaturvedi is a Director-Senior Technical Architect in IDfy and has about 11+ years of experience in building distributed systems, AI-powered workflows and internal developer tools in various startups and product-based companies.

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy