Utkarsh Kanwat

Your Agent Can Run as Long as You Can Check It

Submitted Oct 3, 2026

Session title
Your Agent Can Run as Long as You Can Check It

One-line summary
Long-running AI agents don’t fail because the model is weak; they fail because nothing outside the agent checks its work. This session is about building that checking layer as platform infrastructure.

What problem are you addressing?
Coding agents can now work for hours on one job, but most teams still babysit them every few minutes. Every step an agent takes builds on its last output, so small mistakes compound over a long run, and the agent’s own “done” is the least reliable signal in the system. Bigger models and bigger context windows reduce the error per step but don’t remove it.

What makes this hard is that the fix is infrastructure, not prompting: tools have to be up and healthy before the agent starts, state has to live outside the context window, work has to happen in isolated workspaces, and every claim of progress has to be checked against the real state of the system. That is a platform problem, and it decides how long a team can safely leave an agent alone.

Audience
Platform engineers, developers and engineering leaders running or planning agents on long tasks.

Level
Intermediate

Takeaways

  • A simple way to estimate how long your agent can safely run unattended: how often you can check its work, and how much each check costs.
  • A checklist for the checking layer: state on disk, health-checked tools, isolated workspaces, checks that read the real system, and a separate judge for “done”.

What will you share?
Architecture and design decisions from an open-source harness for long-running coding agents (github.com/ukanwat/overtime): booting the environment and waiting for the tool server before the agent starts, keeping state in files so each session starts from the record, parallel lanes in isolated workspaces, and validators that compare before and after. Production experience from building coding agents at AutonomyAI. Failure modes from real long runs, and a short live demo.

My experience with this problem
Production system, open-source project, and hard-earned engineering lessons. I build coding agents at AutonomyAI, maintain the open-source harness above, and contributed a warm process pool for environment servers to NVIDIA NeMo Gym.

What approaches failed or disappointed?

  • Trusting the agent’s report. Runs that said “done” were often not done.
  • Counting output. One session produced 586 new assets and looked like success by count; rendering the scene showed pitch-black streets.
  • Tests the agent wrote for its own change tend to confirm what it already did.
  • Squeezing old context instead of handing off to a fresh session from a written record.

About me
Utkarsh Kanwat, AI Research Engineer at AutonomyAI; Electrical Engineering, IIT Bombay. My essay “Why I’m Betting Against AI Agents in 2025 (Despite Building Them)” reached the front page of Hacker News and was covered by The Register and Analytics India Magazine. LeadDev panelist, Sep 2026: https://www.youtube.com/watch?v=bUqFjOgH5ks

Draft slides
https://drive.google.com/file/d/1LeETRXu5U-Jk8mWaTB__Vo9AonfPQiEp/view?usp=sharing (comments on)

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy