The Fifth Elephant 2026 Winter Edition

Call for Data Problems and Proposals

Utkarsh Kanwat

Your Agent Can Run as Long as You Can Check It

Submitted Oct 3, 2026

Submission type: Talk/session proposal - 30-40 mins

About the session
An AI agent that runs for hours produces a lot of data about its own work: logs, status reports, test results, counts of things it built. Most of that is the agent describing itself, and it is the least reliable evidence you have. Every step builds on the last output, so small errors compound, and the agent’s own “done” says little about whether the work is right. The question this talk is about is an evaluation question: what evidence actually tells you an agent’s long run worked, and what only looks like it does.

I’ll share what I learned measuring long runs of coding agents, at AutonomyAI and in an open-source harness where an agent works on one large project for many hours (github.com/ukanwat/overtime). Including what failed: one session produced 586 new assets and looked like success by every count, and rendering the result showed pitch-black streets. Tests the agent wrote for itself tended to confirm what it had already done. What worked better was evidence that reads the real state of the system: snapshots before and after, invariants, rendered output, and a separate judge with fresh context. I’ll end with what I still can’t measure well.

Takeaways

  • Which evidence about an agent’s work is trustworthy (it reads the world) and which isn’t (it reads the agent’s account).
  • A simple way to estimate how long an agent can safely run unattended: how often you can check it, and what each check costs.

Audience
Data scientists, ML engineers and evaluation folks working with LLM agents, and anyone deciding whether to trust an agent’s output.

Bio
Utkarsh Kanwat is an AI Research Engineer at AutonomyAI, where he builds coding agents. He studied Electrical Engineering at IIT Bombay. His essay “Why I’m Betting Against AI Agents in 2025 (Despite Building Them)” reached the front page of Hacker News and was covered by The Register and Analytics India Magazine. He was a LeadDev panelist in 2026 (https://www.youtube.com/watch?v=bUqFjOgH5ks) and has published deep learning research with IIT Bombay and Tata Memorial Centre.

Draft slides
https://drive.google.com/file/d/1LeETRXu5U-Jk8mWaTB__Vo9AonfPQiEp/view?usp=sharing (comments on)

What I don’t know yet
How to check work that has no cheap ground truth, like research or design, where a render or a test doesn’t exist. I’d welcome critique and ideas on that.

Match-making tags

  • I can help with: experience, technique, critique
  • I need help with: ideas, critique
  • I’d like to meet: practitioners, tool builders, researchers

Topic tags: evaluation, safety, AI agents

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

Jumpstart better data engineering and AI futures