Nov 2026
23 Mon
24 Tue
25 Wed
26 Thu
27 Fri
28 Sat 10:30 AM – 01:00 PM IST
29 Sun
Utkarsh Kanwat
Submitted Oct 3, 2026
Coding agents can now work for hours without help, but in most other kinds of work an agent still needs a person to step in every so often. It’s the same models in both cases, so the model isn’t what decides how long an agent can run. The checks around it are. Every step an agent takes builds on its last output, and small mistakes pile up unless something outside the agent catches them. The agent saying it’s done doesn’t count, and neither do tests it wrote for itself.
This talk comes from building coding agents in production at AutonomyAI and from an open-source harness where a coding agent works on one large project for many hours (github.com/ukanwat/overtime). I’ll show what a real check looks like, why coding agents got long runs first, and how to make your own task cheaper to check: state snapshots, invariants, rendered output and a separate judge. One example: a run that produced 586 new assets in a session looked like success by count, and a render check showed the streets were pitch black.
Takeaways
Audience
Engineers and tech leads who are moving agents from prototype to production, and anyone running agents on long tasks.
Bio
Utkarsh Kanwat is an AI Research Engineer at AutonomyAI, where he builds coding agents. He studied Electrical Engineering at IIT Bombay. His essay “Why I’m Betting Against AI Agents in 2025 (Despite Building Them)” reached the front page of Hacker News and was covered by The Register and Analytics India Magazine. He was a panelist at LeadDev in 2026, and has published deep learning research with IIT Bombay and Tata Memorial Centre.
Draft slides
https://drive.google.com/file/d/1LeETRXu5U-Jk8mWaTB__Vo9AonfPQiEp/view?usp=sharing (comments on)
Video
Recent panel at LeadDev (Sep 2026): https://www.youtube.com/watch?v=bUqFjOgH5ks
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}