Ravi Kiran Vemila

@raviknacklabs

Harness or ceremony: what held when agents ran unattended in production

Submitted Aug 21, 2026

Description

Every client project at KnackLabs wanted the same thing: an agent that works like a coworker. It knows who it is talking to, whether the message comes from Slack, Teams, or inside the product. It remembers past conversations with that person. It has the access an employee would have, and only that.
So it became Gantry: a self-hosted agent runtime that teams embed through an SDK, call over an API, or reach from the chat tools they already use, with voice planned next.

This talk goes through the harness pieces that held in production and the ones that turned out to be ceremony. Cases include an LLM permission
classifier that cancelled a tool call at 08:53 and allowed the same call at 10:01, jobs that pause on a denial and resume after a crash, and audit
records that still make sense weeks later. Each case covers what was tried, how it failed, and what replaced it.

Takeaways

Build the agent like you would onboard a coworker: one identity across every channel, memory that follows the person rather than the chat window, and access granted up front rather than asked for at each step. Put all of that in the runtime, not the prompt, so it holds when nobody is watching.

Audience

Engineers and founders running agents in production or getting there. No ML background needed.

Bio
Ravi Kiran Vemula, VP of Engineering at KnackLabs. Builds Gantry, an MIT-licensed agent runtime being prepared for open-source release, and runs production agents on it for enterprise client work.

Presentation
https://docs.google.com/presentation/d/1f2VkonSMdkPsM112jUCj7ZqRyNvEFjxwSNlfg8RE1nA/edit?usp=sharing

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

Jumpstart better data engineering and AI futures