Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Kushagar Sharma
@NhkWolf
Submitted Oct 5, 2026
One-line summary
How we built and productionized Hermes for a customer in the AI-powered industrial operations and fleet intelligence space: an infrastructure agent that investigates real systems, reasons about changes, and creates production-ready GitOps PRs while keeping execution deterministic, auditable, and safe.
What problem are you addressing?
Making an AI agent actually useful for infrastructure is very different from building an impressive demo. Giving an LLM powerful production tools is easy; making it safe and reliable enough that engineers can genuinely delegate infrastructure work to it is much harder.
Hermes uses Slack as its interface, read-only infrastructure tools for investigation, Vault-backed secrets, and Git as the mutation boundary. It can independently understand requests, inspect infrastructure, reason about changes, and create PRs without being able to silently mutate production.
Intended audience
Platform Engineering / SRE / DevOps / Infrastructure / Engineers building agents
Level
Intermediate / Advanced
Practical takeaways
What will you share?
The production architecture behind Hermes: Slack → agent → tools → GitOps control plane, including read-only investigation, Vault-backed secrets, PR-only mutations, audit trails, reconciliation loops, and automated rightsizing.
I’ll also show examples of Hermes taking natural-language infrastructure requests, investigating the system, reasoning about the required change, and turning them into correct, reviewable infrastructure PRs.
Experience with this problem
Production system / Internal engineering project / Hard-earned engineering lesson.
Hermes handles real infrastructure workflows through the same GitOps processes used by engineers, making the agent part of the existing engineering control plane rather than creating a separate, privileged AI execution path.
What failed or disappointed?
Direct agent write access creates unnecessary blast radius and makes auditing and recovery harder. We also learned that making everything agentic is counterproductive. Deterministic problems should stay deterministic.
What would you do differently today?
Start by defining trust boundaries: what can be probabilistic, what must remain deterministic, and where human approval belongs.
Trade-offs considered
We deliberately traded some autonomy for safety. Hermes can investigate extensively and propose changes autonomously, while infrastructure mutations still flow through Git, review, and deterministic reconciliation.
How can this help others?
It provides a practical architecture for moving infrastructure agents beyond demos and into real production engineering workflows:
Let agents investigate, reason, and propose. Keep execution deterministic.
Current state
Production experience
Tags
#agents #sre #platformengineering #gitops #kubernetes #llmops #security #governance #casestudy
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}