Over the past few months we’ve been evaluating AI SRE agents with several of our customers: OpenSRE, HolmesGPT, K8sGPT and a handful of others. We started with the things any SRE would check before trusting one of these on a real incident. That only took us so far, so we built a benchmark on top of the open-source SREGym framework, where the agent’s diagnosis is graded against what actually broke. Out of that work came a plug-and-play framework for evaluating AI SRE agents, which we’ll share.
Everything ran on a production-grade setup, so the results show how these agents behave under real conditions, including where they fall short. You’ll leave with numbers, both qualitative and quantitative, on what to expect from these agents in a real production environment.
Teams adopting AI SRE agents need answers to questions that get harder as the agent gets closer to on-call. We set out to answer questions like does it find the right root cause, does it say so when it can’t see something, what does an investigation cost, what access does it need, and does it know anything about our systems, building security guardrails at the agent level through the Agent gateway.
What problems we faced during the experiments:
- AI agents sometimes did not failed loudly and were difficult to debug. A broken integration can sometimes produces a confident root cause, and an investigation that is going wrong looks the same as one that is going right until we look at it from production lense.
- Continously hydrating the knowledge base so that agent does not need to cold start everytime it starts a new investigation.
To compare approaches fairly, we built one reference setup to showcase our learnings:
- iximiuz Labs challenges for the first, typed-prompt experiments as a preliminary analysis.
- A kind cluster with kube-prometheus-stack and demo microservices under load, where real alertmanager alerts flood start the investigations.
- SREGym for the later ones, which for the fault injection and anlysis into a live system and has an LLM judge grading the diagnosis.
- How different model effects the diagnosis.
- Tekenomics of the incidence resolution.
- Platform Engineering, SRE, and DevOps engineers who run or are about to run AI SRE agents in production.
- Engineering leaders deciding whether to bring AI into the SRE path, whether for incident responses, and assessing quality RCA backed by evidences.
Intermediate. Assumes basics of Agentic loops, MCPs, LLM tool callings
A comparison of different AI SRE tools, design considerations when building your own AI SRE agents, cost control patterns, how to evaluate these AI SRE agents for production quantitatively.
- How the open-source AI SRE agents compare with each other, including HolmesGPT, OpenSRE and K8sGPT: their architecture, strengths, weaknesses and limitations.
- The design trade-offs between general-purpose coding agents like Claude Code or Codex (with MCPs) and purpose-built AI SRE agents for solving production problems, along with our qualitative and quantitative results from the evaluation framework.
- A demo of HolmesGPT and K8sGPT running on the SREGym evaluation framework, and a walkthrough of the tokenomics: where the tokens go and what each investigation costs.
- How SREGym evaluates agents, how the LLM-as-a-judge scoring works, and the benchmark numbers we got.
The open source repo has locally runnable kubernetes and docker setups accompanied by demo screenshots which can be used during the talk.
https://github.com/one2nc/ai_sre
- Research / investigation
- Hard earned engineering lesson
- Broken integrations produced confident answers.
- During some of the evaluations, even when the tool integrations were broken the agent still produced a confident looking result.
- A stronger model did not help, because the missing piece was context.
- Our own tests made the agent look better than it was.
- When we picked the faults using our local prometheus, grafana stack on a kind cluster, HolmesGPT got 21 of 21 root causes right.
- SREGym’s decoys and noise brought that to 8 to 10 of 21. The agent latched onto the first anomaly and spent the rest of the run defending it.
- Guarding the API calls to upstream systems during investigation
- How to secure/rate limit upstream APIs for e.g. Kubernetes control plane if agents get stuck in a loop and keep making expensive API calls to fetch the information.
A way to run and evaluate AI SRE agents for production scale.
#observability #agents #Tokenomics #sre #platformengineering #AI SRE #finops
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}