Your on-call agent has amnesia: building AI that learns from production incidents
Every AI agent run starts from zero, so a correction an engineer gives during one incident is gone by the next. This talk covers how we are building a learning loop into a self-hosted enterprise AI platform, tested first on incident response, and how an agent’s track record should decide when it may fix things on its own.
On-call engineers get better with every incident. They learn which alert always follows a deploy of one service, which dashboard misleads during batch jobs, and which query to run first. AI agents don’t. Each investigation starts from the same blank context. When an engineer corrects the agent (for example, “it’s not the database, it’s the retries from checkout”), the correction is lost when the session ends.
Anyone who works with AI runs into this. You explain something once, and next session you explain it again. We ran into it building Operate, a self-hosted enterprise AI platform whose first solution is DevOps/SRE. Production incidents are where it hurts most, because acting on a wrong memory is expensive.
What makes it hard:
- Incidents are rare and each one is different. One incident is weak evidence, and an agent that generalises from it learns superstition.
- Knowledge goes stale when the code or infrastructure changes. A root cause found in March can be wrong after a refactor in June, so a memory should expire when the thing it describes changes.
- The inputs are untrusted. Logs, alert payloads and issue text are written by systems, and sometimes by attackers. If an agent learns from what it reads, a crafted log line becomes a lasting instruction. It’s prompt injection with a save button.
- Putting every past run into the context raises cost and makes answers worse.
- Memory has to live in the platform. If each agent keeps its own notes, every new agent starts with amnesia and each needs its own review, audit and deletion. If the platform keeps them, lessons have to stay scoped so that one agent’s lesson doesn’t steer another agent’s work.
- Learning and acting compound. An agent that remediates on its own based on a wrong lesson does damage faster than one that only suggests.
This matters because teams are giving agents production access now. If they keep repeating the same mistakes, they add toil instead of removing it. Enterprises running AI in their own infrastructure also need that learning to be reviewable and auditable.
SRE, Platform Engineering, DevOps, and engineering leaders who are building or evaluating AI agents for operations, especially those running AI inside their own infrastructure.
Intermediate
- Put learning in the platform, on the run record, so every agent gets it. A lesson is a small, typed record. It carries provenance (the run it came from and the person who confirmed it), a scope (agent, service and environment), and a trigger that invalidates it. Lessons are reviewed like code, versioned, auditable and easy to revert.
- A trust ladder for self-healing. The rungs are observe, suggest, act with approval, and act and report. Each type of action climbs a rung only after it builds a good track record, and drops after one bad outcome.
- Architecture and design decisions: the full loop. A run produces a reflection, the reflection produces a candidate lesson, a human reviews it, the next run retrieves it, and the outcome feeds back. It is built on the platform’s run record, so any agent can use it. The talk covers why we separate what happened (episodes), what is true about the system (facts) and what to do (procedures), and which of these the agent may write without review.
- Failure modes and debugging: memory poisoning through log injection, stale lessons after deploys, overfitting to one incident, and context bloat. For each, how it shows up and how we defend against it.
- What failed for us: runbooks injected into prompts, and a fixed investigation pipeline. We built both, removed both, and rebuilt a single-purpose AI SRE into a platform. That rebuild is what made learning possible.
- Measurements: we replay seeded failures against the same system with memory on and off, and compare time to root cause, tool calls, tokens and whether the agent got the answer right.
- Live demo: an agent investigates the same type of failure twice, and the second time it uses what an engineer taught it the first time. Then a deploy changes the code that lesson was about, and the lesson is invalidated. Finally, the trust ladder for one action, opening a fix PR: the agent goes from attaching a patch, to opening the PR after approval, to opening it on its own and reporting in Slack. One bad PR drops it back a rung. A human always merges.
The examples come from Operate, the self-hosted enterprise AI platform we build. A company runs it in its own infrastructure, connects its tools once (Datadog, databases, GitHub, GitLab) and runs agents that each do one job, with every run recorded. DevOps/SRE is its first solution, with agents that investigate Datadog errors and GitHub/GitLab issues. The talk includes no product walkthrough, and the design applies to any agent stack.
- Production system: Operate runs in production with two pilot clients. Its DevOps/SRE agents investigate Datadog errors and GitHub/GitLab issues inside each client’s own infrastructure. More than 80 open-source repositories also use it to investigate GitHub issues. Every run is recorded with a full timeline.
- Internal engineering project: we recently rebuilt Operate from a single-purpose AI SRE into a platform of tools, agents and recorded runs.
- Hard-earned engineering lesson: our first approach to giving the agent knowledge failed, and we removed it (details below).
- Experiment or prototype: the learning loop. We’re building it now and will present it with measured results.
- Runbooks as prompt text. Our first version let admins write “skills”: editable runbooks that were inserted into the agent’s prompt. They were static text written for one investigation pipeline. Nothing in the system updated them after an incident, even though that’s when the knowledge is freshest, so keeping them current depended on someone remembering to do it. We removed them. Our takeaway was that knowledge someone has to remember to update isn’t really in the system.
- A fixed investigation pipeline. Every incident went through investigation, verification and resolution stages in that order, run by agents built only for that pipeline. Any new job would have needed a new pipeline. We replaced it with a platform where tools are connected once, agents ship as versioned packages that each do one job, and every execution is recorded as a run. That run record turned out to be the raw material for learning.
- Breadth before depth. We wired up twelve alert sources through webhooks, and almost none of them were used. We cut back to one source and built it properly.
- Build the platform primitives (tools, agents, runs) first and the domain solution on top. Then learning, audit and access control are built once for every agent.
- Treat the run record as the source of truth from day one, and derive knowledge from it rather than asking people to write knowledge separately.
- Capture corrections when they happen. An engineer disagreeing with the agent is one of the best learning signals you get.
- Design memory like a codebase, with proposals, review, versions and rollback, before designing retrieval.
- Grant autonomy one type of action at a time.
- Platform memory vs per-agent memory. One store, one review flow and one audit trail serve every agent. In return, every lesson needs a scope (agent, service, environment) so that a lesson from incident response doesn’t steer an unrelated chat answer. We gave up the simplicity of each agent owning its own notes.
- Retrieval of explicit records vs fine-tuning. Explicit records can be inspected and reverted, and they work with any model, including the self-hosted models that enterprise installs often require. We gave up whatever implicit generalisation fine-tuning might offer.
- Structured lessons vs retrieving raw transcripts. Structured lessons add a reflection step to every run, which costs tokens and time. What we get is provenance, scope and deletion. Raw transcripts are cheap to store and hard to invalidate, and similar-looking incidents are often different underneath.
- Human review vs autonomous learning. Review slows learning and adds work. The agent records episodes freely and proposes facts and procedures, but anything that can change an action needs a person’s approval. We gave up speed for safety.
- Calendar expiry vs change-based invalidation. Invalidating on change needs links to deploys and commits, which is more plumbing. But a year-old lesson about a service that hasn’t changed is still correct, and a day-old lesson about a service that was just refactored may not be.
- Build vs buy. General-purpose agent memory libraries are designed for chat personalisation. Incident knowledge needs service scoping and change-based invalidation, and in a self-hosted enterprise install it has to stay inside the customer’s infrastructure.
- A design approach: the lesson record (provenance, scope, invalidation trigger) and the review flow around it. It works with any agent framework and belongs in the platform layer.
- A pattern to adopt: a trust ladder that grants autonomy per type of action, based on track record.
- Mistakes to avoid: maintaining runbooks by hand as prompt text, building memory separately into each agent, and letting an agent learn from untrusted input without review.
- A way to evaluate approaches: replay seeded failures with memory on and off, and measure correctness as well as speed.
In progress. The platform (tools, agents, recorded runs) is in production with two pilot clients running its DevOps/SRE agents, and more than 80 open-source repositories use it. The learning loop is being built now, and we’ll present it at the conference with measured results, including whatever didn’t work.
#agents #sre #incidentresponse #observability #selfhosted #enterpriseai #governance #memory #selfhealing #security #platformengineering #failurestory #demo #workinprogress
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}