Session title
Agents Are Easy, Memory Is Hard
One-line summary
Our SRE agent answers about 95% of the support requests it gets, but when someone corrected it in one thread, the next thread never knew. This talk covers how we gave it a memory we can check and undo, and why the easier versions didn’t work.
What problem are you addressing?
- We run an agent in our platform team’s support channel. People ask it for Grafana or Jenkins access, DNS changes, or help with things that are broken right now, like a stuck rollout, a failing build or an alert that keeps firing. It looks at the live systems with read-only access and replies with what’s wrong, how to fix it and the command to run. The one thing it does on its own is run a small access pipeline. It also picks up PagerDuty incidents. Workers scale with the size of the queue, so up to five requests run at the same time.
- It answers about 95% of what comes in, so getting a reply isn’t the issue. Getting the right one is. When someone corrected the agent, the correction stayed in that thread. A teammate once told it to check every cloud account before saying an IP isn’t ours, and the next question started from scratch.
- We tried the two obvious fixes first. Adding lessons to the prompt means a code change and a deploy each time, and every request then carries every lesson. Pulling in similar old answers brought back stale text that didn’t say it was old, and once the agent quoted a different thread as if it were the current one. Neither approach had a way to take a wrong entry back out.
Who is the intended audience?
SRE, platform engineering and DevOps people, and anyone building or running LLM agents.
Level
Intermediate
List one or two practical takeaways
- Store how to investigate a type of request, not the answers. A checked playbook helps with requests you haven’t seen. A cached answer only helps if the same question comes back, and it goes stale.
- Don’t trust whatever proposes a memory. Re-run each candidate lesson from scratch before you store it, and build the way to remove it at the same time as the way to add it.
What will you share?
- How the agent works. The four ways to reach it (a mention, a channel message, a thread reply, a PagerDuty incident), the request flow, and the autoscaling: at least 2 workers, at most 5, scaled on queue depth. Several workers can run at once safely because replies are idempotent and the queue’s visibility timeout is longer than our time limit.
- How the feedback loop works. Once a thread goes quiet, we draft lessons from it. Each one is checked with a fresh run, and only the ones that pass are saved. Similar questions get them as a hint. A thumbs-down or a contradiction removes a lesson, and only one service is allowed to write.
- What went wrong, with numbers. Corrections that got merged into one lesson, a thorough run that passed a wrong lesson, old lessons hiding new ones, about one in six playbooks being near-duplicates, and a vector index that found about 1 of 190 real matches.
- One lesson from start to finish, from the thread it came from to the answer it later shaped, with the architecture diagrams.
What is your experience with this problem?
Production system, real incidents and failures, internal engineering project, hard-earned lessons.
It’s in daily use. Saved playbooks match roughly three in ten live questions, and we’ve traced wording in real answers back to the exact lesson behind it.
What approaches failed, disappointed, or created unexpected problems?
- Putting every lesson in the system prompt.
- Retrieving raw old threads as memory.
- One merged lesson per correction. One test covered one of four points, so the whole lesson failed.
- Letting the classifier see existing lessons. It proposed fewer new ones.
- Scaling workers on CPU. CPU only goes up while a run is active, not while a request is waiting.
- An approximate vector index on a small table.
- Using the prompt to limit what the agent can write. It once told a user it had saved something to memory when it had no way to.
What will you do differently today?
I’d decide what a lesson is, who can write it and how it gets removed before saving anything. I’d set the similarity threshold from real duplicate pairs instead of guessing. And I’d record from day one which lesson shaped each answer, so a bad rating can point at it.
What trade-offs did you consider?
- Playbooks instead of cached answers: slower on exact repeats, better on new requests.
- A fresh re-run for every candidate: it costs a full investigation, but it doesn’t depend on whatever proposed the lesson.
- Draft pull requests for missing tools, reviewed by a person: slower, but the agent can’t give itself permissions.
- A cap of five workers: lower peak throughput, but cost and load on the systems it queries stay bounded.
- Exact search instead of an approximate index until the table reaches thousands of rows.
How can this help other practitioners?
It’s a worked example of agent memory: what a lesson is, who can write it, how it’s checked and how it’s removed. It also has a warning about measuring. A high response rate says nothing about accuracy, so measure answer quality separately.
Current state
Production experience
Tags
#agents #sre #observability #memory #platformengineering #devops #kubernetes #casestudy #failurestory #governance
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}