Gaurav Maheshwari

@mgaurav

Surviving the query-pocalypse!

Submitted Sep 26, 2026

Proposal template

We have a new user persona - AI Agents. What does it mean for platform availability, reliability & performance?

We have been building systems for human as a user for a long time, which led to emergence of patterns to measure availability, reliability, SLIs/SLOs/SLAs in a certain manner. For example, a user seeing a page loading for more than 30s judges the application to be too slow and may churn. Nowadays, more and more consumption of any application is moving to AI agents consuming the application via CLI/MCP and humans reading the final summary agents write. Combine that with agents relentless-ness of getting the task done, and forgetful-ness of what they did in last session, introduces a set of new challenges in application design:

  • An expensive API is not distinguishable from a fast API for agents.
  • Agents exercise various query shapes, they will keep trying what gets their task done.
  • Agents can learn to use inefficient query shapes (memory layers), even if application improves, older query shapes could continue arriving.
  • Agents don’t report bugs, just workaround issues.

In a way, AI agents makes Hyrum’s law become applicable for your application at much sooner rate. This talk covers insights of how agent queries differ from human queries taking examples from a real world Observability platform and what we have done in the platform from availability & reliability point of view. I believe this is a relevant topic for many applications whose primary user is becoming AI agents.

A rough set of items I have in mind are:

  1. What to monitor? - APIs returning empty results (agents got params wrong?), wider queries (.*, wide time ranges), errors (obviously!)
  2. Error messages - give hints in error messages so that Agents can correct themselves. On that note, give performance hints as well!
  3. Consider fast path and slow path based on the source of the query - is a human waiting for the query or it is an AI agent? E.g. we route queries to lambda for higher parallelization and performance, do we use lambdas for AI agent originated queries.

Intended Audience

Platform Engineers, Application Engineers, SRE/DevOps

Level

Beginner (easy to understand talk)

Takeaways

  1. Design applications considering AI agents as the user.
  2. Monitor human usage v/s AI usage differently.

{What will you share?

  • Architecture or design decisions
  • Production experience
  • Failure modes and debugging

{What is your experience with this problem?

  • Production system
  • Hard-earned engineering lesson

{What approaches failed, disappointed, or created unexpected problems?}

{What will you do differently today?}

{What trade-offs did you consider? For example,

  • Why one architecture over another?
  • Why build vs. buy?
  • What did you optimize for?
  • What did you knowingly give up?
  • What alternatives did you consider?}

{How can this help other practitioners?

  • A design approach
  • A debugging technique
  • A pattern to adopt
  • A mistake to avoid
  • A way to evaluate competing approaches
  • A set of operational practices
  • A new way of thinking about the problem}

{Current state - choose one:

  • Production experience
  • In progress
  • Experimental/prototype
  • Open source
  • Retrospective/lessons from a previous system}

{Add tags to help us understand your proposal. Tags may describe domains, techniques, systems, or formats, for example: #observability #inference #agents #security #governance #finops #sre #devops #platformengineering #mlops #databases #scalability #kubernetes #developerplatforms #failurestory #demo #workshop #casestudy #workinprogress}

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy