Shubham Shrivastav

@shubham_shrivastav

Can an AI agent say "don't ship"? Building a release gate you can trust

Submitted Oct 6, 2026

Session title

An AI agent as a release gate: a ship or don’t-ship call you can trust

One-line summary

An AI agent reviews a pull request, opens its preview build in a real browser, walks the user flows, and ends in one verdict with evidence. This talk shows it live on a real PR and covers the guardrails and cost rules that make that verdict safe to gate a release on.

What problem are you addressing

Teams already let AI write code. The harder question is whether an AI agent can decide if that code ships. A release gate is only useful if you trust it, and an AI agent is easy to distrust: it can misread a screen, run out of time, or sound confident about a check that never ran.

The difficulty:

  • A false pass is worse than no result. If the gate says “ship” when it did not really check, people ship on it. So the agent has to know when to say “I could not tell”.
  • Browser checks are slow and messy. Real pages have custom dropdowns, filters, multi-step forms and loading states. The agent has to act on them and then prove each step worked from screenshots.
  • Many checks run per PR (code review, a browser walk of the flows, and smaller audits). They have to roll up into one call a human can act on, without the cheap checks outvoting the one that matters.
  • Token cost is real, and it does not spread out evenly. One agent holds most of it, and which agent that is changes over time.

The constraints: one verdict per PR (ship, ship with warnings, block, or inconclusive), evidence a reviewer can check, every model call on zero-retention endpoints, and a cost per run low enough to run on every PR.

Why it matters: agents in deployment workflows are coming fast. The question for platform teams is not “can the agent do it” but “what stops it from being confidently wrong”. This talk is one concrete answer.

Intended audience

Platform, DevOps and release engineers who are adding AI agents to CI/CD. Engineering leads deciding whether to trust an agent’s output as a gate. Anyone running LLM workloads who has to watch token spend.

Level

Intermediate. You should know how a CI pipeline and PR preview builds work. No ML background needed.

One or two practical takeaways

  1. Design the gate to fail closed. Make “inconclusive” a first-class verdict. A timeout, a check that did not finish, or a journey that never ran must never turn into a pass, and passing side checks must never outvote a missing main check.
  2. Measure where the tokens go before you optimise, and re-measure. Cost concentrates in one agent and the hot spot moves. Swap a model only after replaying real production inputs side by side, and check what kinds of bugs the new model catches, not just its score.

What will you share

  • Live demo (a real PR, with a recorded run as backup)
  • Architecture and design decisions
  • Production experience
  • Failure modes and debugging
  • Trade-offs
  • Benchmarks (model replay evals)
  • Before/after

Your experience with the problem

Production system, plus hard-earned lessons from building it.

I am the founder of ShipGuarde, an AI release gate. It runs as a live service: a GitHub Action sends a PR, agents review the diff and open the preview build in a real browser (Playwright), a planner turns a plain-language flow into steps, each step is verified from before and after screenshots, and a judge rolls everything into one verdict on the PR with the evidence attached.

Everything in this talk comes from building and testing that system: adversarial test pages built to break the browser agent, replay evals on real production screenshots, side-by-side audits of agent changes, and cost measured from the production database.

What approaches failed, disappointed, or created unexpected problems

These are failure modes found in development and eval, and the design rules they turned into.

  • A faster browser agent that passed its own tests. A DOM-first fast path (read values back instead of a vision check) made form filling much faster, and it passed its tests and a live run. Adversarial side-by-side audits against the previous version found it was less accurate: wrong fields, wrong values, passes it had not earned. I rebuilt it as a small diff on the last known-good version, one fix per observed failure, and turned all 98 audit scenarios into regression tests. Rule: speed ships only when a side-by-side proves it is no worse.
  • Cheap checks outvoting the expensive one. In a parallel stress sweep, a slow browser job waiting for a slot was written off too early, and the remaining audits alone were enough for a verdict. Rule: an audit never outvotes a journey that did not run. Both verdict engines now floor to inconclusive, and a finalized run is immutable, so a late job can never change a result someone has already read.
  • Letting the judge excuse a failure. An LLM judge can explain away a failed step as “an environment issue” in fluent prose. Rule: whether a failure is infrastructure is for code to decide, never the judge’s prose. A ruling like that can lift a block, but the verdict still floors at inconclusive.
  • Hostile pages. I built a set of pages that are correct for a human but hard for an agent: an approvals queue, a legacy WebForms page, a data grid, a wizard and a payment form. They surfaced real failure modes (dropdowns that look like radio buttons, results hidden by an active filter, checks on a counter that had scrolled out of view, keystrokes swallowed while typing) that unit tests never would. Rounds of fixes ended with every flow passing on one build.

What will you do differently today

  • Start with “inconclusive” as a verdict on day one, not something added later. Fail closed by default.
  • Build the adversarial fixtures and the replay harness before the agent, so every change is judged side by side from the start.
  • Log model, tokens and cost per agent execution from the first run, so cost questions are a query, not a guess.
  • Keep the judge out of decisions that code can make.

What trade-offs did you consider

  • Accuracy over speed. A run with 100 steps that takes a few minutes is fine. A fast pass that might be false is not. The rule: when the agent is unsure, it takes the slower, model-checked path.
  • Show the full plan up front vs start faster. Planning page by page started runs about 20 seconds sooner. I kept the full upfront plan, because a reviewer seeing every step before it runs is worth more than the 20 seconds.
  • Cheaper models, but only after a replay. For the per-step verify call I replayed 88 real production steps across 25 models. A cheaper model agreed with the incumbent on 86 of 88 with zero errors, and it went to production. Some cheaper models were rejected because they produced false passes.
  • Build vs buy for the browser. Off-the-shelf test recorders need scripted selectors. The point here is a plain-language flow that survives UI changes, so the planner, grounding and verify loop is built, and the browser itself (Playwright) is bought.
  • Privacy vs model choice. Every model call is restricted to zero-retention endpoints that do not train on the input, enforced in code. That rules out some cheap providers, and it is a hard constraint, not a setting.
  • What I gave up: raw speed, and the cheapest models on the market.

How can this help other practitioners

Anyone putting an agent in a deployment path faces the same questions: when to trust it, how to stop it from being confidently wrong, and how to keep its cost from creeping. This talk gives concrete patterns that work outside my product:

  • a verdict model with “inconclusive” that fails closed
  • immutable finalized results
  • code-owned rules for what counts as infrastructure failure
  • replay evals on real production inputs before any model swap
  • per-agent cost measurement that you re-run, because the hot spot moves

Current state

Production. A live service with a GitHub Action, plus eval and replay harnesses used for every agent change.

Tags

#agents #governance #finops #observability #devops #platformengineering #developerplatforms #demo #casestudy

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy