Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Harsh Joshi
Submitted Oct 7, 2026
Your AI Agent Is a Distributed System
An agent can crash after changing the world but before recording that it succeeded. This session explores how retries, leases, checkpoints, idempotency and durable state turn an LLM-driven workflow into a system that can recover without repeating its side effects.
The difficult part of an agent is often what happens between model calls: a browser interaction, a remote API request, a delayed response, or a worker restart. These operations cross failure boundaries that a conversation transcript cannot resolve on its own.
Consider a browser agent that submits a request. The submission succeeds, but the browser times out before the agent receives confirmation. Retrying might create a duplicate. Skipping the step might abandon a request that never succeeded. Asking the model to reason harder cannot establish which external effect actually occurred.
The session focuses on three engineering questions:
The constraints are practical: browser interfaces without idempotency keys, APIs with incomplete status information, nondeterministic model decisions, rate limits, human approval delays, and work that outlives one process or request.
These problems matter even at low traffic. One duplicated external action can be more consequential than hundreds of failed read-only steps. The goal is recoverable execution with explicit limits, rather than an unqualified promise of exactly-once behavior.
Platform Engineering, SRE, Infrastructure, Developers and Engineering leaders.
Especially engineers moving agents from interactive demos into background workflows, browser automation, internal platforms or operational tools. Familiarity with APIs and asynchronous jobs is useful; prior expertise in model training is not required.
Intermediate, with advanced discussion of uncertain outcomes, worker ownership and recovery semantics.
The demonstration will use a controlled local application so faults are reproducible and no real customer actions are performed. No production benchmark or incident claim is needed to make the failure mechanics visible.
Relevant experience: production systems, open-source projects and agent-tooling experiments.
I work on platform systems at Coupang and previously worked on payment-gateway systems at Visa and as a founding engineer at Plum. That background informs my approach to failure boundaries, durable state and the cost of repeated side effects.
I also build open-source AI tooling, including pixelpi, a browser-agent project, alongside tools for working with coding agents. The session connects those two areas: established distributed-systems reasoning and the execution problems introduced by model-driven workflows.
The agent-specific reference implementation and fault-injection demo are proposed work for this session. This proposal does not claim that the complete architecture has already been deployed at my employer, or attribute a particular production incident to those organisations.
The session will examine the limits of common designs rather than present an unverified personal incident history:
These limitations will be made concrete through controlled failure cases in the reference implementation.
Start by classifying tools as read-only, safely repeatable, externally idempotent or requiring reconciliation. Define action identity and recovery rules before expanding the agent’s autonomy.
Persist the approved intent before execution and the external evidence after execution. Resume from durable workflow state instead of reconstructing progress from a transcript. Make ambiguous outcomes visible, bound retries, and provide a human review path when safe automatic recovery is impossible.
Keep human approval tied to the exact action being executed. A materially changed payload should require a new approval rather than inherit permission from an earlier plan.
The session provides a design review method engineers can apply to any agent framework:
Practitioners should leave able to explain how their agent resumes after a crash, what prevents duplicate effects, and where their guarantees end.
In progress.
The session is grounded in my distributed-systems background and open-source agent tooling. The focused reference implementation and reproducible fault-injection demonstration will be developed for the talk. Any measurements will be reported with their setup and limitations; no unmeasured performance improvement is claimed here.
#agents #distributed-systems #platformengineering #sre #observability #reliability #developerplatforms #durable-execution #idempotency #browserautomation #demo #workinprogress
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}