Harsh Joshi

[Add a catchy title or a Work-in-Progress (WIP) title]

Submitted Oct 7, 2026

Session title

Your AI Agent Is a Distributed System

One-line summary

An agent can crash after changing the world but before recording that it succeeded. This session explores how retries, leases, checkpoints, idempotency and durable state turn an LLM-driven workflow into a system that can recover without repeating its side effects.

What problem are you addressing?

The difficult part of an agent is often what happens between model calls: a browser interaction, a remote API request, a delayed response, or a worker restart. These operations cross failure boundaries that a conversation transcript cannot resolve on its own.

Consider a browser agent that submits a request. The submission succeeds, but the browser times out before the agent receives confirmation. Retrying might create a duplicate. Skipping the step might abandon a request that never succeeded. Asking the model to reason harder cannot establish which external effect actually occurred.

The session focuses on three engineering questions:

  • How do we distinguish a failed operation from an operation whose outcome is unknown?
  • What must survive a process restart so another worker can continue safely?
  • How do we prevent overlapping workers or retries from repeating irreversible actions?

The constraints are practical: browser interfaces without idempotency keys, APIs with incomplete status information, nondeterministic model decisions, rate limits, human approval delays, and work that outlives one process or request.

These problems matter even at low traffic. One duplicated external action can be more consequential than hundreds of failed read-only steps. The goal is recoverable execution with explicit limits, rather than an unqualified promise of exactly-once behavior.

Who is the intended audience?

Platform Engineering, SRE, Infrastructure, Developers and Engineering leaders.

Especially engineers moving agents from interactive demos into background workflows, browser automation, internal platforms or operational tools. Familiarity with APIs and asynchronous jobs is useful; prior expertise in model training is not required.

Level

Intermediate, with advanced discussion of uncertain outcomes, worker ownership and recovery semantics.

Practical takeaways

  1. Design an agent workflow around durable work items and explicit action states: planned, awaiting approval, executing, succeeded, failed and outcome unknown. Understand where checkpoints help and where reconciliation is required.
  2. Choose retry and recovery policies based on the operation’s side effects: retry safe reads, use idempotency keys where supported, verify uncertain writes, and escalate when external evidence cannot establish the outcome.

What will you share?

  • Architecture and design decisions: Separate the model’s proposed next action from the executor that owns permissions, durable state, action identity and recovery policy.
  • Code and implementation details: A small durable action ledger, worker leases, bounded retries and checkpointed workflow progress. Discuss how fencing prevents stale workers from committing state, and why a lease alone cannot fence an unsupported external service.
  • Failure modes and debugging: Timeouts after successful writes, crashes around checkpoint boundaries, expired leases, overlapping workers, stale browser state and partially completed workflows.
  • Trade-offs and alternatives: A database-backed worker versus a durable workflow engine; coarse checkpoints versus action-level records; automatic recovery versus human review.
  • Operational practices: Trace workflow ID, action ID, attempt number, worker ownership, external receipt and recovery decision. Distinguish model failures from execution failures.
  • Planned live demo: Interrupt a browser-driven workflow immediately after a simulated external submission. Restart the worker, inspect durable state and external evidence, then reconcile the result rather than blindly submit again. Repeat with an ambiguous result to show why the correct state is sometimes “outcome unknown.”

The demonstration will use a controlled local application so faults are reproducible and no real customer actions are performed. No production benchmark or incident claim is needed to make the failure mechanics visible.

What is your experience with this problem?

Relevant experience: production systems, open-source projects and agent-tooling experiments.

I work on platform systems at Coupang and previously worked on payment-gateway systems at Visa and as a founding engineer at Plum. That background informs my approach to failure boundaries, durable state and the cost of repeated side effects.

I also build open-source AI tooling, including pixelpi, a browser-agent project, alongside tools for working with coding agents. The session connects those two areas: established distributed-systems reasoning and the execution problems introduced by model-driven workflows.

The agent-specific reference implementation and fault-injection demo are proposed work for this session. This proposal does not claim that the complete architecture has already been deployed at my employer, or attribute a particular production incident to those organisations.

What approaches failed, disappointed, or created unexpected problems?

The session will examine the limits of common designs rather than present an unverified personal incident history:

  • Treating chat history as execution state: A model’s account of an action is not an authoritative receipt that the external system committed it.
  • Retrying the whole agent loop: Re-running reasoning can produce a different plan while repeating effects from the previous attempt.
  • Checkpointing only after a tool returns: A crash between external success and the checkpoint leaves an unresolved write, even when the checkpoint store itself is reliable.
  • Adding a queue and assuming duplicates are solved: Redelivery improves availability but still requires action identity and deduplication or reconciliation.
  • Treating an expired lease as proof the previous worker stopped: A stalled worker can resume. Internal ownership checks and external side-effect controls are separate concerns.

These limitations will be made concrete through controlled failure cases in the reference implementation.

What will you do differently today?

Start by classifying tools as read-only, safely repeatable, externally idempotent or requiring reconciliation. Define action identity and recovery rules before expanding the agent’s autonomy.

Persist the approved intent before execution and the external evidence after execution. Resume from durable workflow state instead of reconstructing progress from a transcript. Make ambiguous outcomes visible, bound retries, and provide a human review path when safe automatic recovery is impossible.

Keep human approval tied to the exact action being executed. A materially changed payload should require a new approval rather than inherit permission from an earlier plan.

What trade-offs did you consider?

  • Small worker versus workflow engine: A database-backed worker makes the mechanics easy to inspect. A durable workflow engine offers stronger orchestration primitives, but cannot by itself make every external write exactly once. The session uses a small implementation to explain the boundary, not to advocate rebuilding a workflow platform.
  • Fine-grained state versus implementation complexity: Action-level records add storage and bookkeeping, but make recovery and auditing more precise. Read-only exploration can use coarser checkpoints; external writes need greater care.
  • Availability versus duplicate prevention: Automatically retrying every timeout may finish more workflows while creating duplicates. Pausing uncertain actions reduces immediate completion in exchange for explicit control over side effects.
  • Deterministic execution versus flexible planning: The model remains free to propose useful actions. Once an action is approved and recorded, retries should preserve its identity and payload rather than ask the model to reinvent it.
  • Generality versus external guarantees: Supported APIs can provide idempotency keys and status queries. Browser workflows may provide neither. Recovery policy must reflect that difference instead of claiming a universal guarantee.

How can this help other practitioners?

The session provides a design review method engineers can apply to any agent framework:

  • Identify every point where an external effect and a local state update can diverge.
  • Define what evidence establishes success and what remains unknowable.
  • Decide which actions may be retried, reconciled or escalated.
  • Test recovery by interrupting execution at those boundaries.
  • Evaluate agent frameworks and workflow engines by their execution guarantees, not only their planning quality.

Practitioners should leave able to explain how their agent resumes after a crash, what prevents duplicate effects, and where their guarantees end.

Current state

In progress.

The session is grounded in my distributed-systems background and open-source agent tooling. The focused reference implementation and reproducible fault-injection demonstration will be developed for the talk. Any measurements will be reported with their setup and limitations; no unmeasured performance improvement is claimed here.

Tags

#agents #distributed-systems #platformengineering #sre #observability #reliability #developerplatforms #durable-execution #idempotency #browserautomation #demo #workinprogress

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy