Mahendra

[Add a catchy title or a Work-in-Progress (WIP) title]

Submitted Sep 29, 2026

How Thousands of Offline-First AI Agents Can Learn Together

One-line summary -
AI is moving from centralized models toward agents that can run local models, maintain local memory and operate independently on laptops, edge devices and disconnected environments. This talk explores the infrastructure needed for thousands of such agents to operate offline, persist what they learn locally, and eventually share knowledge with each other through synchronization.

What problem are you addressing? What real engineering challenge, question, or experience is this session about?

The next generation of AI applications may not always be a request sent to a cloud model. Smaller and increasingly capable models can run locally, allowing AI agents to operate on developer machines, edge devices, private infrastructure and intermittently connected environments.

But running the model locally is only one part of making an agent truly local.

An agent also needs durable state: memories, observations, context, tool results, decisions, user data and application state. If that state lives in a remote database or service, the agent remains dependent on connectivity even if its model runs locally.

A local-first architecture solves that problem by putting the model, agent runtime and persistent state together. SQLite is particularly interesting here because it provides transactional, embedded and serverless storage without requiring another service.

But this creates a new distributed-systems problem at a different scale.

Instead of one central database, we may have thousands or millions of independent local databases — one for each agent, device, application or deployment.

Each agent may discover something useful while operating independently. If agents remain isolated, that knowledge remains isolated. If we want them to learn from each other, we need a mechanism for moving relevant knowledge between them despite intermittent connectivity, failures, retries, conflicting updates and different rates of operation.

The central question of this talk is therefore:

How do we build a population of offline-first AI agents that can operate independently, retain what they learn locally, and collectively benefit from what other agents have learned?

This is not simply a model-training problem. It is a data and distributed-systems problem.

The talk explores a design in which local models and agents operate against durable local databases, changes are captured and synchronized asynchronously, and central infrastructure can consolidate and distribute selected knowledge back to other agents.

Who is the intended audience?

Platform Engineering / SRE / DevOps / Infrastructure / Developers / Engineering leaders

Level:

Intermediate

List one or two practical takeaways.

  • How to architect an AI agent that remains useful and stateful even when disconnected, using a local model and durable embedded state.
  • How to build the data infrastructure for thousands of independent agents to exchange knowledge and collectively improve without requiring continuous connectivity to a central AI service.

What will you share?

The talk will walk through an end-to-end architecture:

Local Model → Agent Runtime → Local Durable State → Synchronization → Shared Knowledge → Other Agents

We will cover:

  • Running AI models locally and what that changes architecturally.
  • Why a local model does not automatically make an application local-first.
  • Using an embedded database such as SQLite as the durable state layer for an agent.
  • Persisting agent memories, observations, context, tool results, decisions and application state.
  • Capturing local changes without requiring application developers to build their own synchronization pipeline.
  • Operating normally when the network is unavailable.
  • Synchronizing local state when connectivity returns.
  • Scaling the architecture from one offline agent to thousands of independent local databases.
  • Moving knowledge from individual agents into shared central infrastructure and eventually back to other agents.
  • Handling retries, duplicate delivery, partial failures, backpressure and schema evolution.
  • Understanding the difference between synchronizing agent knowledge/state and synchronizing model weights.
  • Why collective learning does not necessarily require continuously retraining or synchronizing the underlying model.
  • What information should remain local and what information should become shared knowledge.
  • A live demonstration with multiple local agents operating independently, accumulating state and synchronizing selected information through central infrastructure.

The talk will use SyncLite, an open-source embedded data synchronization project, as the concrete implementation and engineering case study.

What is your experience with this problem? Choose any that apply:

  • Production system
  • Open-source project
  • Experiment or prototype
  • Research / investigation
  • Hard-earned engineering lesson

SyncLite (https://github.com/syncliteio/SyncLite/) is an open-source Apache 2.0 project built around embedded data synchronization.

The project originated from a practical problem: applications and edge systems increasingly have local databases, but organizations still need to consolidate their data into central systems.

SyncLite has been explored through PoCs and production deployments, including deployments handling tens of millions of records per day.

That work led to a broader question: if applications can have durable local state and synchronize that state reliably, could the same infrastructure become a data layer for locally running AI agents?

The local-model, local-state and collective-learning architecture is an ongoing area of experimentation and development.

What approaches failed, disappointed, or created unexpected problems?

The obvious architecture for an AI agent is to put the model and its state in the cloud. The agent sends requests to a remote model and reads/writes its memory in a remote database.

This is operationally convenient, but makes connectivity a fundamental dependency.

Another approach is to run the model locally while keeping memory and application state remote. This provides local inference but not true local autonomy.

A third approach is to build application-specific APIs, queues, streaming systems and synchronization logic. This can work, but pushes distributed-systems complexity into every agent.

There is also a temptation to think about collective learning primarily as synchronizing model weights. That is a very different and significantly more complex problem, and it misses a simpler opportunity: agents can share the experiences, observations, structured knowledge and outcomes they accumulate without necessarily modifying their underlying models.

We also found that the word “synchronization” hides important engineering semantics.

A successful local transaction does not mean that the central system has received the data. A remote destination can fail after local commit. Retries can create duplicates. Different agents operate independently, so there may be no meaningful global ordering of their transactions.

These are distributed-systems problems that become particularly important when the number of agents grows from tens to thousands.

What will you do differently today?

Treat the AI agent as a normal local application.

The model runs locally. The agent runtime runs locally. Its durable state lives locally. The agent should not have to understand whether the network is currently available.

Local transactions are committed normally to the embedded database. A durable change log records what happened. Synchronization happens independently in the background.

When connectivity becomes available, changes can flow toward shared infrastructure. Relevant information can subsequently be made available to other agents.

This separates four concerns that are often coupled together:

  1. Inference — where the model runs.
  2. State — where the agent’s durable memory and application data live.
  3. Synchronization — how changes move between independent instances.
  4. Knowledge sharing — what information is selected and made useful to other agents.

This separation allows us to explore a system where agents remain autonomous locally while the population of agents becomes increasingly knowledgeable collectively.

What trade-offs did you consider?

There are several fundamental trade-offs:

  • Local models vs. cloud models — autonomy and privacy versus centralized compute and capabilities.
  • Local state vs. centralized state — availability and independence versus immediate access to shared state.
  • Eventual consistency vs. global consistency — resilience and scalability versus immediate convergence.
  • Structured state vs. vector-only memory — transactional application state and explicit knowledge versus specialized semantic retrieval.
  • Knowledge synchronization vs. model synchronization — moving useful experiences and data versus synchronizing model parameters.
  • Application simplicity vs. infrastructure complexity — keeping distributed-systems concerns outside the agent versus building more infrastructure underneath it.
  • Selective sharing vs. sharing everything — privacy and data minimization versus collective knowledge.
  • Per-agent replay vs. global ordering — independent autonomous operation versus centralized ordering and coordination.

The architecture deliberately does not attempt to make thousands of agents behave like a single globally consistent distributed transaction system.

The goal is different: allow each agent to remain available and useful locally while providing a reliable mechanism for knowledge to eventually flow between agents.

How can this help other practitioners?

As AI moves toward smaller local models and increasingly autonomous agents, the infrastructure surrounding the model becomes as important as the model itself.

This talk provides a way to reason about that infrastructure.

Instead of thinking of an agent simply as:

Model + Prompt + Vector Database

we can think about a local-first agent as:

Local Model + Durable Local State + Local Experience + Synchronization + Shared Knowledge

This architecture can help practitioners evaluate systems where agents operate on laptops, edge devices, field systems, factories, vehicles, private infrastructure or intermittently connected environments.

It also provides a practical framework for thinking about collective intelligence without requiring centralized execution: thousands of agents can independently observe, act and accumulate knowledge, while synchronization provides the mechanism through which useful knowledge can eventually propagate across the population.

The broader engineering lesson is that once AI becomes distributed, AI infrastructure becomes a distributed data problem as much as a model problem.

Current state - choose one:

Production experience / Open source / In progress

Add tags:

#agents #ai #localai #localmodels #collectiveintelligence #offlinefirst #databases #sqlite #distributed-systems #edgesystems #scalability #opensource #developerplatforms #failurestory #casestudy #demo

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy