Ayushman Bhattacharya

Ayushman Bhattacharya

@elixpo

What Should an AI Search System Remember, Reuse, and Recompute?

Submitted Oct 3, 2026

Session title

What Should an AI Search System Remember, Reuse, and Recompute?

One-line summary

An AI search system does not have one cache problem. Conversation state,
semantically repeated questions, and previously processed sources have different
scopes, lifetimes, correctness risks, and reasons to be recomputed.

What problem are you addressing?

The first version of our open-source answer engine connected browser-based web
search directly to provider-routed LLM synthesis. It worked, but sessions lost
earlier context, rephrased questions reran the complete pipeline, and popular
URLs were repeatedly embedded across sessions.

Calling all of this “caching” hides three different boundaries. Conversation
history belongs to one session. A semantically similar answer may be trustworthy
only briefly and in the right context. A URL embedding can be shared, but only
while it still represents the source. Combining them makes privacy, freshness,
invalidation, and failure handling difficult to reason about.

This session follows a request through the three-layer design we built: a
Redis-backed session context window with disk overflow, a session-scoped
semantic query cache, and a shared URL embedding cache. It examines scope, keys,
TTLs, similarity thresholds, bypass rules, and recovery on one commodity CPU
server. Final answer synthesis uses remote inference, and the talk will keep
that boundary explicit.

Intended audience

Platform engineers, SREs, infrastructure engineers, and developers building or
operating AI search, RAG, agent, or retrieval-backed systems.

Level

Intermediate

One or two practical takeaways

  1. A model for separating conversational state, semantic response reuse, and
    source-level computation by scope, freshness, and failure cost.
  2. A checklist for choosing keys, TTLs, bypass rules, invalidation,
    observability, and persistence without creating one global cache.

What will you share?

  • The request lifecycle and boundaries among Redis databases, a hot conversation window, disk overflow, semantic matching, and shared URL embeddings.
  • The evolution from an uncached pipeline and lossy history truncation to the current architecture.
  • Results from an evaluated 8-vCPU deployment: 89.3% aggregate Redis keyspace hit rate, approximately 0.1 ms Redis reads, and a 1.38 MB measured Redis footprint. The hit rate does not mean 89.3% of questions avoided an LLM call.
  • Failure and bypass cases involving stale answers, context-sensitive questions, changed sources, and unavailable hot storage.
  • Open-source code, an arXiv preprint, a live system, and an evaluation dataset.

What is your experience with this problem?

  • Production system
  • Open-source project
  • Research / investigation
  • Hard-earned engineering lesson

I co-developed and operate OreoLook, the open-source answer engine used as the case study, and co-authored the arXiv preprint documenting its architecture and evaluation. Its live Hugging Face Space and evaluation dataset are public.

What approaches failed, disappointed, or created unexpected problems?

The original pipeline had no cache or persistent session layer. It lost context,
repeated the full pipeline for paraphrases, and recomputed URL embeddings.

A rolling-window policy later kept hot memory bounded by discarding older
messages, but failed when someone returned and referred to an earlier result. We
also considered one Redis namespace for every cache concern, then rejected it
because the three forms of state needed different scopes, TTLs, monitoring, and
failure handling.

Zlib compressed small conversation archives better, but a pure-Python canonical
Huffman codec kept deployment simpler. We optimized for dependency simplicity,
not the best compression ratio.

What will you do differently today?

I would define ownership, freshness, privacy, and failure boundaries before
selecting storage. I would also add per-layer measurements from the first
deployment rather than trying to interpret aggregate Redis statistics later.

I would treat retention, encryption, and deletion as initial design decisions.
Compression saves storage; it does not protect sensitive conversation data.

What trade-offs did you consider?

  • Separate Redis databases improve monitoring, selective flushing, and failure isolation, but add configuration and are not a security boundary.
  • Disk overflow preserves resumable conversations with bounded Redis memory, but adds retention, recovery, and data-protection responsibilities.
  • Huffman simplified deployment, while zlib compressed our small archives 5 to 10 percentage points better.
  • A bounded linear semantic scan avoided another service, but would not suit an unbounded global index.
  • Short TTLs and session scoping reduce reuse but lower the risk of stale or cross-context answers.

How can this help other practitioners?

The same decision order applies to AI search, RAG, and agent systems: identify
what is reused, who may reuse it, how long it remains valid, what evidence must
travel with it, and what happens when it is wrong or unavailable.

Attendees will receive a three-scope architecture, a request-lifecycle diagram,
and a checklist for deciding whether a result belongs in conversation state, a
semantic cache, a source-processing cache, or should be recomputed.

Current state

Production experience

Tags

#platformengineering #mlops #databases #latency #reliability #caching #casestudy #failurestory

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy