Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Submitted Oct 3, 2026
What Should an AI Search System Remember, Reuse, and Recompute?
An AI search system does not have one cache problem. Conversation state,
semantically repeated questions, and previously processed sources have different
scopes, lifetimes, correctness risks, and reasons to be recomputed.
The first version of our open-source answer engine connected browser-based web
search directly to provider-routed LLM synthesis. It worked, but sessions lost
earlier context, rephrased questions reran the complete pipeline, and popular
URLs were repeatedly embedded across sessions.
Calling all of this “caching” hides three different boundaries. Conversation
history belongs to one session. A semantically similar answer may be trustworthy
only briefly and in the right context. A URL embedding can be shared, but only
while it still represents the source. Combining them makes privacy, freshness,
invalidation, and failure handling difficult to reason about.
This session follows a request through the three-layer design we built: a
Redis-backed session context window with disk overflow, a session-scoped
semantic query cache, and a shared URL embedding cache. It examines scope, keys,
TTLs, similarity thresholds, bypass rules, and recovery on one commodity CPU
server. Final answer synthesis uses remote inference, and the talk will keep
that boundary explicit.
Platform engineers, SREs, infrastructure engineers, and developers building or
operating AI search, RAG, agent, or retrieval-backed systems.
Intermediate
I co-developed and operate OreoLook, the open-source answer engine used as the case study, and co-authored the arXiv preprint documenting its architecture and evaluation. Its live Hugging Face Space and evaluation dataset are public.
The original pipeline had no cache or persistent session layer. It lost context,
repeated the full pipeline for paraphrases, and recomputed URL embeddings.
A rolling-window policy later kept hot memory bounded by discarding older
messages, but failed when someone returned and referred to an earlier result. We
also considered one Redis namespace for every cache concern, then rejected it
because the three forms of state needed different scopes, TTLs, monitoring, and
failure handling.
Zlib compressed small conversation archives better, but a pure-Python canonical
Huffman codec kept deployment simpler. We optimized for dependency simplicity,
not the best compression ratio.
I would define ownership, freshness, privacy, and failure boundaries before
selecting storage. I would also add per-layer measurements from the first
deployment rather than trying to interpret aggregate Redis statistics later.
I would treat retention, encryption, and deletion as initial design decisions.
Compression saves storage; it does not protect sensitive conversation data.
The same decision order applies to AI search, RAG, and agent systems: identify
what is reused, who may reuse it, how long it remains valid, what evidence must
travel with it, and what happens when it is wrong or unavailable.
Attendees will receive a three-scope architecture, a request-lifecycle diagram,
and a checklist for deciding whether a result belongs in conversation state, a
semantic cache, a source-processing cache, or should be recomputed.
Production experience
#platformengineering #mlops #databases #latency #reliability #caching #casestudy #failurestory
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}