Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
Sai Khalandar Mokula
Submitted Oct 8, 2026
Why LLM citations break across multiple retrievers: the one architectural rule to fix attribution and unlock precise groundedness evals
Multiple retrievers each number chunks S1, S2, S3. After a later LLM merge, those numbers no longer mean the same document — and if you let the model type URLs instead, you cannot eval at chunk grain. We built a request-scoped gid registry so the model only emits handles, and we grade each claim against the chunk it actually cited. The mechanism is simple; skipping it is how grounded-looking answers ship ungrounded.
Citation guides assume one retriever and one LLM call. Our production path is not that: several retrievers fan out on the same question, an LLM consolidates, another LLM writes the customer-facing document, then we render PDF.
Prior assumption: every chunk has URLs, so generate with [S1] and map back.
Approach 1 — local IDs. Each retriever labels chunks S1..Sn (SIds). After generation, map S1 to a URL. That mapping is wrong as soon as two retrievers reuse the same local index for different text.
Possible workaround one might think of :
Why not using different Prefixing IDs per retriever
i.e :
Retriever 1 uses chunk ids of form SIds i.e. S1, S2, S3..
Retriever 2 uses chunk ids of form HIds i.e. H1, H2, H3..
Retriver 3 uses chunk ids of form Dids i,e D1, D2, D3..
....
Retriever N
looks like a fix right?
this works until a source/retriever you don’t own shows up and ignores the convention. In distributed systems, it is highly impossible for one team to own all the retrivers. So every team that owns and exposes a retriever must follow our above convention - else plugging in that retriever breaks our citation generation.
Approach 2 — let the LLM type URLs.
It would work, but with the following issues.
Problems:
Invalid or drifted links. Worse for eval: one URL is many chunks, so “did this URL support the claim?” is the wrong question. The excerpt came from a chunk.
Approach 3 — what we shipped. Ingest chunks from all retrievers on this request, de-dupe, assign a global chunk id (gid) plus gid→URLs, rewrite local S# to gid, ask the LLM for gids only, expand URLs after generation. Evals compare claim text to gid chunk text.
Evals supported with our approach:
Example: On-demand section(one such generated document/text) eval (not on the user hot path), one generated section: identity 26/26, support 15/21, faithfulness 13/15 (skipped where support failed). Every citation was in the retrieved pool; six claims still failed — including a capability the cited chunk never states. A source-list dashboard would have called that fully grounded.
Why the three evals exist. A footnote can be a real retrieved source and still be wrong. We split that so the failure tells you where to work:
na) — you cannot be faithful to a chunk that doesn’t back the sentence.The example Identity 100% with support 71% is the operational signal: the registry held; the answer was still not grounded. That is a guardrail (the model cannot invent a source URL), governance (the pointer survives hops), and a check (a real footnote that does not support the claim still fails).
Platform / MLOps / applied-AI engineers shipping RAG or agents with more than one retrieval source or more than one generation step. Engineering leads who have to answer whether generated text can go in front of a customer.
Intermediate
S# across retrievers: same label, different chunks, wrong URL after merge.Treat the LLM as untrusted for citations. It only emits opaque handles against a registry built from what was retrieved on this request. URLs are a render step. Identity holds by construction; support and faithfulness are still judged, because a valid gid can still point at the wrong evidence.
Production experience
#governance #observability #platformengineering #mlops #casestudy #failurestory #demo
I am Sai Khalandar working at Nutanix as Software Developer in the SaaS Organization.
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}