Sai Khalandar Mokula

Why LLM citations break across multiple retrievers: the one architectural rule to fix attribution and unlock precise groundedness evals

Submitted Oct 8, 2026

Session title

Why LLM citations break across multiple retrievers: the one architectural rule to fix attribution and unlock precise groundedness evals

One-line summary

Multiple retrievers each number chunks S1, S2, S3. After a later LLM merge, those numbers no longer mean the same document — and if you let the model type URLs instead, you cannot eval at chunk grain. We built a request-scoped gid registry so the model only emits handles, and we grade each claim against the chunk it actually cited. The mechanism is simple; skipping it is how grounded-looking answers ship ungrounded.

What problem are you addressing?

Citation guides assume one retriever and one LLM call. Our production path is not that: several retrievers fan out on the same question, an LLM consolidates, another LLM writes the customer-facing document, then we render PDF.

Prior assumption: every chunk has URLs, so generate with [S1] and map back.

Approach 1 — local IDs. Each retriever labels chunks S1..Sn (SIds). After generation, map S1 to a URL. That mapping is wrong as soon as two retrievers reuse the same local index for different text.

Possible workaround one might think of :
Why not using different Prefixing IDs per retriever
i.e :
Retriever 1 uses chunk ids of form SIds i.e. S1, S2, S3..
Retriever 2 uses chunk ids of form HIds i.e. H1, H2, H3..
Retriver 3 uses chunk ids of form Dids i,e D1, D2, D3..
....
Retriever N

looks like a fix right?
this works until a source/retriever you don’t own shows up and ignores the convention. In distributed systems, it is highly impossible for one team to own all the retrivers. So every team that owns and exposes a retriever must follow our above convention - else plugging in that retriever breaks our citation generation.

Approach 2 — let the LLM type URLs.
It would work, but with the following issues.
Problems:
Invalid or drifted links. Worse for eval: one URL is many chunks, so “did this URL support the claim?” is the wrong question. The excerpt came from a chunk.

Approach 3 — what we shipped. Ingest chunks from all retrievers on this request, de-dupe, assign a global chunk id (gid) plus gid→URLs, rewrite local S# to gid, ask the LLM for gids only, expand URLs after generation. Evals compare claim text to gid chunk text.

Evals supported with our approach:
Example: On-demand section(one such generated document/text) eval (not on the user hot path), one generated section: identity 26/26, support 15/21, faithfulness 13/15 (skipped where support failed). Every citation was in the retrieved pool; six claims still failed — including a capability the cited chunk never states. A source-list dashboard would have called that fully grounded.

Why the three evals exist. A footnote can be a real retrieved source and still be wrong. We split that so the failure tells you where to work:

  • Identity (deterministic): is the cited index in the retrieved pool for this request? Fail = invented or corrupted pointer (generation/carry-through).
  • Support (judge): does that chunk actually contain the claim? Fail = real link, wrong evidence (retrieval or binding).
  • Faithfulness (judge): given the right chunk, did we distort it? Fail = right fact, sloppy rewrite. We skip this when support failed (na) — you cannot be faithful to a chunk that doesn’t back the sentence.

The example Identity 100% with support 71% is the operational signal: the registry held; the answer was still not grounded. That is a guardrail (the model cannot invent a source URL), governance (the pointer survives hops), and a check (a real footnote that does not support the claim still fails).

Who is the intended audience?

Platform / MLOps / applied-AI engineers shipping RAG or agents with more than one retrieval source or more than one generation step. Engineering leads who have to answer whether generated text can go in front of a customer.

Level

Intermediate

List one or two practical takeaways.

  1. Don’t namespace retriever IDs and don’t let the model type URLs. Assign a request-scoped gid, have the LLM emit only that handle, expand URLs after generation. Simple, and we treat it as mandatory once you have more than one retriever or more than one LLM hop.
  2. Persist the chunk text. Grade identity / support / faithfulness at chunk index so a complete source list cannot hide “real link, wrong paragraph.”

What will you share?

  • The three approaches, and why 1 and 2 die.
  • How gids flow: retrieve → normalize → consolidate → render.
  • The three evals: definitions, why URL-level scoring lies, the 100% / 71% split, and how we roll up across accounts (claim-weighted, not one blended “quality %”).
  • Trade-offs: build vs buy, chunk-boundary duplicates, dropping unknown handles, evals on-demand not on the hot path.
  • Live walkthrough: generate → carry → judge on one section. Not a product pitch.

What is your experience with this problem?

  • Production system
  • Internal engineering project
  • Hard-earned engineering lesson

What approaches failed, disappointed, or created unexpected problems?

  • Local S# across retrievers: same label, different chunks, wrong URL after merge.
  • Convention on the retriever: not scalable when you don’t own the next one.
  • Model-typed URLs: hallucinated/broken links; URL→chunk for eval is many-to-one, so support looks better than it is.
  • URLs in the prompt: summary/merged nodes can carry many URLs per chunk; that’s token cost and worse reasoning for a mapping the backend already has.

What will you do differently today?

Treat the LLM as untrusted for citations. It only emits opaque handles against a registry built from what was retrieved on this request. URLs are a render step. Identity holds by construction; support and faithfulness are still judged, because a valid gid can still point at the wrong evidence.

What trade-offs did you consider?

  • Build vs buy: we looked for a product, library, or SDK that (1) works on a self-hosted model with no vendor citation API, (2) keeps citation identity across multiple unowned retrievers and later LLM hops, and (3) evals at chunk grain, not URL grain. Hosted citation APIs assume one model call and one retrieved set. Attribution libraries score a single-shot answer. Nothing we found did pass-through identity plus claim-vs-chunk eval on this shape, so we built a small registry rather than wrap an incomplete buy.
  • Dedup vs exact evidence: identity includes retriever-local id + content, so identical boilerplate in two docs stays two gids. We would rather attribute the document the model actually saw than merge by text.
  • Chunk boundaries: different retrievers chunk different spans; we don’t regex-merge them. Some duplicate context in the consolidator is better than a brittle aligner.
  • Unknown handle: registry drops it. The sentence stays, uncited. We preferred an obvious hole over a forged footnote.
  • The “Eval (Latency vs. Cost)”": Running multi-judge LLM evals (Identity, Support, Faithfulness) synchronously across dozens of sub-queries adds unacceptable user latency. We scoped quick evals to section-level UI triggers, while extending the future scope to include full document citation eval to with help of asynchronous background queue architecture.

How can this help other practitioners?

  • Steal the pattern: gid pass-through, model never types URLs — simple, and a must if citations have to survive more than one hop.
  • Don’t score grounding at rendered-URL grain.
  • Use identity vs support vs faithfulness to split “illegal pointer” from “wrong chunk” from “misread the chunk.”

Current state

Production experience

Tags

#governance #observability #platformengineering #mlops #casestudy #failurestory #demo

Bio

I am Sai Khalandar working at Nutanix as Software Developer in the SaaS Organization.

Comments

{{ gettext('Login to leave a comment') }}

{{ gettext('Post a comment…') }}
{{ gettext('New comment') }}
{{ formTitle }}

{{ errorMsg }}

{{ gettext('No comments posted yet') }}

Hosted by

We care about site reliability, cloud costs, security and data privacy