What I Learned Building a Citation-Based RAG System

Table of contents
- Lesson 1: A citation is only as good as the chunk boundary
- Lesson 2: Embedding similarity and citation-worthiness are different things
- Lesson 3: Rerank before you cite, not after
- Lesson 4: A citation needs its neighbors, not just itself
- Lesson 5: Provenance has to be captured at ingest, not reconstructed later
- Lesson 6: A wrong citation is worse than no citation
- What I’d tell someone starting from scratch
- FAQ
- Related resources
A citation-based RAG system answers questions with a link back to the exact source, so a reader can check the claim instead of trusting it blind. I built one, Nexus, a self-hosted knowledge base for Slack, code, wikis and tickets. The hardest part turned out to have nothing to do with retrieval. It was making sure every citation actually supported the sentence it was attached to.
Six things I learned the hard way, in the order I hit them.
Lesson 1: A citation is only as good as the chunk boundary
A citation points at a chunk of text, not a document. Cut a sentence in half, or drop the heading that explains what the paragraph is even about, and the citation is technically correct and completely useless. The reader clicks through and still can’t tell if the claim holds up.
This is why Nexus chunks code by language-aware boundaries (class, then method, then block) instead of a fixed token window, and why Slack gets distilled into a {question, summary, resolution} record instead of embedding raw messages. A chunk has to be a complete thought before it’s worth citing.
Lesson 2: Embedding similarity and citation-worthiness are different things
Cosine similarity measures “these two pieces of text use similar words.” It does not measure “this text is good evidence for that claim.” A short Slack reply like “sounds good, thanks!” regularly scores closer to a query than the long, detailed answer three messages above it, simply because short text has less noise to dilute the match.
Nexus runs four separate scorers over every Slack thread: full-text, embedding, inverse document frequency, and age decay. None of them is reliable by itself, which is the whole point of running four. A citation system that leans on embedding similarity alone will confidently cite the wrong message more often than feels acceptable once real users start clicking the links.
Lesson 3: Rerank before you cite, not after
The fix for lesson 2 isn’t a better embedding model. It’s a second pass: take the fused candidates and score each one against the actual question with a separate, smaller model, from 0 to 10. Nexus caches this score per (model, query, document) and caps the whole thing at a four-second time limit. If it’s slow, the system just falls back to the pre-rerank order instead of making the user wait.
Here’s why the order matters: a document can rank first because it shares vocabulary with the query while actually answering a different question. Reranking is what catches that mismatch before it turns into a citation someone trusts and then finds out is wrong.
Lesson 4: A citation needs its neighbors, not just itself
Chunking splits text from its surroundings. A matched paragraph can lose the heading above it, the precondition two sentences earlier, or the caveat right after it, and that’s exactly the context a reader needs to judge whether the citation actually backs the claim.
Nexus pulls in the neighboring chunks of every finalist after reranking, once the order is already locked in. The citation points at the expanded passage, never the isolated match on its own. Out of everything here, this is the one lesson that has nothing to do with retrieval quality. It’s purely about what happens when a reader actually clicks.
Lesson 5: Provenance has to be captured at ingest, not reconstructed later
Every embedding row in Nexus carries its source id, timestamp, and channel or repo right alongside the vector, written at ingest time by the connector that read it. Trying to reconstruct “where did this come from” later, from the content alone, doesn’t work. Two near-identical paragraphs from different Slack channels, with different access permissions, look exactly the same once they’re just text sitting in a vector index.
Get this wrong and you don’t find out until a citation links to a channel the asker can’t actually see, or a stale answer cites a source that’s since been deleted. Both are findable in testing. Neither is fun to find in production.
Lesson 6: A wrong citation is worse than no citation
This is the one that changed how I think about the whole feature. A plain, uncited answer gets read with the skepticism it deserves; the reader knows to go verify it themselves. A cited answer gets trusted on the strength of the citation alone, often without anyone actually opening the link to check. And if that citation turns out to be wrong (a real page that just doesn’t say what the answer claims), it doesn’t just fail to help. It actively launders a wrong answer into a credible-looking one.
Lessons 2 through 4 aren’t optional polish because of this one. Hybrid retrieval, reranking, and context expansion all exist because a bad citation costs more than a bad answer with no citation attached.
flowchart LR
accTitle: Where a bad citation gets caught, or doesn't
accDescr: A query is fused across retrievers, which can rank a wrong document first on vocabulary overlap. Reranking against the actual question catches most of that before citing. Without reranking, the fused order goes straight to the reader as a confidently wrong citation.
Q[Query]
FUSE[Fused candidates]
RR{Rerank against<br/>the real question}
GOOD[Citation the<br/>reader can trust]
BAD[Wrong citation,<br/>trusted anyway]
Q --> FUSE
FUSE --> RR
RR -->|catches mismatch| GOOD
FUSE -.skip reranking.-> BAD
What I’d tell someone starting from scratch
- Design the chunk boundary first. Everything downstream (retrieval, reranking, citing) inherits whatever unit of text you chose at ingest time.
- Never trust a single retrieval signal. Full-text, embeddings, recency and a reranker each catch a failure mode the others miss.
- Rank first, then widen the window. Rerank the tight chunk, then cite the expanded one.
- Carry provenance and permissions from the connector itself. You can’t infer access control from text that’s already in the index.
- Measure citation accuracy on purpose. If nobody is checking whether citations actually support their claims, the system is optimizing for “looks cited,” not “is correct.”
FAQ
Do I need a reranker to build a citation-based RAG system?
Not to start, but plan for one early. Vector search alone ranks documents by vocabulary overlap, which regularly promotes a wrong-but-similar match over the right-but-differently-worded one. A reranker is the cheapest fix once that starts producing citations that don’t hold up.
How do you actually measure whether citations are correct?
Spot-check a sample of real answers against their cited sources by hand, on a schedule, not just when something looks wrong. Nexus also supports A/B testing its optional graph retriever against plain hybrid search with a built-in eval runner, which is the same idea applied to a specific retrieval technique rather than the whole pipeline.
Is a citation system just a RAG system with footnotes added?
No. Footnotes are a UI decision made after retrieval is already done. A citation system has to treat “can I actually prove this” as a design constraint from ingest onward: chunk boundaries that preserve meaning, provenance captured at the source, reranking before anything reaches the reader. Bolt footnotes onto ordinary RAG instead, and you just get confident-looking links to the wrong evidence.


