Skip to content
diviteb

Guide · 5 chapters

RAG systems that don't hallucinate

Chunking, hybrid retrieval, citations, and the eval suite to keep the answers honest.

Chapter 01

What hallucination usually is

When a RAG system hallucinates, the cause is almost always one of: chunks too big to be specific, chunks too small to be coherent, retrieval missing the relevant chunk entirely, or a prompt that doesn't constrain the model to the retrieved context.

Chapter 02

Chunking strategies, per content type

There is no universal chunking strategy. Markdown docs chunk well by heading. Transcripts chunk well by speaker turn. Code chunks well by function. PDFs are a problem of their own. We default to a hybrid: structural chunking with a sliding-window fallback for long sections.

Chapter 03

Hybrid retrieval — dense + BM25

Dense vectors are great for semantic similarity but miss exact-keyword matches that BM25 handles well. We run both and fuse the scores. The fusion ratio is tuned per corpus during eval — typically start near 70/30 dense/sparse for technical docs and 50/50 for support transcripts, then let the evals move it.

Chapter 04

Citations and ranked sources

Every answer surfaces the chunks it drew from, ranked by retrieval score. The user (or your reviewer) can verify the source. This is non-negotiable — without citations, you can't tell whether the model hallucinated or whether the source itself is wrong.

Chapter 05

Evals that catch hallucination

We grade RAG answers on two axes: correctness (did the answer match the ground truth) and groundedness (was the answer derivable from the cited chunks). A correct but ungrounded answer is a false win — the model got lucky.

Run this with your team

Book the workshop version.

A half-day workshop with your team — same content, your codebase. We tailor the chapters to where your team is today.