Guide · 5 chapters
RAG systems that don't hallucinate
Chunking, hybrid retrieval, citations, and the eval suite to keep the answers honest.
Chapter 01
What hallucination usually is
When a RAG system hallucinates, the cause is almost always one of: chunks too big to be specific, chunks too small to be coherent, retrieval missing the relevant chunk entirely, or a prompt that doesn't constrain the model to the retrieved context.
Chapter 02
Chunking strategies, per content type
There is no universal chunking strategy. Markdown docs chunk well by heading. Transcripts chunk well by speaker turn. Code chunks well by function. PDFs are a problem of their own. We default to a hybrid: structural chunking with a sliding-window fallback for long sections.
Chapter 03
Hybrid retrieval — dense + BM25
Dense vectors are great for semantic similarity but miss exact-keyword matches that BM25 handles well. We run both and fuse the scores. The fusion ratio is tuned per corpus during eval — typically start near 70/30 dense/sparse for technical docs and 50/50 for support transcripts, then let the evals move it.
Chapter 04
Citations and ranked sources
Every answer surfaces the chunks it drew from, ranked by retrieval score. The user (or your reviewer) can verify the source. This is non-negotiable — without citations, you can't tell whether the model hallucinated or whether the source itself is wrong.
Chapter 05
Evals that catch hallucination
We grade RAG answers on two axes: correctness (did the answer match the ground truth) and groundedness (was the answer derivable from the cited chunks). A correct but ungrounded answer is a false win — the model got lucky.
More guides
Multi-tenant SaaS, end to end
Postgres RLS, RBAC, metered billing, and a SOC 2-ready audit trail — wired in before the first tenant signs.
Building production AI agents
Tool use against real APIs, eval harnesses, observability, and the kill-switches you'll want.
Core Web Vitals — the playbook we run
LCP, INP, CLS — how we diagnose, fix, and lock in. With the GitHub Action we use to enforce.
Run this with your team
Book the workshop version.
A half-day workshop with your team — same content, your codebase. We tailor the chapters to where your team is today.