Retrieval is a loop, not a lookup.
A complete six-stage agentic RAG loop in ~200 readable lines of pure-stdlib Python, plus a zero-config demo whose toy corpus contains a planted contradiction — so you can watch the adversary stage catch it instead of confidently asserting whichever chunk ranked higher. Distilled from the retrieval system we run in production (see the companion embedding and reranker books).
Retrieve top-k, stuff the context, answer — this works until the day two documents disagree, the top chunk is stale, or the question needs an angle nobody retrieved. Single-shot RAG has no mechanism to notice any of that. It answers fluently either way. The failure is silent, which is the worst kind.
0 PLAN decompose the question into sub-questions
1 FAN-OUT retrieve per sub-question (multiple angles, deduped)
2 DEEP-READ read top docs fully; extract claims WITH exact quotes
3 GAP CRITIC "enough to answer?" — new sub-questions, loop until dry
4 ADVERSARY per claim, actively search for CONTRADICTING evidence
5 SYNTHESIZE cited answer + confidence + explicit "could not verify" list
Two design decisions carry most of the value:
- The gap critic loops (bounded by
MAX_ROUNDStotal retrieval rounds): retrieval that can ask for what it's missing beats retrieval that hopes round one was enough. - The adversary stage treats every claim as a suspect: for each extracted claim it searches the corpus again looking for disagreement. Disputed claims are surfaced as conflicts in the answer — never stated as fact.
Every stage appends one line to a JSONL trace. An observable loop is the difference between a demo and something you can actually debug.
export LLM_BASE=http://localhost:11434/v1 # any OpenAI-compatible endpoint
export LLM_MODEL=<your-model> # LLM_KEY too if needed
python3 demo/run_demo.pyThe toy corpus documents a fictional kettle. The warranty policy and the support FAQ flatly disagree (12 months vs 24 months — no effective-date escape hatch, a genuine conflict). Ask the default question — "How long is the Orbit Kettle warranty, and what does it cover?" — and watch:
- single-shot behavior would pick one number and sound sure;
- the adversary stage flags the conflict (
disputed ≥ 1); - the final answer cites both documents and refuses to pick a side, listing the conflict under Could not verify instead of asserting either.
The retriever is a two-method hook — implement search(query, k) and
get(doc_id), pass your object to rag.loop.run(), done. The shipped
MarkdownStore is a deliberately boring BM25-lite over a folder of .md
files: retrieval quality is your hook to improve (that's where vector
search, hybrid ranking, and a fine-tuned embedding model plug in — see the
companion training books for how we built ours, including the eval-leakage
and scoring-head traps).
- Write-back closes the loop: answers worth keeping become new corpus documents, so the next question benefits. (Kept out of the minimal implementation on purpose — add it once your claims/citation quality is trustworthy, not before.)
- Sufficiency critics need a dry-out guard: an unbounded "retrieve more"
loop will happily burn your budget.
MAX_ROUNDS+ "no fresh docs" both stop it. - Adversarial verification is the cheapest honesty you can buy: one extra search-and-judge pass per claim, and contradictions stop being silent.
- The retriever matters less than the loop at small corpus sizes — and more than everything else at large ones. Instrument before optimizing.
RyanAI Lab · Distilled from a production retrieval stack. Updated 2026-09. Issues welcome.
