Skip to content

About

A six-stage agentic RAG loop in ~200 lines of pure-stdlib Python — gap critic, adversarial verification, cited synthesis. Retrieval is a loop, not a lookup.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Agentic RAG From Scratch

Agentic RAG — Retrieval is a loop, not a lookup.

License Zero deps RyanAI Lab

Retrieval is a loop, not a lookup.

A complete six-stage agentic RAG loop in ~200 readable lines of pure-stdlib Python, plus a zero-config demo whose toy corpus contains a planted contradiction — so you can watch the adversary stage catch it instead of confidently asserting whichever chunk ranked higher. Distilled from the retrieval system we run in production (see the companion embedding and reranker books).

Why single-shot RAG fails quietly

Retrieve top-k, stuff the context, answer — this works until the day two documents disagree, the top chunk is stale, or the question needs an angle nobody retrieved. Single-shot RAG has no mechanism to notice any of that. It answers fluently either way. The failure is silent, which is the worst kind.

The six stages

0 PLAN         decompose the question into sub-questions
1 FAN-OUT      retrieve per sub-question (multiple angles, deduped)
2 DEEP-READ    read top docs fully; extract claims WITH exact quotes
3 GAP CRITIC   "enough to answer?" — new sub-questions, loop until dry
4 ADVERSARY    per claim, actively search for CONTRADICTING evidence
5 SYNTHESIZE   cited answer + confidence + explicit "could not verify" list

Two design decisions carry most of the value:

  • The gap critic loops (bounded by MAX_ROUNDS total retrieval rounds): retrieval that can ask for what it's missing beats retrieval that hopes round one was enough.
  • The adversary stage treats every claim as a suspect: for each extracted claim it searches the corpus again looking for disagreement. Disputed claims are surfaced as conflicts in the answer — never stated as fact.

Every stage appends one line to a JSONL trace. An observable loop is the difference between a demo and something you can actually debug.

Run the demo (zero config beyond an LLM endpoint)

export LLM_BASE=http://localhost:11434/v1   # any OpenAI-compatible endpoint
export LLM_MODEL=<your-model>               # LLM_KEY too if needed
python3 demo/run_demo.py

The toy corpus documents a fictional kettle. The warranty policy and the support FAQ flatly disagree (12 months vs 24 months — no effective-date escape hatch, a genuine conflict). Ask the default question — "How long is the Orbit Kettle warranty, and what does it cover?" — and watch:

  1. single-shot behavior would pick one number and sound sure;
  2. the adversary stage flags the conflict (disputed ≥ 1);
  3. the final answer cites both documents and refuses to pick a side, listing the conflict under Could not verify instead of asserting either.

Use your own documents

The retriever is a two-method hook — implement search(query, k) and get(doc_id), pass your object to rag.loop.run(), done. The shipped MarkdownStore is a deliberately boring BM25-lite over a folder of .md files: retrieval quality is your hook to improve (that's where vector search, hybrid ranking, and a fine-tuned embedding model plug in — see the companion training books for how we built ours, including the eval-leakage and scoring-head traps).

What we learned running this in production

  • Write-back closes the loop: answers worth keeping become new corpus documents, so the next question benefits. (Kept out of the minimal implementation on purpose — add it once your claims/citation quality is trustworthy, not before.)
  • Sufficiency critics need a dry-out guard: an unbounded "retrieve more" loop will happily burn your budget. MAX_ROUNDS + "no fresh docs" both stop it.
  • Adversarial verification is the cheapest honesty you can buy: one extra search-and-judge pass per claim, and contradictions stop being silent.
  • The retriever matters less than the loop at small corpus sizes — and more than everything else at large ones. Instrument before optimizing.

RyanAI Lab · Distilled from a production retrieval stack. Updated 2026-09. Issues welcome.

About

A six-stage agentic RAG loop in ~200 lines of pure-stdlib Python — gap critic, adversarial verification, cited synthesis. Retrieval is a loop, not a lookup.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages