Skip to content

Retrieval-quality evaluation on SciFact, opt-in hybrid search, fusion-score fix - #2

Merged
HarshitGadge merged 1 commit into
mainfrom
retrieval-eval
Sep 28, 2026
Merged

HarshitGadge merged 1 commit into
mainfrom
retrieval-eval

Conversation

@HarshitGadge

Copy link
Copy Markdown
Owner

Adds a retrieval-quality evaluation, uses its results to change the defaults, adds optional hybrid search, and corrects an inaccurate claim in the README.

Retrieval-quality evaluation (eval/)

  • eval/retrieval_eval.py runs the service's own chunker, embedder, ChromaDB index, BM25 index and agent over BEIR SciFact (5,183 abstracts, 300 labelled queries, 17,332 chunks). It reports nDCG@10, Recall@10/100 and MRR@10, with paired-bootstrap CIs. make eval-data && make eval; results are committed in eval/results/scifact.json.
  • Vector search scores nDCG@10 0.727, in line with BAAI's published 0.713 for this model on SciFact.

Defaults changed because of it

  • MMR is now off by default (RAG_MMR_LAMBDA=1.0). At 0.5 it cost 5 points of nDCG@10: the service goes from 0.663 to 0.713 (+0.050, 95% CI +0.031 to +0.069), and recall@10 from 0.734 to 0.806.

Hybrid search (opt-in, RAG_HYBRID_SEARCH=true)

  • New in-memory BM25 index (app/retrieval/lexical.py). It is rebuilt from ChromaDB at startup and kept in sync on upsert/reset, and fused with vector results by weighted RRF (RAG_KEYWORD_WEIGHT, default 0.5).
  • On SciFact it raises recall@100 from 0.940 to 0.962 but doesn't improve nDCG@10 (−0.017, CI −0.039 to +0.004), so it stays off by default. It's meant for corpora with exact identifiers.

Bug fix

  • Rank fusion overwrote hits' cosine scores with RRF scores (~0.03). Those are always below the agent's 0.35 weak-coverage threshold, so every decomposed (multi-part) question retried, and the retry's results replaced the decomposed ones. I confirmed this on the previous commit: rounds == 2 for a two-part question. Scores now stay cosine; the fusion score goes in SearchHit.fused_score. A regression test covers it.

Correction: the embedding model is not int8

  • fastembed's Qdrant/bge-small-en-v1.5-onnx-Q has an empty quantization config, and all 149 weight tensors are FLOAT16. It's a graph-optimised fp16 export.
  • The backend is renamed onnx_optimized; onnx_quantized still works, with a deprecation warning. The README's benchmark text is corrected, including the explanation of the search-stage slowdown, which relied on the int8 claim.
  • Its quality matches the fp32 export on SciFact (nDCG@10 0.727 vs 0.728).

Other

  • MMR is vectorised with numpy: the same selections as the original loop, checked against a reference implementation on 200 random cases in the tests. It also has a fast path when λ=1.
  • RAG_EMBED_MODEL_PATH loads a local model directory, for offline use.
  • numpy is declared as a direct dependency.
  • The README's placeholder clone URL is fixed.
  • 94 tests pass (16 new, in tests/test_hybrid.py), and ruff is clean.

Numbers came from a 2-vCPU cloud machine. The M3 latency table in the README is unchanged apart from the backend label.

🤖 Generated with Claude Code

https://claude.ai/code/session_011FP1EQ2trPyfu1ViFcxRkW


Generated by Claude Code

…nd int8 claim

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011FP1EQ2trPyfu1ViFcxRkW
@HarshitGadge
HarshitGadge merged commit 223dc6b into main Sep 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant