Retrieval-quality evaluation on SciFact, opt-in hybrid search, fusion-score fix - #2
Merged
Merged
Conversation
…nd int8 claim Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011FP1EQ2trPyfu1ViFcxRkW
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds a retrieval-quality evaluation, uses its results to change the defaults, adds optional hybrid search, and corrects an inaccurate claim in the README.
Retrieval-quality evaluation (
eval/)eval/retrieval_eval.pyruns the service's own chunker, embedder, ChromaDB index, BM25 index and agent over BEIR SciFact (5,183 abstracts, 300 labelled queries, 17,332 chunks). It reports nDCG@10, Recall@10/100 and MRR@10, with paired-bootstrap CIs.make eval-data && make eval; results are committed ineval/results/scifact.json.Defaults changed because of it
RAG_MMR_LAMBDA=1.0). At 0.5 it cost 5 points of nDCG@10: the service goes from 0.663 to 0.713 (+0.050, 95% CI +0.031 to +0.069), and recall@10 from 0.734 to 0.806.Hybrid search (opt-in,
RAG_HYBRID_SEARCH=true)app/retrieval/lexical.py). It is rebuilt from ChromaDB at startup and kept in sync on upsert/reset, and fused with vector results by weighted RRF (RAG_KEYWORD_WEIGHT, default 0.5).Bug fix
rounds == 2for a two-part question. Scores now stay cosine; the fusion score goes inSearchHit.fused_score. A regression test covers it.Correction: the embedding model is not int8
Qdrant/bge-small-en-v1.5-onnx-Qhas an empty quantization config, and all 149 weight tensors are FLOAT16. It's a graph-optimised fp16 export.onnx_optimized;onnx_quantizedstill works, with a deprecation warning. The README's benchmark text is corrected, including the explanation of the search-stage slowdown, which relied on the int8 claim.Other
RAG_EMBED_MODEL_PATHloads a local model directory, for offline use.numpyis declared as a direct dependency.tests/test_hybrid.py), and ruff is clean.Numbers came from a 2-vCPU cloud machine. The M3 latency table in the README is unchanged apart from the backend label.
🤖 Generated with Claude Code
https://claude.ai/code/session_011FP1EQ2trPyfu1ViFcxRkW
Generated by Claude Code