Conversation
added 2 commits
September 28, 2026 00:39
Includes previously-uncommitted sigmoid normalization found in working tree (EXPERIMENTS_LOG 2026-09-27: llama.cpp /v1/rerank returns raw logits ~[-11,+11], not Cohere [0,1]; 1/(1+e^-x) in llama_cpp branch, MIN_RERANK_SCORE stays 0.3; 7 sigmoid tests). New: MAX_RERANKER_TOPN env tumbler (PerformanceConfig.reranker_topn_keep, default 0 = current behavior) - union of threshold-passers with top-N by score, so uncalibrated cross-encoder scores (P3 target 0.271<0.3) can survive as recall floor without lowering the absolute cut. calibrate_threshold() (F1-max on labeled holdout) refuses eval sources with ValueError - the anti-overfit rule (no sweep on the 16 frozen eval rules) is encoded as code. Split protocol + sizes (>=10 holdout queries disjoint from eval-16) in threshold_calibration.py docstring. Default 0.3 unchanged pending a real holdout measurement.
…use) Root cause (verified): 3-way RRF rewards multi-tier consensus, so a target found by ONE tier only (P2: BM25 rank 0 for src/core/search/engine.py scores 1/(60+1)) loses to junk present in 2-3 tiers (2-3x the score) and is amputated by rrf_results[:limit] before the reranker ever sees it. MMR is innocent (reorder-only, no drops); bucket weights favour the target (.py 1.0 vs .txt/.md 0.5); query expansion keeps the verbatim query as variants[0]. This resolves the open contradiction in KNOWN_ISSUES (standalone BM25 rank 0 vs hybrid loss): hybrid never returns raw BM25 order. Fix: anchor_tier_winners() in scoring.py - pool = MMR-ordered RRF top-limit (order preserved for the no-reranker path) + missing per-tier top-1 winners appended from the full RRF list (reconstructed as fused entries when outside it), capped at MAX_RERANKER_INPUT. Reranker top_n (=limit) unchanged. Regression test failed before (target cut at rank ~13) / passes after. KNOWN_ISSUES P2 entry updated: root cause established, live validation + holdout calibration remain open.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Root cause P2 (verified, was Open)
3-way RRF rewards multi-tier consensus: the P2 target is found by ONE tier only (BM25 rank 0 for \src/core/search/engine.py\ -> 1/(60+1)≈0.0164) while junk present in 2-3 tiers at mediocre ranks accumulates 2-3x that.
rf_results[:limit]\ (\engine.py:746) then amputates the single-tier winner before the reranker ever sees it.
Ruled out by reading: MMR is reorder-only (no drops); bucket weights favour the target (.py 1.0 vs .txt/.md 0.5); query expansion keeps the verbatim query as variants[0]. This resolves the open contradiction in KNOWN_ISSUES (standalone BM25 rank 0 vs hybrid loss): hybrid never returns raw BM25 order. Regression test failed before (target cut, pool = 10x experiments artifacts) / passes after.
Design
Numbers
Caveats (open)
DO NOT MERGE yet — review requested.