Skip to content

Adopt ①: P2 pre-rerank pool anchors + top-N recall floor + holdout calibration - #52

Open
ManSio wants to merge 2 commits into
mainfrom
adopt/reranker-threshold-and-pool
Open

ManSio wants to merge 2 commits into
mainfrom
adopt/reranker-threshold-and-pool

Conversation

@ManSio

@ManSio ManSio commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Root cause P2 (verified, was Open)

3-way RRF rewards multi-tier consensus: the P2 target is found by ONE tier only (BM25 rank 0 for \src/core/search/engine.py\ -> 1/(60+1)≈0.0164) while junk present in 2-3 tiers at mediocre ranks accumulates 2-3x that.
rf_results[:limit]\ (\engine.py:746) then amputates the single-tier winner before the reranker ever sees it.

Ruled out by reading: MMR is reorder-only (no drops); bucket weights favour the target (.py 1.0 vs .txt/.md 0.5); query expansion keeps the verbatim query as variants[0]. This resolves the open contradiction in KNOWN_ISSUES (standalone BM25 rank 0 vs hybrid loss): hybrid never returns raw BM25 order. Regression test failed before (target cut, pool = 10x experiments artifacts) / passes after.

Design

  • \�nchor_tier_winners()\ (\scoring.py): pool = MMR-ordered RRF top-limit (order preserved for the no-reranker path) + missing per-tier top-1 winners appended from the full RRF list (reconstructed as fused entries when outside it), capped at \MAX_RERANKER_INPUT. Reranker \ op_n\ (=limit) unchanged. No P2 remainder — the fix is complete and cheap.
  • Top-N recall floor \MAX_RERANKER_TOPN\ (default 0 = current behavior exactly): union of threshold-passers with top-N by score. Absolute 0.3 cut untouched.
  • \ hreshold_calibration.calibrate_threshold()\ (F1-max on labeled holdout) raises \ValueError\ on eval sources (\eval/frozen/token_reduction_v3/…) — the sweep-on-eval rejection (EXPERIMENTS_LOG 2026-09-27) is encoded as a test, not a comment.
  • Holdout protocol + sizes (>=10 queries disjoint from frozen eval-16, sha 31f1b0c9) documented in \ hreshold_calibration.py\ docstring. Default 0.3 unchanged pending a real holdout.
  • Commit 1 also carries the previously-uncommitted sigmoid normalization found in the working tree (llama.cpp raw logits -> [0,1]); all 7 sigmoid tests green.

Numbers

  • New suite: 18 tests in \ est_reranker_pool_and_threshold.py\ (P2 regression red->green, anchors/dedup/cap/fallback/limit=0, top-N off/on/cap on logged P3 logits, F1-selection, 6 eval markers -> ValueError).
  • Affected suites: 143 passed (reranker 45 + pool/threshold 18 + searcher + hardening + ubatch + bs_audit); ruff clean; pre-commit gates green on both commits.
  • Holdout/eval before-after: synthetic only. P3 logged values: target 0.271 stays cut at default (documented ranking failure, not scale); with \MAX_RERANKER_TOPN=4\ it returns as 4th by score. Live P2/P3 on the real index + holdout calibration NOT run here.

Caveats (open)

  • Live validation on real index + llama.cpp needed (synthetic mechanism only).
  • Holdout set (>=10 queries outside eval-16) not collected — owner's call.
  • No split needed: P2 remainder is zero; threshold part ships as mechanism with default-preserving behavior.

DO NOT MERGE yet — review requested.

MSCodeBase Agent added 2 commits September 28, 2026 00:39
Includes previously-uncommitted sigmoid normalization found in working
tree (EXPERIMENTS_LOG 2026-09-27: llama.cpp /v1/rerank returns raw logits
~[-11,+11], not Cohere [0,1]; 1/(1+e^-x) in llama_cpp branch, MIN_RERANK_SCORE
stays 0.3; 7 sigmoid tests).

New: MAX_RERANKER_TOPN env tumbler (PerformanceConfig.reranker_topn_keep,
default 0 = current behavior) - union of threshold-passers with top-N by
score, so uncalibrated cross-encoder scores (P3 target 0.271<0.3) can survive
as recall floor without lowering the absolute cut. calibrate_threshold()
(F1-max on labeled holdout) refuses eval sources with ValueError - the
anti-overfit rule (no sweep on the 16 frozen eval rules) is encoded as code.
Split protocol + sizes (>=10 holdout queries disjoint from eval-16) in
threshold_calibration.py docstring. Default 0.3 unchanged pending a real
holdout measurement.
…use)

Root cause (verified): 3-way RRF rewards multi-tier consensus, so a target
found by ONE tier only (P2: BM25 rank 0 for src/core/search/engine.py scores
1/(60+1)) loses to junk present in 2-3 tiers (2-3x the score) and is amputated
by rrf_results[:limit] before the reranker ever sees it. MMR is innocent
(reorder-only, no drops); bucket weights favour the target (.py 1.0 vs
.txt/.md 0.5); query expansion keeps the verbatim query as variants[0].
This resolves the open contradiction in KNOWN_ISSUES (standalone BM25 rank 0
vs hybrid loss): hybrid never returns raw BM25 order.

Fix: anchor_tier_winners() in scoring.py - pool = MMR-ordered RRF top-limit
(order preserved for the no-reranker path) + missing per-tier top-1 winners
appended from the full RRF list (reconstructed as fused entries when outside
it), capped at MAX_RERANKER_INPUT. Reranker top_n (=limit) unchanged.
Regression test failed before (target cut at rank ~13) / passes after.
KNOWN_ISSUES P2 entry updated: root cause established, live validation +
holdout calibration remain open.
@coderabbitai

coderabbitai Bot commented Sep 27, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 86bd06f3-a26c-4306-9ac9-fb8a866501db


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant