fix(search): module-doc chunker + P3 anchor + reranker scale + measurement hygiene (P3 gate GREEN) - #58
fix(search): module-doc chunker + P3 anchor + reranker scale + measurement hygiene (P3 gate GREEN)#58ManSio wants to merge 5 commits into
Conversation
…ract) llama.cpp returns logits ~[-11,+11] vs threshold calibrated for [0,1]; filter ate 70-97%. Threshold 0.3 untouched. Does NOT fix P3 ranking — see KNOWN_ISSUES P3 entry.
…r_p3/RESULT-chunker.md)
Probe/result JSON dumps and experiments results/work files contain verbatim query terms and dominate BM25, invalidating gates. Canonical exclusion via SystemArtifacts.is_measurement_artifact, wired into FileGuard.should_skip_file used by all indexer walks.
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Conflict map vs #62/#63 (for your rebase; owner decision: #58 merges first, I adapt after). KEEP YOURS (superset, no action needed on my side):
SWAP ON REBASE (mine wins, #63 already open):
NO-OP overlaps (disjoint regions, merge cleanly after rebase):
DB note: purge of 772 garbage files already executed against the shared index (verified 0 left); no re-purge needed after either merge. |
Mechanism (why P3 flipped FALSIFIED -> PASS)
Three stacked changes, each load-bearing (ablation in Red Team):
src/core/indexing/parser.py): AST chunker dropped module-leveldocstrings index-wide, so the passage the reranker scores highest never
existed as a chunk. Files with a docstring now get it as chunk 0
(
symbol_type=module_docstring,hierarchy_level=module); no-docstringbehavior unchanged. Live index: 371 module-doc chunks, all
chunk_index=0.src/providers/reranker/multi_provider.py): llama.cpp rawlogits (e.g. P3 code chunk -0.99) were filtered against
MIN_RERANK_SCORE=0.3unnormalized. Sigmoid normalization per the scale contract.
src/core/search/engine.py:_anchor_module_head_chunks_async):appends each pool file's chunk-0 head pre-rerank (tail, no re-sort; capped by
_O1_ANCHOR_TOTAL/MAX_RERANKER_INPUT; doc/data files ineligible). P2 precedent.src/core/system_artifacts.py+file_guard.py,cherry-picked from
fix/index-hygiene-final-gate496aa04d, clean apply):probe/result JSON dumps +
experiments/**/results|workexcluded from the index(they contain verbatim query terms and dominated BM25).
Numbers (FULL reindex + fresh-process gate, rev
f1c32d05)Index (standalone
indexer.index_project+ physical clear, same call the MCPintel_trigger_reindex(mode=full)makes — MCP server was down, see caveats):11026 rows (was 14832), 738 files, IVF_FLAT finalized (was flat,
num_indices=0; now 11026/11026 indexed, COSINE), junkresults/workchunks2152 -> 0, probe/result-JSON chunks 2092 -> 0.
Gate
scripts/p3_holdout_gate.py(FTS prebuild + discarded warm-up + per-casereranker-cache clear), 3/3 fresh-process PASS, bit-identical rows:
vs FALSIFIED baseline (junk index, no chunker/sigmoid):
P3 None, R2 None, H1 None, H3 None, H4 2, H8 1, H11 2, H12 None.No unexpected boosts, no doc boosts, no reranker-cache voids, no degraded rows
(
model=llama.cpp-rerankerevery query). FTS-timeout flake: 0 occurrences in3 gate runs + ~10 probe processes (all stderr clean; the
timed_outfail-loudpath is present, just never triggered). Unit suites: 34/34 pass
(
test_chunker_module_doc5,test_measurement_hygiene20,test_p3_module_head_anchor9).Red team (GREEN gate attacked, 4 attacks + stability work)
(
artifact_gc.py:0, score 0.77,module_head_anchor=True); prereg clause (i)holds live. Anchor-OFF control: P3 Quick question about your repo #1 falls back to a code chunk (0.27) —
chunker+anchor are load-bearing, not decorative.
H1/H8: target file-rank unchanged; H8's Quick question about your repo #1 competitor (
e17_golden.json,anchor=False) enters via BM25/vector tiers, never via the anchor. Anchor exonerated.ctx_F5S-11_B,judged_raw_t5return hits only from legit indexed manifests(
INVENTORY.json,MANIFEST.json), 0 hits from excludedresults/workpaths.reranker_p3/README.mdmatches P3 at 0.00016 vs gold 0.77 (~4000x gap); DOC control clean all runs.
2 in full-gate order (5/5, bit-identical). Quick question about your repo #1 in gate order is the legitimately
indexed eval file
experiments/bootstrap/e17_golden.json(cypher vocabulary,BM25 rank 8 + tiers → rerank 0.17 vs code chunk 0.11). Mechanism unproven
(prime suspects: RRF pool-boundary interplay + ±5% cross-run reranker score
wobble observed on identical inputs). No gate impact (H-rows report-only), but
honest flag. Guard proposals: (a) extend hygiene exclusion to
experiments/bootstrap/e17_*golden files; (b) gate dumps top-5 per row forauditability; (c) 3x repetition of volatile rows in the gate.
Caveats
src.main) was killed mid-task to pick up the synced extension codeand did not auto-respawn; reindex ran standalone via repo venv with the identical
indexer.index_project+recreate_table_physicalpath (log: 738 files, 528.7s).Owner action needed: reload Zed to bring the MCP server back on the new code.
experiments/planted_break/results.jsonlocal modification left untouched(unrelated dirt, not staged, not in this PR).
None= threshold cuts, same as baseline (no regression); H1 4 andH6 2 reflect genuine topical competition (dedicated FTS5/graph experiment files
outscore the generic gold chunk).