diff --git a/AGENT_DIARY.md b/AGENT_DIARY.md index 95047fae..87051b5c 100644 --- a/AGENT_DIARY.md +++ b/AGENT_DIARY.md @@ -693,3 +693,11 @@ chunk_index -(20_000_000+line), graph_score=0.4 (ниже функций 1.0). E **P-003 — правка не в той ветке.** Сигмоида попала в ONNX-блок вместо `llama_cpp` (oldString оказался уникальным, но не тем), ветка ONNX осиротела, `if scores:` выехал из `try`. Поймано `ast.parse` + просмотром diff. **Правило:** после правки в много-ветвистом коде — `git diff` целиком, а не только «применилось». **P-004 — letter-vs-spirit instruction reading.** GPT-5.5 читает «never» буквально (jitter vs page_one_exit, Nishikanta 2026-09-27): perverse-compliant прочтение проходит фильтр, задуманный смысл — нет. Наш зеркальный кейс — qwen temporal-hint (E4b): БЕЗ хинта 'NOT FOUND AT HEAD' ни одна модель не робастна, т.е. правило работает только в дух-прочтении, буква его не несёт. **Guard:** тестировать граничные прочтения каждого правила (perverse-compliant кейс), а не только задуманное. + +## [2026-09-28] P2 — gold не входит в пул: якоря идентификаторов +**Status:** ✅ Код+тесты+live (ветка `fix/p2-pool-contains-gold`, PR следует). +**Root Cause (числа, fresh-process):** срез `rrf_results[:limit]`, raw_limit=min(limit*2,30); цель P2: BM25#126, FTS#74, dense вне @200. Пул 5→1.7с/10→4.0с/20→7.7с/50→22.9с реранка — глубина 126 (≈50с) отвергнута. O1 не спасает: кандидат `reciprocal_rank_fusion` (df=4), символ цели `hybrid_search_async` (df=100) — буст уходил в scoring.py. +**Fix:** `_anchor_identifier_chunks_async` (single-token FTS, def-first, df-кап 120, docs/data ineligible, капы 2+3+MAX) → P2 rank 1 live (engine.py:18). Red-team: H2-коллизия (BM25 df=327 не якорится), doc-guard (docs/X.md, canary .json), caller-vs-def (live_search_audit vs engine). +**Harness-находка:** `asyncio.run()` на запрос роняет чётные запросы в reranker-passthrough (ms=0) — детерминировано по паритету; гейт идёт одним loop + degraded-флаг. Void-флаг KI этот класс не ловил. +**Guard:** `tests/test_p2_pool_anchors.py` (13) + `scripts/p2_holdout_gate.py` (GATE PASS 15/15); смежные 64 passed; ruff check чист. Формат-откат: `ruff format` давал +447 строк churn — откачен, diff +171/-0. +**P-005 — n=5 не значит «пул был 5».** Passthrough реранкера режет `[:top_n]`, пряча сработавшие якоря (6–8-е места) — трижды неверно выводил «якоря не сработали». **Правило:** судить pool-этап только трейсом пула до реранкера, не финальным n. diff --git a/KNOWN_ISSUES.md b/KNOWN_ISSUES.md index da484075..201c08e1 100644 --- a/KNOWN_ISSUES.md +++ b/KNOWN_ISSUES.md @@ -9,6 +9,7 @@ - **Правило:** все retriever-замеры и A/B-тесты — только в свежем процессе либо с явным сбросом реранкер-кэша (`Searcher._reranker_cache.clear()`). Ключ кэша включает текст запроса (engine.py:1646): повтор того же запроса в том же процессе отдаёт закэшированные скоры, а не измеряет код. - **Эвристика void-замера:** wall <2s на `hybrid_search_async` при ожидании полного пайплайна (embed+BM25+FTS+rerank) = подозрение на cache hit; сверяться с `Searcher._last_rerank_timing` (пусто = реранкер не работал). Холодный FTS-билд (~2.5s) — обратная ловушка: ПЕРВЫЙ замер в свежем процессе молча теряет FTS-тир (2s `wait_for`), нужен discarded warm-up на чужом запросе. +- **Harness-ловушка (2026-09-28, Verified):** `asyncio.run()` на КАЖДЫЙ запрос роняет чётные запросы в reranker-passthrough (`reranker_ms=0`, `model='-'`, возврат пула без скоринга) — детерминировано по паритету позиции, свежая/здоровая инфра, флаги провайдера в норме. Серия обязана идти в ОДНОМ event loop; плюс явный degraded-флаг (`not reranker_ms` → замер недействителен). Void-флаг (`timing=={}`) этот класс НЕ ловит (timing={ms:0,...} ≠ {}). - **Статус:** 🟡 Open (процедурное правило; guard-скрипт `scripts/o1_holdout_gate.py` — fresh-process + warm-up + void-флаг). **21 entries** — compressed per §4.8 R3 (conclusion-first; dedup 2026-09-08, 2026-09-21). Closed entries moved to docs/archive/KNOWN_ISSUES_2026_09.md on 2026-09-27 (R1 size guard; second batch on merge experiment/4a-unit-of-return). @@ -45,12 +46,13 @@ - **Guard:** `tests/test_reranker.py` — `test_sigmoid_matches_reference_values` (4 кейса), `test_sigmoid_is_numerically_stable_at_extremes`, `test_llama_cpp_scores_normalized_to_unit_interval`, `test_llama_cpp_negative_logit_does_not_leave_unit_interval`. Проверено мутацией (`_sigmoid` → identity): **7 failed / 38 passed**; чисто — 45 passed. - **T3 (обобщение):** иных мест с абсолютным порогом по логитам в `src/` нет. `_DEFAULT_THRESHOLD = 0.85` в `duplication.py:37` — порог по Jaccard (по определению в [0,1], `clamp` на строке 136), другой механизм. -## 2026-09-27 — P2: целевой файл не доходит до финального пула; причина не установлена (Open) +## 2026-09-27 — P2: целевой файл не доходит до финального пула (Root Cause установлен, fix в PR) - **Симптом:** для запроса P2 (`hybrid_search_async reciprocal_rank_fusion FTS5 BM25`) целевой `src/core/search/engine.py` не найден. Top-хиты — собственные артефакты эксперимента: `experiments/**/*.txt`, `results.json`, `docs/zh/SEARCH_PIPELINE.md`. Реранкер ни при чём — цели нет в пуле ещё до него. -- **Root Cause: НЕ УСТАНОВЛЕН.** Зафиксировано открытое противоречие: отдельный standalone-прогон BM25 вернул цель на **rank 0**, что несовместимо с утверждением «цель не находится вовсе». Расхождение между standalone BM25 и путём внутри `hybrid_search_async` не изучено. **Причину не утверждать** до разбора построения пула и RRF-слияния. -- **Побочно (Verified):** индекс содержит вывод собственных экспериментов, что загрязняет lexical-выдачу по общим терминам — это отдельная проблема индексации, не фильтра. -- **Что нужно:** разобрать построение pre-rerank пула и слияние RRF; выяснить, почему BM25 rank-0 не доходит до финального пула. +- **Root Cause (Verified live 2026-09-28, fresh process + discarded warm-up):** срез пула — `rrf_results[:limit]` (`engine.py`, `raw_limit=min(limit*2,30)`), а цель многотермовым RRF зарыта глубоко: **BM25#126, FTS#74, dense вне @200** (индекс загрязнён собственными артефактами — дословный текст запроса лежит в `experiments/`). Ни лимит 50, ни O1 пул не чинят: (a) расширение пула до глубины 126 стоило бы ~126×0.4с реранка (~50с) — замерено и отвергнуто (пул 5→1.7с, 10→4.0с, 20→7.7с, 50→22.9с); (b) O1-кандидат — `reciprocal_rank_fusion` (df=4, строго редчайший), а символ цели — `hybrid_search_async` (df=100): exact-совпадения нет, буст уходит в `scoring.py`. Старый standalone-BM25-rank-0 — устаревший замер на незагрязнённом индексе. Per-tier top-1 anchoring (ветка `adopt/reranker-threshold-and-pool`) для текущего индекса refuted: топы тиров — мусор, цель на #74–126. +- **Fix (ветка `fix/p2-pool-contains-gold`):** `_anchor_identifier_chunks_async` (`engine.py`) — single-token FTS-добор exact-символов редких идентификаторов (df≤120: `hybrid_search_async` 100 ✓, `BM25` 327 ✗, `FTS5` 140 ✗) прямо в pre-rerank пул, def-first, docs/data ineligible, капы 2/токен + 3 всего + MAX_RERANKER_INPUT. Live: P2 rank **1** (чанк engine.py:18). Guard: `tests/test_p2_pool_anchors.py` (13) + `scripts/p2_holdout_gate.py` (GATE PASS 15/15, свежий процесс). +- **Побочно (Verified):** индекс содержит вывод собственных экспериментов — отдельная проблема индексации, не фильтра. Harness-урок: `asyncio.run()` на запрос роняет чётные запросы в reranker-passthrough (ms=0) — гейт идёт одним loop + явный degraded-флаг; void-флаг KI это не ловил (timing≠{}). +- **Остаточное:** df-кап 120 эвристичен и привязан к текущему индексу (100 vs 140 — тонкая граница); P2 rank=1 требует живого реранкера (без него цель в пуле, но не в топе). Валидация O1-гейта (`o1_holdout_gate.py`) тем же harness-багом занижена — не чинилось (чужой мёрджнутый файл). ## 2026-09-25 — Падения не фиксировались: zombie-job + глушение исключений + нет ledger (Fixed) / Open (server hard-death) diff --git a/scripts/p2_holdout_gate.py b/scripts/p2_holdout_gate.py new file mode 100644 index 00000000..986b3183 --- /dev/null +++ b/scripts/p2_holdout_gate.py @@ -0,0 +1,154 @@ +#!/usr/bin/env python3 +"""P2 pool-anchor holdout gate (live, fresh-process): P2 + H1-H12 + N/doc controls. + +Usage: + python scripts/p2_holdout_gate.py [--project D:/Project/MSCodeBase] + +Pattern follows scripts/o1_holdout_gate.py (fresh process, discarded warm-up, +void-flag), PLUS a blocking FTS prebuild: the cold FTS5 to_pandas build takes +~2.9s live, exceeding the 2s tier budget — without a prebuild the FIRST +measured queries race a half-built index and results flake run to run +(measured 2026-09-28: gate-exact repro rank 1, 3/3 with prebuild rank 1, +o1-gate run without prebuild rank None at wall=0.48s). + +Case table is imported from o1_holdout_gate (O1-holdout-v1, single source of +truth — no duplicated query list to rot). + +HARNESS LESSON (2026-09-28, verified): all queries run inside ONE event loop. +The o1-gate pattern (`asyncio.run()` per query) silently degrades every +even-positioned query to reranker passthrough (provider-None, reranker_ms=0): +asyncio primitives (locks/semaphores/client) bound to the first, now-closed +loop misbehave on alternating fresh loops. A per-query loop makes P2 (position +0) fail deterministically even with correct code. Reranker-cache discipline is +kept via explicit `searcher._reranker_cache.clear()` per case. + +PASS RULES: P2 rank==1; expect_no_boost cases unboosted; doc chunks never +boosted; no reranker-cache void measurements; no degraded rows +(reranker_ms falsy = reranker did not run = invalid measurement, fails loudly +instead of judging rank on passthrough order). +""" + +import argparse +import asyncio +import sys +import time +import traceback +from pathlib import Path + +if sys.stdout.encoding and sys.stdout.encoding.lower() != "utf-8": + try: + sys.stdout.reconfigure(encoding="utf-8") + except Exception: # noqa: BLE001 — encoding guard must never fail startup + pass + +ROOT = Path(__file__).resolve().parent.parent +if str(ROOT) not in sys.path: + sys.path.insert(0, str(ROOT)) + +from scripts.o1_holdout_gate import ( # noqa: E402 — imported for the case table + HOLDOUT, + P2, + build_searcher, + rank_of, +) + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("--project", default=str(ROOT)) + args = ap.parse_args() + project = Path(args.project) + try: + import subprocess # noqa: PLC0415 + + rev = ( + subprocess.run( + ["git", "rev-parse", "--short", "HEAD"], + cwd=str(ROOT), + capture_output=True, + timeout=10, + ) + .stdout.decode() + .strip() + ) + except Exception: # noqa: BLE001 — git metadata is best-effort diagnostics + rev = "unknown" + print(f"P2-anchor holdout gate | rev={rev} | project={project} | fresh process") + searcher = build_searcher(project) + rows = asyncio.run(_run_all(searcher)) + print("---") + fails = [] + p2 = rows[0] + if p2["rank"] != 1: + fails.append(f"P2 rank={p2['rank']} (expected 1)") + for row in rows: + case = next(c for c in [P2, *HOLDOUT] if c["id"] == row["id"]) + if case.get("expect_no_boost") and row["n_boost"]: + fails.append(f"{row['id']} unexpectedly boosted") + if case.get("expect_no_doc_boost") and row["n_doc_boost"]: + fails.append(f"{row['id']} doc chunk boosted") + if row["void"]: + fails.append(f"{row['id']} reranker-cache void (wall<2s, no timing)") + if row["degraded"]: + fails.append(f"{row['id']} reranker degraded (reranker_ms=0, passthrough)") + if fails: + print("GATE FAIL: " + "; ".join(fails)) + return 1 + print("GATE PASS") + return 0 + + +async def _run_all(searcher): + """Whole sequence on ONE event loop (see HARNESS LESSON above).""" + # Blocking FTS prebuild (discarded): cold build ~2.9s > 2s tier budget. + t0 = time.perf_counter() + await asyncio.to_thread(searcher._build_fts5_index) + print(f"FTS prebuild (discarded): {time.perf_counter() - t0:.2f}s") + # Discarded warm-up on a disjoint query (cache key includes query text). + await searcher.hybrid_search_async("warmup cold start primer", limit=3) + rows = [] + for case in [P2, *HOLDOUT]: + searcher._reranker_cache.clear() + t0 = time.perf_counter() + results = await searcher.hybrid_search_async(case["query"], limit=5) + wall = time.perf_counter() - t0 + boosted = [r for r in results if r.get("identifier_boost")] + doc_boosted = [ + r + for r in boosted + if str((r.get("metadata") or {}).get("file", "")) + .lower() + .endswith((".md", ".markdown", ".rst", ".txt", ".ipynb")) + ] + rank = rank_of(results, case["target"]) if case.get("target") else None + timing = dict(getattr(searcher, "_last_rerank_timing", None) or {}) + void = wall < 2.0 and timing == {} + degraded = not timing.get("reranker_ms") + rows.append( + { + "id": case["id"], + "rank": rank, + "target": case.get("target"), + "n_boost": len(boosted), + "n_doc_boost": len(doc_boosted), + "n": len(results), + "wall": round(wall, 2), + "void": void, + "degraded": degraded, + } + ) + print( + f"{case['id']:>3} rank={rank} n={len(results)} target={case.get('target')} " + f"boost={len(boosted)} doc_boost={len(doc_boosted)} wall={wall:.2f}s " + f"rerank_ms={timing.get('reranker_ms')} model={timing.get('model')} " + f"flags={None if searcher._multi_reranker is None else (searcher._multi_reranker.ollama_available, searcher._multi_reranker.llama_cpp_available)}" + ) + return rows + + +if __name__ == "__main__": + try: + sys.exit(main()) + except Exception: # noqa: BLE001 — gate must report, never crash silently + traceback.print_exc() + sys.exit(1) diff --git a/src/core/search/engine.py b/src/core/search/engine.py index 1cf010cb..2323cbc7 100644 --- a/src/core/search/engine.py +++ b/src/core/search/engine.py @@ -242,6 +242,89 @@ def _boost_rare_identifier(pool: List[dict], candidate_norm: str) -> List[dict]: return pool +# P2 pool-entry fix (2026-09-28): identifier-token anchoring for the pre-rerank +# pool. O1 boosts exact symbol matches INSIDE the pool, but for P2 the gold +# chunk (src/core/search/engine.py#hybrid_search_async) never ENTERS the pool: +# multi-term RRF buries it at BM25#126 / FTS#74 (measured live), while the cut +# is rrf_results[:limit] with raw_limit=min(limit*2,30). Widening the pool to +# depth 126 would cost ~126x0.4s rerank time — refuted by measurement. +# Anchors instead resolve exact-symbol chunks DIRECTLY (single-token FTS fetch, +# ~70ms warm) and append them pre-rerank, capped. Red-team defenses (inherited +# from O1's D1/D2 + doc-guard e8ffb671 + H2-collision lesson): +# A1 rarity cap — token anchored only if df <= _O1_ANCHOR_MAX_DF (BM25 stats); +# H2-style ubiquitous terms (BM25 df=327, FTS5 df=140) never anchor. +# A2 exact gate — full-token equality symbol == token, never substring; +# _is_doc_chunk exclusion (docs never anchor, even when they cite symbols). +# A3 cost cap — per-token and total anchor caps + pool never exceeds +# MAX_RERANKER_INPUT; FTS failure degrades to unanchored pool (never breaks +# search). Per-tier top-1 anchoring (PR52-branch idea) was refuted live: +# tier tops are polluted-junk (experiments/*.txt), gold sits at #74-126. +_O1_ANCHOR_MAX_DF = 120 +_O1_ANCHOR_FETCH_LIMIT = 30 +_O1_ANCHOR_PER_TOKEN = 2 +_O1_ANCHOR_TOTAL = 3 + + +# Anchor-eligible files are symbol definitions, which live in real code — +# never in docs or data (canary_set.json carries a bogus `hybrid_search_async` +# symbol from fallback scope-splitting; docs cite symbols without defining). +_ANCHOR_BLOCKED_EXTS = frozenset( + {".md", ".markdown", ".rst", ".txt", ".log", ".ipynb", + ".json", ".jsonl", ".csv", ".toml", ".yaml", ".yml"} +) + + +def _is_anchor_eligible(r: Dict) -> bool: + """Code file whose extension can host a symbol definition (P2/A2).""" + if _is_doc_chunk(r.get("metadata") or {}): + return False + fname = str((r.get("metadata") or {}).get("file", "") or "") + ext = "." + fname.rsplit(".", 1)[-1].lower() if "." in fname else "" + return ext not in _ANCHOR_BLOCKED_EXTS + + +def _is_symbol_definition(text: str, tok_norm: str) -> bool: + """True if the chunk text DEFINES the symbol (def/class line, P2/A2). + + Distinguishes the defining chunk (engine.py#hybrid_search_async) from + caller chunks that merely reference it (live_search_audit.py) — both + carry the same symbol_name in FTS metadata. + """ + try: + return bool(re.search( + rf"^\s*(?:async\s+def|def|class)\s+{re.escape(tok_norm)}\b", + text or "", re.MULTILINE | re.IGNORECASE, + )) + except re.error: # noqa: BLE001 — bad token never breaks search + return False + + +def _rare_identifier_tokens(query: str, df_of, max_df: int) -> List[str]: + """Identifier-shaped query tokens passing the rarity cap (P2/A1). + + Multi-token queries only (single identifiers are served by + _boost_exact_name_matches). Order-preserving, deduplicated, normalized. + """ + idents = _extract_identifier_tokens(query) + if not idents: + return [] + qtokens = _tokenize(query, _O1_TOKENIZER_RE) + if len(set(qtokens)) < 2: + return [] + out: List[str] = [] + for raw in idents: + norm = _identifier_norm(raw) + if norm in out: + continue + try: + df = int(df_of(norm) or 0) + except Exception: # noqa: BLE001 — broken stats never break search + continue + if df <= max_df: + out.append(norm) + return out + + def _prepend_code_name_matches( results: List[dict], pool: Optional[List[dict]], query: str, limit: int ) -> List[dict]: @@ -656,6 +739,85 @@ def df_of(term: str) -> int: logger.debug(f"[O1] rare-identifier boost: '{candidate}' x{_O1_BOOST_FACTOR:g}") return _boost_rare_identifier(pool, candidate) + def _df_of(self): + """Document-frequency callable over EXISTING BM25 stats (shared O1/anchors).""" + try: + self._build_bm25_index() + except Exception: # noqa: BLE001 — degraded BM25 never breaks search + return None + bm25 = getattr(self, "_bm25", None) or {} + if not bm25: + return None + + def df_of(term: str) -> int: + n = 0 + for doc_terms in bm25.values(): + if term in doc_terms: + n += 1 + return n + + return df_of + + async def _anchor_identifier_chunks_async( + self, pool: List[dict], query: str, limit: int + ) -> List[dict]: + """P2: anchor exact-symbol chunks for rare identifier tokens (defenses A1-A3). + + For each identifier-shaped query token with df <= _O1_ANCHOR_MAX_DF, a + single-token FTS fetch retrieves exact symbol matches that multi-term + RRF buried (P2 gold at FTS#74). Definition chunks outrank caller chunks + sharing the symbol_name; non-code files are ineligible. Matches are + appended pre-rerank (tail, no re-sort — the reranker re-sorts by its own + scores), deduplicated by file:chunk_index, capped at _O1_ANCHOR_TOTAL + and MAX_RERANKER_INPUT. Any failure degrades to the unanchored pool. + limit<=0 keeps the empty contract (no anchors on a zero pool). + """ + if limit <= 0: + return pool + df_of = self._df_of() + if df_of is None: + return pool + tokens = _rare_identifier_tokens(query, df_of, _O1_ANCHOR_MAX_DF) + if not tokens: + return pool + have = set() + for r in pool: + meta = r.get("metadata") or {} + have.add(f"{meta.get('file', '?')}:{meta.get('chunk_index', 0)}") + added = 0 + for tok in tokens: + if added >= _O1_ANCHOR_TOTAL or len(pool) >= MAX_RERANKER_INPUT: + break + try: + fetched = await self._fts5_search_async(tok, limit=_O1_ANCHOR_FETCH_LIMIT) + except Exception: # noqa: BLE001 — degraded FTS never breaks search + continue + per = 0 + # Definition-first ordering: the defining chunk (def/class line in + # code) outranks caller chunks sharing the same symbol_name and + # FTS order; non-code files (docs/data with bogus or citing + # symbols) are ineligible outright. + exact = [r for r in fetched + if _identifier_exact_match(r, tok) and _is_anchor_eligible(r)] + exact.sort(key=lambda r: (not _is_symbol_definition( + str(r.get("text", "") or ""), tok),)) + for r in exact: + if added >= _O1_ANCHOR_TOTAL or len(pool) >= MAX_RERANKER_INPUT: + break + if per >= _O1_ANCHOR_PER_TOKEN: + break + meta = r.get("metadata") or {} + key = f"{meta.get('file', '?')}:{meta.get('chunk_index', 0)}" + if key in have: + continue + pool.append(r) + have.add(key) + added += 1 + per += 1 + if added: + logger.debug(f"[P2-anchor] +{added} exact-symbol chunks for {tokens}") + return pool + async def hybrid_search_async( self, query: str, @@ -896,6 +1058,15 @@ async def hybrid_search_async( if tracer and _mmr_before: tracer.record_mmr(_mmr_before, pre_rerank_results, lambda_param=0.6) + # === P2 (2026-09-28): identifier-token anchors at PRE-RERANK pool === + # O1 boosts exact matches inside the pool but cannot rescue chunks that + # never enter it (P2 gold at BM25#126/FTS#74 vs [:limit] cut). Anchors + # resolve exact-symbol chunks directly (single-token FTS, ~70ms warm) + # and append them pre-rerank; O1 then boosts exact matches as usual. + pre_rerank_results = await self._anchor_identifier_chunks_async( + pre_rerank_results, query, limit + ) + # === O1 (2026-09-28): rare-identifier boost at PRE-RERANK pool (D3) === # Runs here — after sort+cut/MMR, before the reranker — never post-cut: # the x100 guarantees a rare exact symbol match survives the reranker's diff --git a/tests/test_p2_pool_anchors.py b/tests/test_p2_pool_anchors.py new file mode 100644 index 00000000..2c7bc9b4 --- /dev/null +++ b/tests/test_p2_pool_anchors.py @@ -0,0 +1,228 @@ +"""P2 pool-entry anchors: exact-symbol chunks enter the pre-rerank pool. + +Root cause (measured live 2026-09-28, fresh process + discarded warm-up): +P2 query `hybrid_search_async reciprocal_rank_fusion FTS5 BM25`, gold +src/core/search/engine.py — BM25#126, FTS#74, dense absent @200. Pool cut is +rrf_results[:limit] with raw_limit=min(limit*2,30): gold never enters at any +limit<=50. Widening to depth 126 would cost ~126x0.4s rerank (~50s) — refuted. +O1 alone cannot rescue it either: O1 candidate is `reciprocal_rank_fusion` +(df=4, strictly rarest) while gold's symbol is `hybrid_search_async` +(df=100) — no exact match, no boost (live: O1 fired on scoring.py instead). + +Fix: _anchor_identifier_chunks_async — single-token FTS fetch per rare +identifier token (df<=_O1_ANCHOR_MAX_DF), exact-symbol + non-doc filter, +appended pre-rerank, capped. All tests use mocked _fts5_search_async (no live +services); live validation via scripts/o1_holdout_gate.py (fresh process). +""" + +from __future__ import annotations + +from src.core.search.engine import ( + _O1_ANCHOR_MAX_DF, + _O1_ANCHOR_TOTAL, + _is_anchor_eligible, + _is_symbol_definition, + _rare_identifier_tokens, +) + +GOLD = "src/core/search/engine.py" +P2Q = "hybrid_search_async reciprocal_rank_fusion FTS5 BM25" + + +def _chunk(path, sym, idx=0, doc=False): + meta = {"file": path, "chunk_index": idx, "symbol_name": sym} + if doc: + meta["file"] = "docs/NOTE.md" + return {"text": f"def {sym}(): ...", "metadata": meta, "final_score": 0.01} + + +def _df_of(mapping): + def df(t): + return mapping.get(t, 0) + + return df + + +LIVE_DF = { + "hybrid_search_async": 100, + "reciprocal_rank_fusion": 4, + "fts5": 140, + "bm25": 327, + "hybrid": 79, + "search": 1633, + "async": 790, +} + + +def test_rare_tokens_include_gold_symbol_exclude_h2_common(): + toks = _rare_identifier_tokens(P2Q, _df_of(LIVE_DF), _O1_ANCHOR_MAX_DF) + assert "hybrid_search_async" in toks + assert "reciprocal_rank_fusion" in toks + assert "bm25" not in toks # H2-collision guard (df=327) + assert "fts5" not in toks # common acronym (df=140) + + +def test_rare_tokens_single_term_empty(): + assert _rare_identifier_tokens("get_db", _df_of({"get_db": 2}), _O1_ANCHOR_MAX_DF) == [] + + +def test_rare_tokens_no_identifiers_empty(): + assert ( + _rare_identifier_tokens("quantum computing configuration", _df_of({}), _O1_ANCHOR_MAX_DF) + == [] + ) + + +class _FakeSearcher: + """Minimal harness around the real anchor method (no index needed).""" + + def __init__(self, fetched): + from src.core.search.engine import Searcher + + self._anchor = Searcher._anchor_identifier_chunks_async.__get__(self) + self._df = _df_of(LIVE_DF) + self._fetched = fetched + + def _df_of(self): + return self._df + + async def _fts5_search_async(self, query, limit=10): + return list(self._fetched.get(query, [])) + + +def _run(coro): + import asyncio + + return asyncio.run(coro) + + +def test_pool_contains_gold_p2(): + """P2 assertion: gold enters the pool via anchors even when RRF cut it.""" + gold = _chunk(GOLD, "hybrid_search_async", idx=18) + pool = [_chunk("experiments/probe_0.txt", "probe", idx=0)] + s = _FakeSearcher( + { + "hybrid_search_async": [_chunk("src/other.py", "hybrid_search_async", idx=1), gold], + "reciprocal_rank_fusion": [ + _chunk("src/core/search/scoring.py", "reciprocal_rank_fusion", idx=0) + ], + } + ) + out = _run(s._anchor(pool, P2Q, 5)) + files = [(r["metadata"]["file"], r["metadata"]["chunk_index"]) for r in out] + assert (GOLD, 18) in files + + +def test_doc_chunks_never_anchor(): + pool = [_chunk("experiments/probe_0.txt", "probe", idx=0)] + s = _FakeSearcher( + {"hybrid_search_async": [_chunk("docs/X.md", "hybrid_search_async", idx=3, doc=True)]} + ) + out = _run(s._anchor(pool, P2Q, 5)) + assert len(out) == 1 # doc exact-symbol match must not anchor + + +def test_data_files_never_anchor(): + """canary_set.json carries a bogus exact symbol (fallback scope-split) — + data extensions are ineligible even with an exact symbol match.""" + victim = { + "text": "def hybrid_search_async(): ...", + "metadata": { + "file": "src/providers/embedder/canary_set.json", + "chunk_index": 0, + "symbol_name": "hybrid_search_async", + }, + "final_score": 0.04, + } + assert not _is_anchor_eligible(victim) + pool = [] + s = _FakeSearcher({"hybrid_search_async": [victim]}) + assert _run(s._anchor(pool, P2Q, 5)) == [] + + +def test_definition_outranks_caller(): + """Caller chunk listed first in FTS order, but the def chunk anchors first; + with per-token room both enter, def first.""" + caller = { + "text": " await searcher.hybrid_search_async(q)", + "metadata": { + "file": "scripts/live_search_audit.py", + "chunk_index": 4, + "symbol_name": "hybrid_search_async", + }, + "final_score": 0.05, + } + definition = { + "text": " async def hybrid_search_async(self,", + "metadata": {"file": GOLD, "chunk_index": 18, "symbol_name": "hybrid_search_async"}, + "final_score": 0.04, + } + assert _is_symbol_definition(definition["text"], "hybrid_search_async") + assert not _is_symbol_definition(caller["text"], "hybrid_search_async") + pool = [] + s = _FakeSearcher({"hybrid_search_async": [caller, definition]}) + out = _run(s._anchor(pool, P2Q, 5)) + assert out[0]["metadata"]["file"] == GOLD + assert (GOLD, 18) in [(r["metadata"]["file"], r["metadata"]["chunk_index"]) for r in out] + + +def test_substring_symbol_does_not_anchor(): + pool = [] + s = _FakeSearcher( + {"hybrid_search_async": [_chunk("src/a.py", "my_hybrid_search_async_wide", idx=0)]} + ) + out = _run(s._anchor(pool, P2Q, 5)) + assert out == [] # full-token equality only, never substring + + +def test_common_token_fetch_never_anchored_even_if_exact(): + """H2-collision: even if FTS returned an exact match for a common token, + the rarity cap excludes the token before any fetch.""" + pool = [] + s = _FakeSearcher({"bm25": [_chunk("src/core/search/bm25.py", "bm25", idx=0)]}) + out = _run(s._anchor(pool, P2Q, 5)) + assert out == [] # 'bm25' df=327 > cap: no fetch, no anchor + + +def test_total_cap_and_dedup(): + pool = [_chunk(GOLD, "hybrid_search_async", idx=18)] + many = [_chunk(f"src/f{i}.py", "hybrid_search_async", idx=i) for i in range(10)] + s = _FakeSearcher( + { + "hybrid_search_async": many, + "reciprocal_rank_fusion": [ + _chunk("src/core/search/scoring.py", "reciprocal_rank_fusion", idx=0) + ], + } + ) + out = _run(s._anchor(pool, P2Q, 5)) + assert len(out) <= 1 + _O1_ANCHOR_TOTAL + keys = [(r["metadata"]["file"], r["metadata"]["chunk_index"]) for r in out] + assert len(keys) == len(set(keys)) + + +def test_limit_zero_keeps_empty_contract(): + s = _FakeSearcher({"hybrid_search_async": [_chunk(GOLD, "hybrid_search_async", idx=18)]}) + assert _run(s._anchor([], P2Q, 0)) == [] + + +def test_fts_failure_degrades_to_pool(): + pool = [_chunk("experiments/probe_0.txt", "probe", idx=0)] + + class _Broken(_FakeSearcher): + async def _fts5_search_async(self, query, limit=10): + raise RuntimeError("fts down") + + s = _Broken({}) + assert _run(s._anchor(pool, P2Q, 5)) == pool + + +def test_no_bm25_stats_degrades_to_pool(): + pool = [_chunk("experiments/probe_0.txt", "probe", idx=0)] + + class _NoStats(_FakeSearcher): + def _df_of(self): + return None + + s = _NoStats({}) + assert _run(s._anchor(pool, P2Q, 5)) == pool