Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions KNOWN_ISSUES.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,14 @@

---

## 2026-09-29 — FTS-тир молча выпадает по таймауту 2s, ранки плавают run-to-run (Open)

- **Локация:** `src/core/search/engine.py:970-976` (`wait_for(..., timeout=2.0)` вокруг `_fts5_search_async`; `TimeoutError` → tier `[]` + warning в лог).
- **Симптом:** членство FTS-тира в RRF-пуле недетерминировано от прогона к прогону: наблюдалось P2 rank 1↔2, воббл H6/H11/H12 на том же индексе и коде. Холодный FTS-билд (~0.5–2.5s lazy `to_pandas`) съедает бюджет целиком — первый замер в свежем процессе систематически без FTS-тира (частный случай уже учтён warm-up-правилом в записи 2026-09-28, но теплый индекс тоже флипает у границы 2s).
- **Почему это ловушка для гейтов:** silent-degraded — ни флаг в выдаче, ни строка в harness-таблице; воббл выглядит как эффект кода, а не бюджета. `void`-флаг (`timing=={}`) и degraded-флаг (`reranker_ms==0`) этот класс НЕ ловят.
- **Fix options (решение владельца):** (a) явный `fts_timed_out` флаг в результат/трейсер + degraded-строка в гейтах; (b) бюджет/квота вместо жёсткого капа (адаптивный timeout, повтор с урезанным лимитом); (c) прогрев FTS до замеров как обязательный шаг harness (уже частично: discarded warm-up).
- **Статус:** 🟡 Open (процедурное правило до фикса: gate-серия обязана идти в одном процессе с discarded warm-up + фиксировать число FTS-таймаутов по логу `FTS5 search timed out`).

## 2026-09-28 — Pre-commit hook fail-open при потере маркеров (Fixed)

- **Локация:** `.githooks/pre-commit:31-53` (`find_project_root` + `run_script`).
Expand Down
37 changes: 37 additions & 0 deletions experiments/reranker_p3/PREREG.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# P3 module-head anchor — preregistration (2026-09-29)

Status: PREREGISTERED — written BEFORE implementing the anchor. No numbers
invented here; thresholds/caps referenced are the existing ones
(`_O1_ANCHOR_TOTAL`, `MAX_RERANKER_INPUT=30`, `MIN_RERANK_SCORE=0.3` untouched).

## Hypothesis

P3 (`src/core/artifact_gc.py` gold file below threshold) is a pool-entry /
ranking problem of the same family as P2: the chunk the reranker scores
highest (module-head/docstring) never reaches the reranker, while the code
chunk that does reach it scores a negative logit and dies at the threshold.

## Decision rule (verbatim)

"anchor a file's module-head/docstring chunk into the pre-rerank pool iff it
scores above threshold on holdout queries; falsified if (i) docstring chunks
do not outscore code chunks on holdout, (ii) anchoring regresses holdout H-set,
(iii) pool cost exceeds caps".

## Reading of the three clauses

- (i) is checked by the persisted probes
(`rerank_probe_run1.json`: gold_doc #1 on 6/6 holdout-style queries;
`rerank_probe_run2.json`: module_head top on P3/R2, code chunks negative).
- (ii) is checked by the no-regression holdout H1–H12 + P/R + doc-control
gate (fresh-process discipline — reranker-cache trap).
- (iii) is checked structurally: anchors bounded by the existing caps
(`_O1_ANCHOR_TOTAL`, `MAX_RERANKER_INPUT=30`).

## Scope

- Validate on holdout only — NEVER the frozen eval-16.
- `MIN_RERANK_SCORE=0.3` untouched (threshold sweep on the eval set was
explicitly refuted in `EXPERIMENTS_LOG.md:2815-2816`).
- Minimal diff extending the P2 `_anchor_identifier_chunks_async` precedent
(`engine.py:772-805`), same style.
46 changes: 46 additions & 0 deletions experiments/reranker_p3/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# reranker_p3 probes — module-head / docstring anchor evidence

Persisted 2026-09-29 from `%TEMP%\opencode\` (byte-identical copies, hashes in SHA256SUMS).

## What

Two fresh-process direct-reranker probes that bypass retrieval and POST
straight to the reranker. Finding: the gold module-head/docstring chunk of
`src/core/artifact_gc.py` outscores every code chunk on holdout-style queries,
while the `_prune`-body code chunk scores a negative logit (P3 -0.99, R2 -2.60
in the eval path). Motivation for the P3 module-head anchor: guarantee the
docstring chunk a place in the pre-rerank pool.

## When / how run

- Both scripts stdlib-only (`urllib`), run as `python <script> <out.json>`.
- Model: BGE-M3 reranker served on `http://127.0.0.1:8081` (`/rerank`;
probe 1 docstring says `/v1/rerank`, actual POST path in code is `/rerank`).
- No frozen eval-16 queries were used: probe queries are holdout-style
P3/R2 variants (`ArtifactGC _cleanup_old_projects 30d 90d 7d retention_policy`
etc.), disjoint from the frozen eval set.

## Files

- `probe_rerank.py` — probe 1: 6 queries x 5 passages (gold_doc head 800 chars
of `src/core/artifact_gc.py`, gold_code `_prune` body chars [1500:2300],
distractors engine/settings/diary heads). Output `rerank_probe_run1.json`.
- `probe_rerank2.py` — probe 2: every AST-plausible chunk of `artifact_gc.py`
(module_head + per-def chunks with scope header, 800 chars) on P3/R2 queries.
Output `rerank_probe_run2.json`.
- `rerank_probe_run1.json` — 6/6 queries rank `gold_doc` #1 (sigmoid 0.69–0.99).
- `rerank_probe_run2.json` — `module_head` top on both queries
(P3 +0.57, R2 +0.43); all code chunks negative except `prune_stale_artifacts`
on R2 (+0.27).

## Citation rule

Rerank numbers are cited as
`experiments/reranker_p3/rerank_probe_run{1,2}.json` + `EXPERIMENTS_LOG.md:2809`
(raw eval logits P3 -0.99 / R2 -2.60 on the pre-rerank pool) — never
`judged_raw.json` for rerank scores.

## Frozen prompt

None — queries are listed verbatim in the scripts and in `rerank_probe_run1.json`
(`queries.<id>.query`). No LLM judge involved.
85 changes: 85 additions & 0 deletions experiments/reranker_p3/RESULT-chunker.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
# Chunker module-docstring fix — live gate result (2026-09-29): RED (mechanism delivered, retrieval blocks)

Branch: `fix/chunker-module-doc` (off `origin/main` + cherry-picked P3 stack:
anchor `d3d53075`, sigmoid `f2df8160`). Task: emit module docstring as chunk 0.

## 1. Mechanism (implemented, unit-verified)

`src/core/indexing/parser.py`: `_extract_module_docstring` (stdlib `ast`,
.py only) + `chunks.insert(0, …)` in `_parse_with_tree_sitter` (only when the
walk already produced chunks — fallback path untouched), `hierarchy_map`
`module_docstring → module`, symbol `__module_doc__`. Files without a
docstring: byte-identical behavior. Symbols untouched (no graph pollution).

`tests/test_chunker_module_doc.py`: 5 passed. Related suites green:
`test_parser.py` 5, `test_p3_module_head_anchor.py` 9,
`test_chunk_cache.py` + `test_move_chunks.py` + `test_scm_definitions.py` 45.
`ruff check` clean on both files (`ruff format` not enforced repo-wide —
all pre-existing files fail it too; new code matches file style).

## 2. Live reindex (evidence)

- BEFORE: 15,426 rows, 0 `module_docstring` rows; `artifact_gc.py` 6 chunks,
chunk 0 = `function_definition` (`def _dir_has_files`).
- Method: wipe `delete("1 = 1")` (no rmtree — live servers hold the DB open)
+ `indexer.index_project(project)` with working-tree code, embedder :8080.
- AFTER (repaired, see §4): **14,839 rows**, **371 `module_docstring` chunks**,
0 dup `(file_path, chunk_index)`; `artifact_gc.py` 7 chunks,
chunk 0 = `module_docstring` (start_line 1, head = file docstring).
- Direct rerank of the new chunk 0 on the P3 query: **logit +1.21 / sig 0.77**
(probe predicted +0.83; old chunk 0 scored −6.66 and the probe passage did
not exist). Chunk 1 scores −6.52 → correctly cut. The chunker half of the
drill-down is closed.

## 3. Gate (fresh process, `scripts/p3_holdout_gate.py`, rev `f2df8160`)

```
P2 rank=2 n=5 (expected 1 — REGRESSION vs FALSIFIED baseline rank 1)
P3 rank=None n=2 (expected 1) | R2 rank=3 n=3 (expected 1; was None)
H1 3 n=3 | H2 None n=1 | H3 3 n=3 | H4 2 n=2 | H5 1 n=5 | H6 3 n=5
H7 None n=1 | H8 1 n=5 | H9 None n=5 | H10 None n=1 | H11 1 n=4 | H12 2 n=3
N1 None (clean) | DOC None (clean)
```

Controls clean throughout: no unexpected boosts, no doc boosts, no
reranker-cache voids, no degraded rows, no timeouts
(`model=llama.cpp-reranker` every row — sigmoid fix confirmed live).

## 4. Analysis (why RED)

1. **Retrieval never surfaces any `artifact_gc.py` chunk for the P3 query.**
FINAL n=2 are both JSON experiment artifacts
(`reranker_p3/rerank_probe_run1.json:0`, `noderag/results/results.json:1`).
The P3 query's distinctive terms (`_cleanup_old_projects`,
`retention_policy`) occur verbatim in OUR OWN probe/result JSON
(verified by grep) and nowhere in `artifact_gc.py` — they dominate BM25
while the 0.77 doc chunk sits outside the pool. The P3 pool-anchor can
only anchor chunk 0 of files *already in the pool*; with zero pool chunks
from the target file it cannot fire. Design-level finding, not an
implementation bug. Fixing it means broadening retrieval/anchor scope —
beyond this task's minimal chunker brief.
2. **P2 rank 2 is run-to-run noise, not a chunker regression.** A follow-up
fresh-process P2 run put `engine.py:35` back at rank 1; that run logged
`FTS5 search timed out (>2s), skipping FTS5 tier` — tier membership
flip-flops (~2.18s prebuild vs 2s budget) and moves ranks 1↔2. Same noise
class explains H6 2→3 / H11 2→1 / H12 None→2 wobbles. H-set: no systematic
demotion (H5=1, H8=1, H11=1 hold; R2 None→3 and H12 None→2 improved).
3. **Infra incidents during this run (caveats):**
- Double-write: post-reindex table held 29,134 rows (every chunk twice,
identical ids). In-run mechanism unproven; repaired deterministically
(pandas dedup by id → wipe → single re-add, round-trip validated on a
scratch table first: vectors/text identical). AFTER = 14,839 verified.
- `_safe_ivf_index` (optimize/create_index) aborted: PID lock held by live
MCP servers (pid 2944/13172). Table now has NO vector index (flat exact
scan — correct, slightly slower; fine for the gate). Owner should run
`intel_trigger_reindex(full)` or Reload-Window Zed so the servers
re-open the table (their handles predate the wipe) and finalize IVF.
- Unrelated dirt `experiments/planted_break/results.json` was already
modified before this task — left untouched, unstaged.

## Verdict

RED per gate rule (P3 None, P2 2, R2 3 — all expected 1). No PR opened.
Chunker fix stands on its own (unit + live verified); P3 rank 1 needs a
retrieval-side follow-up (anchor scope / BM25 competition from experiment
JSON artifacts), proposed as the next drill-down level, not this branch.
56 changes: 56 additions & 0 deletions experiments/reranker_p3/RESULT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
# P3 module-head anchor — live result (2026-09-29): FALSIFIED

Mechanism implemented per `PREREG.md` (`_anchor_module_head_chunks_async`,
`engine.py`, P2-style, caps `_O1_ANCHOR_TOTAL` / `MAX_RERANKER_INPUT=30`,
`MIN_RERANK_SCORE=0.3` untouched). Live holdout validation
(`scripts/p3_holdout_gate.py`, fresh process, FTS prebuild + discarded warm-up
+ per-case reranker-cache clear, BGE-M3 on :8081, embed on :8080) REJECTS the
prereg expectation: **P3 rank None (expected 1), R2 rank None (expected 1)**.

## What was verified live (mechanism works as coded)

Pre-rerank pool capture for the P3 query
(`ArtifactGC _cleanup_old_projects 30d 90d 7d retention_policy`): pool n=8,
`src/core/artifact_gc.py:0` present with `module_head_anchor=True`, pool cap
respected. The anchor does what the spec says.

## Why the hit does not happen (two load-bearing findings)

1. **chunk_index 0 is not a docstring chunk — for ANY file.** The AST chunker
drops module-level docstrings index-wide (verified: `artifact_gc.py` has 6
chunks, all function-scoped; chunk 0 = `def _dir_has_files`, len 279;
`engine.py` chunk 0 = `def _cache_key`; `graph.py`, `llama_runner.py`
likewise first-function). The probe passage that scores +0.83 (`gold_doc`,
first 800 chars of the file) does not exist as an indexed chunk. The real
chunk 0 scores **-6.66** on the P3 query — correctly cut. No pool-anchor can
place a chunk the index never built. Prereg clause (i)/(ii) fire.
2. **The EXPERIMENTS_LOG:2813 sigmoid fix is absent from the code.**
`grep sigmoid src/` = zero hits; the llama_cpp path stores raw logits in
`reranker_score` (`multi_provider.py:500-502`, `apply_scores`) and filters
them against `MIN_RERANK_SCORE=0.3` (`multi_provider.py:670`). Live winner
for P3 was a junk JSON chunk at logit +1.65; FINAL n=1. Out of scope here
(threshold untouched by design), but it explains the all-None holdout rows.

## Full holdout table (branch, rev e5756f55 + anchor)

```
P2 rank=1 n=2 (no P2 regression)
P3 rank=None n=1 | R2 rank=None n=1
H1 None n=1 | H2 None n=1 | H3 None n=1 | H4 2 n=2 | H5 1 n=5 | H6 2 n=3
H7 None n=1 | H8 1 n=5 | H9 None n=5 | H10 None n=1 | H11 2 n=2 | H12 None n=1
N1 None (negative control, clean) | DOC None (doc-control, clean)
```

Controls clean throughout: no unexpected boosts, no doc boosts, no
reranker-cache voids, no degraded rows (reranker ran every query,
`model=llama.cpp-reranker`). H-rank Nones are threshold cuts (finding 2),
not anchor demotions — the anchor is append-only pre-rerank and cannot demote
a chunk that passes the threshold.

## Verdict

FALSIFIED per prereg rule. No PR opened (nothing to propose). Next step, if
wanted, is index-level, not pool-level: emit a module-docstring chunk at index
time (parser/indexer change + reindex), then re-run this gate. The mocked
suite (`tests/test_p3_module_head_anchor.py`, 9 passed) and the gate script
stay on the branch as the reusable harness for that attempt.
6 changes: 6 additions & 0 deletions experiments/reranker_p3/SHA256SUMS
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
0a8718e9c1293f85c7493fb4062bfcbad9075be2d3f3851b990a8be67320ff7e rerank_probe_run1.json
e40f153d8638d911024eefc476be7d534ba79d5ca04047daff5115d0e0c9b34b rerank_probe_run2.json
fcaf3a725af6c366ace02292a3226d741797b62e120cef4fcc16a1ba4433177c probe_rerank.py
706f64b4cd6538cf637c7c42b93d0a94884d638c528e849c864ecbe421a723f4 probe_rerank2.py
1bb952b9f911d32a88dc44ce2dc2fd8ef584405c78fcedaf53aa220f5369a6f6 README.md
ea9681e6dbef57fed0b2a5010a38debf70dad8de9c6ba1e113d1ac765d1f3307 PREREG.md
71 changes: 71 additions & 0 deletions experiments/reranker_p3/probe_rerank.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
"""Fresh-process direct reranker probe for OPEN item #3 (P3/R2 gold below threshold).
Bypasses retrieval: POSTs straight to llama.cpp /v1/rerank on :8081.
Stdlib only. Raw JSON -> same dir as this script (%TEMP%\\opencode\\).
"""
import io, json, math, sys, urllib.request

sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")

BASE = "http://127.0.0.1:8081"
OUT = sys.argv[1] if len(sys.argv) > 1 else "rerank_probe_out.json"

REPO = "D:\\Project\\MSCodeBase"


def load(rel, n=800):
with open(REPO + "\\" + rel, encoding="utf-8") as f:
return f.read()[:n].strip()


def score(query, passages):
payload = json.dumps({"query": query, "texts": passages}).encode()
req = urllib.request.Request(BASE + "/rerank", data=payload,
headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=120) as r:
return json.loads(r.read().decode())


def sig(x):
return 1.0 / (1.0 + math.exp(-x))


GOLD_DOC = load("src\\core\\artifact_gc.py", 800) # module docstring head
GOLD_CODE = open(REPO + "\\src\\core\\artifact_gc.py", encoding="utf-8").read()[1500:2300] # _prune body
D_ENGINE = load("src\\core\\search\\engine.py", 800)
D_SETTINGS = load("src\\config\\settings.py", 800)
D_DIARY = load("AGENT_DIARY.md", 800)

QUERIES = {
"P3_orig": "ArtifactGC _cleanup_old_projects 30d 90d 7d retention_policy",
"R2_orig": "ArtifactGC 30d projects 90d telemetry 7d logs retention",
"short": "ArtifactGC",
"keyword": "ArtifactGC cleanup old projects retention days",
"defform": "What is the ArtifactGC retention policy for old projects telemetry logs",
"real_symbols": "ArtifactGC prune_stale_artifacts project_max_age_days telemetry_max_age_days",
}

PASSAGES = {"gold_doc": GOLD_DOC, "gold_code": GOLD_CODE, "distr_engine": D_ENGINE,
"distr_settings": D_SETTINGS, "distr_diary": D_DIARY}

print(f"gold_doc len={len(GOLD_DOC)} gold_code len={len(GOLD_CODE)} "
f"max_passage={max(len(p) for p in PASSAGES.values())}", flush=True)

result = {"queries": {}, "passage_lens": {k: len(v) for k, v in PASSAGES.items()}}
for qname, q in QUERIES.items():
names = list(PASSAGES)
raw = score(q, [PASSAGES[k] for k in names])
logits = [d["score"] for d in sorted(raw, key=lambda d: d["index"])]
ranked = sorted(zip(names, logits), key=lambda t: t[1], reverse=True)
result["queries"][qname] = {
"query": q,
"logits": {n: lg for n, lg in zip(names, logits)},
"sigmoid": {n: round(sig(lg), 4) for n, lg in zip(names, logits)},
"rank": [n for n, _ in ranked],
}
print(f"== {qname} ==", flush=True)
for n, lg in ranked:
print(f" {n:14s} logit={lg:+.4f} sig={sig(lg):.4f}", flush=True)

with open(OUT, "w", encoding="utf-8") as f:
json.dump(result, f, indent=1)
print(f"WROTE {OUT}", flush=True)
50 changes: 50 additions & 0 deletions experiments/reranker_p3/probe_rerank2.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
"""Probe run 2: score every AST-plausible chunk of artifact_gc.py on P3/R2 queries.
Identifies which gold chunk reproduces eval logit -0.99 (P3) / -2.60 (R2).
"""
import io, json, math, re, sys, urllib.request

sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8")
BASE = "http://127.0.0.1:8081"
OUT = sys.argv[1]
SRC = open("D:\\Project\\MSCodeBase\\src\\core\\artifact_gc.py", encoding="utf-8").read()


def score(query, passages):
payload = json.dumps({"query": query, "texts": passages}).encode()
req = urllib.request.Request(BASE + "/rerank", data=payload,
headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=120) as r:
return json.loads(r.read().decode())


def sig(x):
return 1.0 / (1.0 + math.exp(-x))


# Split into module-head + per-def chunks, mimicking AST chunking w/ scope header
parts = re.split(r"(?m)^(?=def |^class )", SRC)
chunks = {}
for i, p in enumerate(parts):
m = re.match(r"(def |class )(\w+)", p)
name = m.group(2) if m else "module_head"
hdr = f"// Scope: other | function | src.core.artifact_gc\n" if m else "// Scope: module\n"
chunks[name] = (hdr + p)[:800].strip()

print(f"n_chunks={len(chunks)} lens=" + ",".join(f"{k}:{len(v)}" for k, v in chunks.items()), flush=True)
QUERIES = {
"P3_orig": "ArtifactGC _cleanup_old_projects 30d 90d 7d retention_policy",
"R2_orig": "ArtifactGC 30d projects 90d telemetry 7d logs retention",
}
result = {}
for qname, q in QUERIES.items():
names = list(chunks)
raw = score(q, [chunks[k] for k in names])
logits = [d["score"] for d in sorted(raw, key=lambda d: d["index"])]
print(f"== {qname} ==", flush=True)
for n, lg in sorted(zip(names, logits), key=lambda t: t[1], reverse=True):
print(f" {n:22s} logit={lg:+.4f} sig={sig(lg):.4f}", flush=True)
result[qname] = {n: lg for n, lg in zip(names, logits)}

with open(OUT, "w", encoding="utf-8") as f:
json.dump(result, f, indent=1)
print(f"WROTE {OUT}", flush=True)
Loading
Loading