diff --git a/AGENT_DIARY.md b/AGENT_DIARY.md index 87051b5c..157c5ea8 100644 --- a/AGENT_DIARY.md +++ b/AGENT_DIARY.md @@ -1,5 +1,7 @@ ## Key Historical Decisions +- **F5 relang clean-stack + purge + PR52-resolution (2026-09-29):** индекс был на 20.3% из мусора (3127/15426 чанков, `experiments/**/results|work`) → новый слой `SystemArtifacts.is_experiment_output` + purge 772 файлов/2152 чанков (штатный prune отказал бы: 52.4% файлов > safety-guard 50%), проверка — 0 осталось. Чистый замер B×5: **RU 26/80=32.5% vs EN 30/80=37.5%, CI пересекаются — эффекта языка нет**; 6/16 запросов флипаются all-or-nothing (язык меняет какие, не сколько). Конфликтный PR #52 закрыт как superseded: tier-anchor пропущен (P2 уже закрыт #54 в той же точке), спасены сигмоида/top-N/holdout-калибровка (PR #63); мои FTS-hoist+guard cherry-pick в PR #62. Артефакты: `results/f5relang/`, `scripts/purge_experiment_outputs.py`. + - **stale_after + discriminator for memory notes (2026-09-27):** `src/core/intelligence/staleness.py` + `store.check_staleness()` + CLI. 17/17 tests. stale_after (date) → STALE; discriminator (command, exit≠0) → EXPIRED. Backward compat (no fields → ACTIVE). Implements final promise from hooks thread. Артефакты: `experiments/stale_after/`. - **redact.py: scrub personal paths (2026-09-27):** `scripts/redact.py` + `tests/test_redact.py` (5/5 pass). Drive paths → ``, username → ``, clean text unchanged. Implements "redact.py at delivery points + planted key test" promise. Артефакты: `scripts/redact.py`, `tests/test_redact.py`. diff --git a/KNOWN_ISSUES.md b/KNOWN_ISSUES.md index f1b9c326..1f2f7a16 100644 --- a/KNOWN_ISSUES.md +++ b/KNOWN_ISSUES.md @@ -5,6 +5,13 @@ --- +## 2026-09-29 — Индекс вычищен от мусора + relang: эффекта языка нет (Fixed/Closed) + +- **Purge (Fixed):** 772 файла / 2152 чанка (`experiments/**/results|work`, было 20.3% индекса) удалены one-time скриптом `scripts/purge_experiment_outputs.py` (штатный prune отказал бы: 52.4% файлов > safety-guard 50%). Проверка: 0 осталось. Guard на будущее — PR #62 (`SystemArtifacts.is_experiment_output`). +- **Relang (Closed):** B×5 на чистом стеке — RU 26/80=32.5% vs EN 30/80=37.5%, CI пересекаются → эффекта языка нет. 6/16 запросов флипаются all-or-nothing (язык меняет какие, не сколько). Старый EN-замер на сломанном стеке невалиден. Артефакты: `results/f5relang/`, `f5/RESULTS_RELANG.md`. +- **PR #52 (Closed как superseded):** tier-anchor пропущен (P2 закрыт #54 в той же точке); спасены сигмоида/top-N/holdout-калибровка → PR #63. FTS-hoist+guard → PR #62. +- **Objective (Done 2026-09-29):** перемер на чистом индексе (`results/f5/objective_clean.json`) — A hit@1 4/16, hit@3 5/16, hit@10 6/16 (=), B top-1 4/16. Топ двинут на 1 запрос (шум n=16): purge значимо не повлиял. + ## 2026-09-28 — Шкала реранкера + top-N floor (salvage из PR #52, tier-anchor пропущен) - **Спасено из конфликтного PR #52:** `_sigmoid`-нормализация логитов llama.cpp → [0,1] (без неё MIN_RERANK_SCORE=0.3 отсекал 70–97% выдачи), top-N recall floor `reranker_topn_keep` (default 0 = выключено), holdout-калибровка порога с запретом eval-источников кодом. Guard: 4 sigmoid-теста + 13 тестов top-N/калибровки. @@ -43,6 +50,13 @@ **24 entries** — compressed per §4.8 R3 (conclusion-first; dedup 2026-09-08, 2026-09-21). Closed entries moved to docs/archive/KNOWN_ISSUES_2026_09.md on 2026-09-27 (R1 size guard; second batch on merge experiment/4a-unit-of-return). +## 2026-09-28 — Холодный FTS-билд превышал 2s-бюджет и молча выпадал (Fixed) + _get_ext_dir указывал в src/ (Fixed) + +- **FTS (c, flaky A/B):** замер — холодный `to_pandas`-билд всего индекса = **2.47s > 2.0s** `wait_for` в `engine.py:671`. Первый поиск в свежем процессе молча терял FTS-тир → пилот 18/20 vs 8/20 на тех же запросах. **Fix:** build вынесен из-под таймаута (идемпотентен, double-checked lock), 2s остались только на сам поиск (~0.05s). Guard `test_fts5_timeout_does_not_break_search` зелёный. +- **llama-пути (b):** `llama_install.py:_get_ext_dir` брал 3 `parent` от `__file__` вместо 4 → указывал в `src/`, ветка «режим разработки» была мёртвой, модели резолвились в пустой `%LOCALAPPDATA%/mscodebase/models`. На вопрос «падает или не успевает»: после простоя restart **пытается** (`idle-unload recovery`), но падал по отсутствию файлов, не по таймингу. **Fix:** off-by-one исправлен + `multilingual-e5-small-Q8_0.gguf` (132MB) докопирован из расширения в `models/` (git-ignored). Live-check `smoke_e2e.py`: **SMOKE E2E PASSED** (embed dim=384, rerank top=1, поиск по индексу). +- **Побочно (Verified, не чинено — решение владельца):** холодный топ захламлён артефактами (`judged_raw*.json`, `work/ctx_*.txt` в выдаче) — живое подтверждение индексного мусора (P2-смежное). Чистка индекса сменит ретрив-базисы. +- **Статус:** ✅ Fixed (пути + FTS-холод). + ## 2026-09-27 — Import-time os.environ mutation in scripts breaks xdist workers (Fixed) - **Симптом:** 6 plugin-тестов (`test_plugins_subprocess/registry`) падали под `-n auto` с `ModuleNotFoundError: No module named 'src'` в runner-subprocess, серийно (`-n0`) — зелёные. diff --git a/experiments/4A_unit_of_return/f5/RESULTS.md b/experiments/4A_unit_of_return/f5/RESULTS.md index e06335a8..7d7fd732 100644 --- a/experiments/4A_unit_of_return/f5/RESULTS.md +++ b/experiments/4A_unit_of_return/f5/RESULTS.md @@ -4,6 +4,15 @@ G6 `OVERLAP: PASS`; автор — context-clean субагент). **Harness:** `scripts/f5_retrieve_arms.py`. **Сырьё:** `f5/results_fresh.json`. +## Перемер на чистом индексе (2026-09-29, `results/f5/objective_clean.json`) + +После purge 772 мусорных файлов (20.3% чанков): **A hit@1 4/16 (0.25) · hit@3 5/16 (0.3125) · +hit@10 6/16 (0.375) · B top-1 4/16 (0.25)**. Против базы 26.09 (5/16 · 6/16 · 6/16 · 5/16): +hit@10 идентичен (те же 6 gold достижимы), топ-позиции −1 запрос (в пределах шума n=16). +**Вердикт: purge топ не двинул значимо** — мусор душил лексику, но top-1 устоял. +B-контексты стали короче (avg 8184 vs 19747 chars — другие top-1 доки). +Старая база ниже сохранена как протухшая, цитировать только новую. + ## Referent (каждое число) | Поле | Значение | diff --git a/experiments/4A_unit_of_return/f5/RESULTS_RELANG.md b/experiments/4A_unit_of_return/f5/RESULTS_RELANG.md new file mode 100644 index 00000000..284b95e7 --- /dev/null +++ b/experiments/4A_unit_of_return/f5/RESULTS_RELANG.md @@ -0,0 +1,38 @@ +# F5 relang — RU vs EN на чистом стеке (2026-09-29) + +**Вопрос:** влияет ли язык промпта на точность (B-плечо, whole-doc)? +**Дизайн:** 16 EN-переводов frozen-вопросов (`results/f5relang/queries_en.jsonl`, reference не тронуты), +B × 5 trials × 2 языка = 160+160 вызовов, один чистый индекс (после purge 772 мусорных файлов), +один стек, один день. Флаг харнесса `--queries-file` (варианты вне `frozen/`). + +## Результат + +| язык | correct | rate | Wilson95 | +|---|---|---|---| +| RU | 26/80 | 0.325 | 0.232–0.434 | +| EN | 30/80 | 0.375 | 0.277–0.485 | + +**Вердикт: эффекта языка нет** (CI пересекаются почти полностью). +Старый EN-замер 38.75% на сломанном стеке (BM25-only, мёртвый эмбеддер) — невалиден, не цитировать. + +## Нюанс (ключевая находка) + +6/16 запросов флипаются между языками all-or-nothing +(F5S-02: RU 5/5 vs EN 0/5; F5S-03/04/08 наоборот; F5S-13/15 частично). +Язык меняет *какие* запросы проходят (ретрив на другом языке достаёт другие доки), +а не *сколько*. Бимодальность подтверждает тезис Reader Capacity: ретрив решает, +читатель подчиняется. + +## Побочно + +- RU на чистом индексе (32.5%) ≈ RU на мусорном (34.4%, trials=10): B-плечо к мусору + почти иммунно. Выигрыш от purge ждём в objective hit@k — **objective-база протухла, + нужен перемер `f5_retrieve_arms`**. +- Сырьё: `results/f5relang/{ru,en}/judged_raw.json`, `aggregate.json`. + +## Оговорки + +- Run-to-run дисперсия жива (пилот 18/20 vs 8/20 на тех же запросах ранее) — + точечные сравнения ±1 запрос не интерпретировать, только агрегаты с CI. +- Замеры по новым правилам: свежий процесс, warm-up, void-флаг (`reranker_ms=0` → + замер недействителен), один event loop на серию. diff --git a/experiments/4A_unit_of_return/results/f5/objective_clean.json b/experiments/4A_unit_of_return/results/f5/objective_clean.json new file mode 100644 index 00000000..0fcfbc11 --- /dev/null +++ b/experiments/4A_unit_of_return/results/f5/objective_clean.json @@ -0,0 +1,614 @@ +{ + "queries": 16, + "mode": "quality", + "limit": 10, + "populations": [ + "code", + "prose" + ], + "metrics": { + "A_hit@1": { + "k": 4, + "n": 16, + "rate": 0.25, + "ci95": [ + 0.1018, + 0.495 + ] + }, + "A_hit@3": { + "k": 5, + "n": 16, + "rate": 0.3125, + "ci95": [ + 0.1416, + 0.556 + ] + }, + "A_hit@10": { + "k": 6, + "n": 16, + "rate": 0.375, + "ci95": [ + 0.1848, + 0.6136 + ] + }, + "B_top1_is_gold": { + "k": 4, + "n": 16, + "rate": 0.25, + "ci95": [ + 0.1018, + 0.495 + ] + }, + "C_oracle": { + "k": 16, + "n": 16, + "rate": 1.0, + "ci95": [ + 0.8064, + 1.0 + ] + }, + "D_closed_book": { + "k": 0, + "n": 16, + "rate": 0.0, + "ci95": [ + 0.0, + 0.1936 + ] + } + }, + "avg_context_chars": { + "A": 5425, + "B": 8184, + "C": 7337, + "D": 0 + }, + "records": [ + { + "id": "F5S-01", + "population": "code", + "gold_file": "src/core/rate_limiter.py", + "returned_files": [ + "experiments/4A_unit_of_return/frozen/f5/queries.md" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 779 + }, + "B_whole_doc": { + "top1_file": "experiments/4A_unit_of_return/frozen/f5/queries.md", + "hit": false, + "context_chars": 1568 + }, + "C_oracle": { + "hit": true, + "context_chars": 14768 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-02", + "population": "code", + "gold_file": "src/core/redact.py", + "returned_files": [ + "src/core/redact.py", + "src/core/redact.py", + "scripts/redact.py", + "src/core/redact.py", + "scripts/redact.py", + "scripts/redact.py", + "scripts/console_flash_monitor.py", + "src/mcp/context.py", + "experiments/canary_shadow/exp_canary_attack.py", + "scripts/e2e_quality_search.py" + ], + "A_chunks": { + "hit@1": true, + "hit@3": true, + "hit@10": true, + "context_chars": 6121 + }, + "B_whole_doc": { + "top1_file": "src/core/redact.py", + "hit": true, + "context_chars": 4570 + }, + "C_oracle": { + "hit": true, + "context_chars": 4570 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-03", + "population": "code", + "gold_file": "src/core/gitignore_parser.py", + "returned_files": [ + "experiments/4A_unit_of_return/frozen/f5/queries.md" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 779 + }, + "B_whole_doc": { + "top1_file": "experiments/4A_unit_of_return/frozen/f5/queries.md", + "hit": false, + "context_chars": 1568 + }, + "C_oracle": { + "hit": true, + "context_chars": 4990 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-04", + "population": "code", + "gold_file": "src/core/embedder_lease.py", + "returned_files": [ + "experiments/4A_unit_of_return/frozen/f5/queries.md", + "src/core/quiet_break_gate.py", + "src/core/search/fts5_index.py", + "src/providers/embedder/remote_embedder.py", + "scripts/console_flash_monitor.py", + "src/core/embedder_lease.py", + "src/providers/embedder/remote_embedder.py", + "src/core/doc_llm_verifier.py", + "src/core/indexing/indexer.py", + "src/providers/embedder/remote_embedder.py" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": true, + "context_chars": 5285 + }, + "B_whole_doc": { + "top1_file": "experiments/4A_unit_of_return/frozen/f5/queries.md", + "hit": false, + "context_chars": 1568 + }, + "C_oracle": { + "hit": true, + "context_chars": 2232 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-05", + "population": "code", + "gold_file": "src/core/reindex_ledger.py", + "returned_files": [ + "src/core/reindex_ledger.py" + ], + "A_chunks": { + "hit@1": true, + "hit@3": true, + "hit@10": true, + "context_chars": 842 + }, + "B_whole_doc": { + "top1_file": "src/core/reindex_ledger.py", + "hit": true, + "context_chars": 2552 + }, + "C_oracle": { + "hit": true, + "context_chars": 2552 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-06", + "population": "code", + "gold_file": "src/core/quiet_break_gate.py", + "returned_files": [ + "src/core/quiet_break_gate.py", + "src/core/quiet_break_gate.py", + "src/core/graph.py", + "src/core/quiet_break_gate.py", + "experiments/universal-engine/e05_action_receipt.py", + "src/core/quiet_break_gate.py", + "scripts/revision_gate.py", + "experiments/arclux_cycles_inventory.py", + "scripts/revision_gate.py", + "src/core/graph.py" + ], + "A_chunks": { + "hit@1": true, + "hit@3": true, + "hit@10": true, + "context_chars": 7357 + }, + "B_whole_doc": { + "top1_file": "src/core/quiet_break_gate.py", + "hit": true, + "context_chars": 11770 + }, + "C_oracle": { + "hit": true, + "context_chars": 11770 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-07", + "population": "code", + "gold_file": "src/providers/reranker/reranker_scoring.py", + "returned_files": [ + "src/providers/reranker/reranker_scoring.py", + "experiments/evalmut/probe_evalmut_transfer.py", + "src/providers/reranker/reranker_scoring.py", + "experiments/token_reduction/run_experiment.py", + "scripts/smoke_memory.py", + "scripts/revision_gate.py", + "src/core/intelligence/verify_on_read.py", + "scripts/diag_quality_hang.py", + "experiments/misc_probes/sandbox_lancedb_rmtree.py", + "src/core/error_handler.py" + ], + "A_chunks": { + "hit@1": true, + "hit@3": true, + "hit@10": true, + "context_chars": 7010 + }, + "B_whole_doc": { + "top1_file": "src/providers/reranker/reranker_scoring.py", + "hit": true, + "context_chars": 7358 + }, + "C_oracle": { + "hit": true, + "context_chars": 7358 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-08", + "population": "code", + "gold_file": "src/core/error_envelope.py", + "returned_files": [ + "scripts/check_lsp_health.py", + "src/mcp/tools/meta_tools.py", + "src/mcp/server_factory.py", + "src/mcp/tools/meta_tools.py", + "src/core/bootstrap_entities.py", + "src/mcp/server_factory.py", + "src/core/error_handler.py", + "src/mcp/tools/meta_tools.py", + "src/core/version_manager.py", + "src/mcp/tools/codebase_tool.py" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 5016 + }, + "B_whole_doc": { + "top1_file": "scripts/check_lsp_health.py", + "hit": false, + "context_chars": 9766 + }, + "C_oracle": { + "hit": true, + "context_chars": 4271 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-09", + "population": "prose", + "gold_file": "docs/en/GRACEFUL_DEGRADATION.md", + "returned_files": [ + "src/core/indexing/startup_diagnostics.py", + "src/core/autonomous_fix.py", + "src/core/search/graph_adapter.py", + "scripts/console_flash_monitor.py", + "src/core/search/agentic_search.py", + "experiments/1V_memory_contamination/burst_rename_audit.py", + "src/core/indexing/parser.py", + "docs/research/lsp-archive/lsp_main.py", + "src/core/log_manager.py", + "scripts/console_flash_monitor.py" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 6308 + }, + "B_whole_doc": { + "top1_file": "src/core/indexing/startup_diagnostics.py", + "hit": false, + "context_chars": 10028 + }, + "C_oracle": { + "hit": true, + "context_chars": 7394 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-10", + "population": "prose", + "gold_file": "docs/en/SEARCH_PIPELINE.md", + "returned_files": [ + "experiments/2E_evidence_ladder/graph_context_builder.py", + "src/core/bootstrap_entities.py", + "src/core/redact.py", + "experiments/embed_real_path_vs_raw/exp_feed_queue.py", + "src/mcp/server_tools.py", + "docs/ru/SEARCH_PIPELINE.md", + "experiments/context_engine/results_tasks_v3.json", + "docs/zh/SEARCH_PIPELINE.md", + "experiments/context_engine/results_tasks_v3.json", + "src/core/bootstrap_entities.py" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 9219 + }, + "B_whole_doc": { + "top1_file": "experiments/2E_evidence_ladder/graph_context_builder.py", + "hit": false, + "context_chars": 21343 + }, + "C_oracle": { + "hit": true, + "context_chars": 9413 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-11", + "population": "prose", + "gold_file": "docs/TRUST_BOUNDARY.md", + "returned_files": [ + "experiments/universal-engine/e05_action_receipt.py", + "src/core/quiet_break_gate.py", + "scripts/db_health.py", + "experiments/bootstrap/a2_test_nodes.py", + "experiments/embed_real_path_vs_raw/exp_feed_queue.py", + "src/core/search/branch_aware_index.py", + "src/core/artifact_paths.py", + "src/mcp/server_tools.py", + "src/core/graph.py", + "experiments/2E_evidence_ladder/temporal_facts_generator.py" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 6383 + }, + "B_whole_doc": { + "top1_file": "experiments/universal-engine/e05_action_receipt.py", + "hit": false, + "context_chars": 9363 + }, + "C_oracle": { + "hit": true, + "context_chars": 2531 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-12", + "population": "prose", + "gold_file": "docs/en/TELEMETRY.md", + "returned_files": [ + "scripts/summarize_1L_categories.py", + "src/providers/reranker/reranker_scoring.py", + "scripts/check_lsp_health.py", + "src/plugins/loader.py", + "experiments/bootstrap/e17_smoke_answers.json", + "experiments/stale_after/results.json", + "src/main.py", + "experiments/bootstrap/e17_judge/mapping.json", + "experiments/bootstrap/e17_shards/shard_1.answers.json", + "src/config/settings.py" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 8409 + }, + "B_whole_doc": { + "top1_file": "scripts/summarize_1L_categories.py", + "hit": false, + "context_chars": 14043 + }, + "C_oracle": { + "hit": true, + "context_chars": 11160 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-13", + "population": "prose", + "gold_file": "docs/adr/0003-verify-on-read.md", + "returned_files": [ + "src/core/action_receipt.py", + "src/core/intelligence/layer.py", + "src/core/intelligence/verify_on_read.py", + "experiments/1V_memory_contamination/memory_contamination_results_v3_retraction.json", + "src/providers/embedder/remote_embedder.py", + "experiments/stale_after/results.json", + "src/core/indexing/parser.py", + "src/core/instruction_scan.py", + "src/mcp/server_tools.py", + "experiments/bootstrap/e17_smoke_answers.json" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 10034 + }, + "B_whole_doc": { + "top1_file": "src/core/action_receipt.py", + "hit": false, + "context_chars": 16358 + }, + "C_oracle": { + "hit": true, + "context_chars": 12852 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-14", + "population": "prose", + "gold_file": "docs/en/HANDFOFF.md", + "returned_files": [ + "tests/conftest.py", + "src/core/platform_utils.py", + "scripts/check_lsp_health.py", + "adapters/zed/zed_config.py", + "experiments/context_engine/results_tasks_v3.json", + "src/sources/local_fs/windows.py", + "src/core/platform_utils.py", + "src/core/project_resolution.py", + "src/core/platform_utils.py", + "docs/ru/ZED_WINDOWS_QUIRKS.md" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 5052 + }, + "B_whole_doc": { + "top1_file": "tests/conftest.py", + "hit": false, + "context_chars": 2551 + }, + "C_oracle": { + "hit": true, + "context_chars": 6479 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-15", + "population": "prose", + "gold_file": "docs/en/ZED_WINDOWS_QUIRKS.md", + "returned_files": [ + "docs/ru/ZED_WINDOWS_QUIRKS.md", + "docs/en/ZED_WINDOWS_QUIRKS.md" + ], + "A_chunks": { + "hit@1": false, + "hit@3": true, + "hit@10": true, + "context_chars": 786 + }, + "B_whole_doc": { + "top1_file": "docs/ru/ZED_WINDOWS_QUIRKS.md", + "hit": false, + "context_chars": 12526 + }, + "C_oracle": { + "hit": true, + "context_chars": 12446 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + }, + { + "id": "F5S-16", + "population": "prose", + "gold_file": "docs/en/SYSTEM_REQUIREMENTS.md", + "returned_files": [ + "experiments/bootstrap/e17_smoke_answers.json", + "src/core/embedder/onnx_client.py", + "experiments/stale_after/results.json", + "src/providers/reranker/llama_install.py", + "experiments/bootstrap/e17_shards/shard_1.answers.json", + "src/providers/reranker/llama_install.py", + "experiments/bootstrap/e17_judge/mapping.json", + "install.py", + "experiments/bootstrap/e17_shards/shard_12.answers.json", + "src/core/commit_memory.py" + ], + "A_chunks": { + "hit@1": false, + "hit@3": false, + "hit@10": false, + "context_chars": 7420 + }, + "B_whole_doc": { + "top1_file": "experiments/bootstrap/e17_smoke_answers.json", + "hit": false, + "context_chars": 4017 + }, + "C_oracle": { + "hit": true, + "context_chars": 2608 + }, + "D_closed_book": { + "hit": false, + "context_chars": 0 + } + } + ] +} \ No newline at end of file diff --git a/experiments/4A_unit_of_return/results/f5relang/aggregate.json b/experiments/4A_unit_of_return/results/f5relang/aggregate.json new file mode 100644 index 00000000..12394a96 --- /dev/null +++ b/experiments/4A_unit_of_return/results/f5relang/aggregate.json @@ -0,0 +1,58 @@ +{ + "ru": { + "correct": 26, + "n": 80, + "rate": 0.325, + "wilson95": [ + 0.232, + 0.434 + ], + "code_rate": 0.5, + "per_query": { + "F5S-01": 0, + "F5S-02": 5, + "F5S-03": 0, + "F5S-04": 0, + "F5S-05": 5, + "F5S-06": 5, + "F5S-07": 5, + "F5S-08": 0, + "F5S-09": 0, + "F5S-10": 0, + "F5S-11": 0, + "F5S-12": 0, + "F5S-13": 1, + "F5S-14": 0, + "F5S-15": 5, + "F5S-16": 0 + } + }, + "en": { + "correct": 30, + "n": 80, + "rate": 0.375, + "wilson95": [ + 0.277, + 0.485 + ], + "code_rate": 0.75, + "per_query": { + "F5S-01": 0, + "F5S-02": 0, + "F5S-03": 5, + "F5S-04": 5, + "F5S-05": 5, + "F5S-06": 5, + "F5S-07": 5, + "F5S-08": 5, + "F5S-09": 0, + "F5S-10": 0, + "F5S-11": 0, + "F5S-12": 0, + "F5S-13": 0, + "F5S-14": 0, + "F5S-15": 0, + "F5S-16": 0 + } + } +} \ No newline at end of file diff --git a/experiments/4A_unit_of_return/results/f5relang/en/judged_raw.json b/experiments/4A_unit_of_return/results/f5relang/en/judged_raw.json new file mode 100644 index 00000000..2b523e2c --- /dev/null +++ b/experiments/4A_unit_of_return/results/f5relang/en/judged_raw.json @@ -0,0 +1,676 @@ +{ + "config": { + "outdir": "experiments/4A_unit_of_return/results/f5relang/en", + "n": 16, + "ids": "", + "arms": "B", + "trials": 5, + "judge_repeats": 1, + "parallel": 4, + "reader_model": "opencode-go/longcat-2.0", + "judge_model": "opencode-go/qwen3.7-plus", + "timeout": 400, + "seed": 42, + "reader_instr_file": null, + "queries_file": "experiments/4A_unit_of_return/results/f5relang/queries_en.jsonl", + "dry_run": false + }, + "records": [ + { + "id": "F5S-01", + "population": "code", + "gold_file": "src/core/rate_limiter.py", + "reference": "Защита от race conditions: threading.Lock (НЕ asyncio.Lock) —", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-02", + "population": "code", + "gold_file": "src/core/redact.py", + "reference": "Ловит только ИЗВЕСТНЫЕ ФОРМЫ по структурному префиксу.", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-03", + "population": "code", + "gold_file": "src/core/gitignore_parser.py", + "reference": "# Используем 'gitignore' вместо 'gitwildmatch' (deprecated)", + "arms": { + "B": { + "answers": [ + "Because 'gitwildmatch' is deprecated. The code explicitly uses the 'gitignore' format instead, as noted in the docstring and inline comment.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The parser uses the `'gitignore'` format because `'gitwildmatch'` is deprecated. Switching to `'gitignore'` eliminates 1740 GitWildMatchPattern deprecation warnings in tests.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Because `'gitwildmatch'` is deprecated — switching to `'gitignore'` eliminates 1740 `GitWildMatchPattern` deprecation warnings in tests and speeds up .gitignore checks during indexing.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The parser uses the 'gitignore' format because 'gitwildmatch' is deprecated — this avoids the deprecation warnings and is the non-deprecated alternative.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The `'gitwildmatch'` format is deprecated. Using the `'gitignore'` format eliminates 1740 GitWildMatchPattern deprecation warnings and avoids re-parsing `.gitignore` on every file.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "{\"verdict\": \"correct\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-04", + "population": "code", + "gold_file": "src/core/embedder_lease.py", + "reference": "Fix: an indexing process creates a short-lived *lease* file refreshed on every", + "arms": { + "B": { + "answers": [ + "The problem is solved with a cross-process **lease file**: an indexing process creates/short-lived lease file (`embedder.lease`) refreshed on every progress tick, and the idle watchdog checks for a fresh lease before killing — if any recent lease exists (own or another process's reindex), it must NOT kill the embedder.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The indexing process creates a short-lived lease file refreshed on every progress tick, and the idle watchdog is prohibited from killing the embedder while any fresh lease exists.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "A short-lived *lease* file is created on every progress tick by an indexing process; the idle watchdog must not kill the embedder while any fresh lease exists (its own reindex or another process's).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "A short-lived lease file is created by the indexing process and refreshed on every progress tick. The idle watchdog checks for this lease and must NOT kill the embedder while any fresh lease exists (from its own reindex or another process's).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The fix uses a short-lived **lease file** that an indexing process refreshes on every progress tick. The idle watchdog is required to check this lease and must **NOT kill the embedder while any fresh lease exists** (its own reindex or another process's).\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-05", + "population": "code", + "gold_file": "src/core/reindex_ledger.py", + "reference": "- MUST NOT raise, ever (a recording channel cannot break indexing).", + "arms": { + "B": { + "answers": [ + "The main constraint is that the module MUST NOT raise, ever — a recording channel cannot break indexing. It also must be append-only and use a single artifact path convention, but the non-negotiable rule is swallowing every error by design.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The module **MUST NOT raise, ever** — it swallows all errors by design, because a recording channel cannot be allowed to break indexing.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The module MUST NOT raise any exception — it is designed to swallow every error so that a recording channel can never break indexing.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The main constraint is that the module **MUST NOT raise, ever** — as stated in the design constraints, \"a recording channel cannot break indexing.\"\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The module **MUST NOT raise, ever** — it swallows every error by design so that the recording channel cannot break indexing.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-06", + "population": "code", + "gold_file": "src/core/quiet_break_gate.py", + "reference": "(fail-open, причина логируется): инфраструктурный сбой не должен блокировать коммит.", + "arms": { + "B": { + "answers": [ + "On infrastructural git or graph failure, the gate returns status `\"unavailable\"` (or `\"empty\"`) with an empty findings list and **does not trigger** (fail-open): the failure is logged but does not block the commit, because an infrastructure outage should not halt development.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "On infrastructural git or graph failure, the gate returns status `\"unavailable\"` with empty findings and does **not** trigger (fail-open) — the reason is logged, but the commit is not blocked.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "On infrastructural git or graph failure, the gate returns status `\"unavailable\"` with empty findings and does **not** block the commit (fail-open). It never throws an exception outward, and the reason is logged.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "On infrastructural git or graph failure, the gate returns status `\"unavailable\"` with empty findings and does **not** trigger (fail-open), logging the reason. This ensures an infrastructural failure does not block the commit.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "On infrastructural failure (missing git, absent PropertyGraph, or unreachable DB), the gate returns status `\"unavailable\"` with empty `findings` and does **not** block the commit — it fails open, logging the reason. This ensures an infrastructure outage never halts development.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-07", + "population": "code", + "gold_file": "src/providers/reranker/reranker_scoring.py", + "reference": "score: конечное число. NaN/Inf отбрасываются", + "arms": { + "B": { + "answers": [ + "`validate_scores` discards items with NaN/Inf scores (line 70-71: `if not math.isfinite(score): continue`). The docstring explains this is intentional — NaN/Inf represent an absence of assessment, and previously min/max clamping would silently convert NaN to 1.0 (maximum score for an unscored chunk).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`validate_scores` discards (filters out) items with NaN or Inf scores — they are not included in the validated list. The code at line 70-71 checks `math.isfinite(score)` and skips non-finite values, preventing NaN from silently becoming 1.0 via clamp.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "validate_scores discards elements where the score is NaN or Inf — they fail the `math.isfinite()` check and are excluded from the validated list entirely, rather than being clamped to 1.0 or 0.0. This is fail-safe: a missing score is preferred over a silently rearranged one.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`validate_scores` discards items with NaN or Inf scores — they are filtered out via `if not math.isfinite(score): continue`. The comment explains that NaN/Inf would otherwise pass `isinstance` checks and silently become 1.0 after min/max clamping, giving unscored chunks a maximum score. This is fail-safe: a lost score is better than a silently reordered one.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`validate_scores` discards items with NaN or Inf scores — they are filtered out via `math.isfinite(score)` check (line 70-71), not clamped. The comment explains NaN/Inf would otherwise pass `isinstance` checks and silently become 1.0/0.0.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"correct\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-08", + "population": "code", + "gold_file": "src/core/error_envelope.py", + "reference": "``next_action`` hint so an agent can branch and recover without guessing", + "arms": { + "B": { + "answers": [ + "The `next_action` field gives a concrete, single-step hint (\"what to do about it\") so an agent can branch and recover from a tool failure without guessing — agents branch on `code`, while humans read `next_action`.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The `next_action` field exists so an agent can branch and recover from a tool failure without guessing — each hint is concrete and single-step. Agents branch on the error code, while humans read `next_action` to understand what to do next.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The `next_action` field provides a concrete, single-step hint that tells an agent how to branch and recover after a tool failure—so it doesn't have to guess what to do next.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The `next_action` field gives an agent a concrete, single-step hint so it can branch and recover from a tool failure without guessing what to do next. Each error code maps to a specific playbook instruction (e.g., fix arguments, trigger a reindex, or back off), and agents branch on `code` while humans read `next_action`.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The `next_action` field exists so an agent can branch and recover from tool failures without guessing. Agents branch on the stable `code`, while humans and agents read the concrete, single-step `next_action` hint to know exactly what to do next.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-09", + "population": "prose", + "gold_file": "docs/en/GRACEFUL_DEGRADATION.md", + "reference": "it **degrades gracefully** through 6 levels,", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n→ Skill \"agentic-eval-lab\"\r\n", + "I don't know. The context only describes one degradation level (falling back to temporary storage when `safe_mkdir` fails with `ArtifactStorageError`), but it says nothing about external providers.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"incorrect\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-10", + "population": "prose", + "gold_file": "docs/en/SEARCH_PIPELINE.md", + "reference": "Ranks are computed **separately** for each channel, starting from 1.", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-11", + "population": "prose", + "gold_file": "docs/TRUST_BOUNDARY.md", + "reference": "**Принцип: содержимое репозитория — это ДАННЫЕ, а не код.**", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-12", + "population": "prose", + "gold_file": "docs/en/TELEMETRY.md", + "reference": "Metrics persist across MCP server restarts — `load_metrics()` reads the saved JSON on startup.", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"incorrect\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-13", + "population": "prose", + "gold_file": "docs/adr/0003-verify-on-read.md", + "reference": "вердикт **INCONCLUSIVE, никогда VERIFIED** (fail-closed:", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-14", + "population": "prose", + "gold_file": "docs/en/HANDFOFF.md", + "reference": "**Solution:** read Zed's SQLite database directly:", + "arms": { + "B": { + "answers": [ + "The MCP server detects the project via the `$ZED_WORKTREE_ROOT` environment variable, which Zed resolves at launch time to the open project's worktree root directory. The server entry configures `PROJECT_PATH = \"$ZED_WORKTREE_ROOT\"` in its `env` dict, so Zed substitutes the actual path when starting the process.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The MCP server uses the `$ZED_WORKTREE_ROOT` environment variable (set as `PROJECT_PATH` at line 308), which Zed automatically injects when launching the context server — it resolves to the root of the currently open project/worktree, so no manual environment variables need to be passed.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The MCP server reads the `PROJECT_PATH` environment variable, which is set to `\"$ZED_WORKTREE_ROOT\"` — Zed automatically substitutes this token with the path to the currently open project's worktree root.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The MCP server detects the project via the `PROJECT_PATH` environment variable, which is set to `$ZED_WORKTREE_ROOT` (line 308 in `zed_config.py`). This is automatically injected by Zed itself, representing the root of the currently open worktree, so no manual environment variables need to be passed.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "The server entry sets `env[\"PROJECT_PATH\"] = \"$ZED_WORKTREE_ROOT\"` — Zed automatically injects this variable pointing to the open project's root, so the MCP server reads it at startup (see `_make_server_entry` at line 308).\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-15", + "population": "prose", + "gold_file": "docs/en/ZED_WINDOWS_QUIRKS.md", + "reference": "**Status:** ⚠️ Always `` on Windows (Zed bug #36019)", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n✱ Grep \"ZED_WORKTREE_ROOT\" in . · 5 matches\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\n \"verdict\": \"incorrect\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-16", + "population": "prose", + "gold_file": "docs/en/SYSTEM_REQUIREMENTS.md", + "reference": "Moving to llama.cpp GGUF reduced MCP memory by 88% and total by 60%.", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + } + ] +} \ No newline at end of file diff --git a/experiments/4A_unit_of_return/results/f5relang/queries_en.jsonl b/experiments/4A_unit_of_return/results/f5relang/queries_en.jsonl new file mode 100644 index 00000000..c316d3fb --- /dev/null +++ b/experiments/4A_unit_of_return/results/f5relang/queries_en.jsonl @@ -0,0 +1,16 @@ +{"evidence_span": "Защита от race conditions: threading.Lock (НЕ asyncio.Lock) —", "gold_file": "src/core/rate_limiter.py", "id": "F5S-01", "population": "code", "question": "Why does the rate limiter use threading.Lock instead of asyncio.Lock, and which incident is this related to?", "question_ru": "Почему в ограничителе скорости используется threading.Lock, а не asyncio.Lock, и с каким инцидентом это связано?"} +{"evidence_span": "Ловит только ИЗВЕСТНЫЕ ФОРМЫ по структурному префиксу.", "gold_file": "src/core/redact.py", "id": "F5S-02", "population": "code", "question": "How reliably does the redact module strip secrets, and by what principle does it catch credential strings?", "question_ru": "Насколько надёжно модуль redact вырезает секреты и по какому принципу он ловит credential-строки?"} +{"evidence_span": "# Используем 'gitignore' вместо 'gitwildmatch' (deprecated)", "gold_file": "src/core/gitignore_parser.py", "id": "F5S-03", "population": "code", "question": "Why does the .gitignore parser use the 'gitignore' format instead of 'gitwildmatch'?", "question_ru": "Почему парсер .gitignore использует формат 'gitignore', а не 'gitwildmatch'?"} +{"evidence_span": "Fix: an indexing process creates a short-lived *lease* file refreshed on every", "gold_file": "src/core/embedder_lease.py", "id": "F5S-04", "population": "code", "question": "How is the problem solved that the idle-watchdog kills the shared embedder during a full reindex?", "question_ru": "Как решается проблема, что idle-watchdog убивает общий embedder во время полной переиндексации?"} +{"evidence_span": "- MUST NOT raise, ever (a recording channel cannot break indexing).", "gold_file": "src/core/reindex_ledger.py", "id": "F5S-05", "population": "code", "question": "What is the main constraint imposed on the durable reindex ledger module so that it cannot break indexing?", "question_ru": "Какое главное ограничение наложено на модуль durable reindex ledger, чтобы он не мог сломать индексацию?"} +{"evidence_span": "(fail-open, причина логируется): инфраструктурный сбой не должен блокировать коммит.", "gold_file": "src/core/quiet_break_gate.py", "id": "F5S-06", "population": "code", "question": "How does the quiet-break gate behave on an infrastructural git or graph failure?", "question_ru": "Как quiet-break gate ведёт себя при инфраструктурном сбое git или графа?"} +{"evidence_span": "score: конечное число. NaN/Inf отбрасываются", "gold_file": "src/providers/reranker/reranker_scoring.py", "id": "F5S-07", "population": "code", "question": "What does validate_scores do with unscored chunks whose score is NaN or Inf?", "question_ru": "Что делает validate_scores с неоценёнными чанками, у которых score равен NaN или Inf?"} +{"evidence_span": "``next_action`` hint so an agent can branch and recover without guessing", "gold_file": "src/core/error_envelope.py", "id": "F5S-08", "population": "code", "question": "Why does the MCP tool error envelope have a next_action field?", "question_ru": "Зачем в конверте ошибок MCP-инструментов есть поле next_action?"} +{"evidence_span": "it **degrades gracefully** through 6 levels,", "gold_file": "docs/en/GRACEFUL_DEGRADATION.md", "id": "F5S-09", "population": "prose", "question": "How many degradation levels is the system designed for, and what happens when external providers are unavailable?", "question_ru": "На сколько уровней деградации рассчитана система и что происходит, когда внешние провайдеры недоступны?"} +{"evidence_span": "Ranks are computed **separately** for each channel, starting from 1.", "gold_file": "docs/en/SEARCH_PIPELINE.md", "id": "F5S-10", "population": "prose", "question": "Why are ranks in the RRF fusion computed separately for each channel rather than with a shared enumerate?", "question_ru": "Почему в RRF-слиянии ранги считаются отдельно для каждого канала, а не общим enumerate?"} +{"evidence_span": "**Принцип: содержимое репозитория — это ДАННЫЕ, а не код.**", "gold_file": "docs/TRUST_BOUNDARY.md", "id": "F5S-11", "population": "prose", "question": "How does the system treat instructions from an indexed foreign repository — as commands or as data?", "question_ru": "Как система относится к инструкциям из проиндексированного чужого репозитория — как к командам или как к данным?"} +{"evidence_span": "Metrics persist across MCP server restarts — `load_metrics()` reads the saved JSON on startup.", "gold_file": "docs/en/TELEMETRY.md", "id": "F5S-12", "population": "prose", "question": "Are per-tool metrics preserved across MCP server restarts?", "question_ru": "Сохраняются ли метрики по инструментам между перезапусками MCP-сервера?"} +{"evidence_span": "вердикт **INCONCLUSIVE, никогда VERIFIED** (fail-closed:", "gold_file": "docs/adr/0003-verify-on-read.md", "id": "F5S-13", "population": "prose", "question": "What verdict does an ACTIVE memory node get if the resolver is unavailable and the check yielded no result?", "question_ru": "Какой вердикт получает ACTIVE-узел памяти, если резолвер недоступен и проверка не дала результата?"} +{"evidence_span": "**Solution:** read Zed's SQLite database directly:", "gold_file": "docs/en/HANDFOFF.md", "id": "F5S-14", "population": "prose", "question": "How does the MCP server detect the project open in Zed on Windows when no environment variables are passed?", "question_ru": "Каким способом MCP-сервер определяет открытый в Zed проект на Windows, если переменные окружения не передаются?"} +{"evidence_span": "**Status:** ⚠️ Always `` on Windows (Zed bug #36019)", "gold_file": "docs/en/ZED_WINDOWS_QUIRKS.md", "id": "F5S-15", "population": "prose", "question": "What is known about the ZED_WORKTREE_ROOT variable on Windows?", "question_ru": "Что известно о переменной ZED_WORKTREE_ROOT на Windows?"} +{"evidence_span": "Moving to llama.cpp GGUF reduced MCP memory by 88% and total by 60%.", "gold_file": "docs/en/SYSTEM_REQUIREMENTS.md", "id": "F5S-16", "population": "prose", "question": "By how much did the migration from ONNX to llama.cpp GGUF reduce the MCP process memory consumption?", "question_ru": "Насколько переход с ONNX на llama.cpp GGUF снизил потребление памяти MCP-процессом?"} diff --git a/experiments/4A_unit_of_return/results/f5relang/ru/judged_raw.json b/experiments/4A_unit_of_return/results/f5relang/ru/judged_raw.json new file mode 100644 index 00000000..fc588fb9 --- /dev/null +++ b/experiments/4A_unit_of_return/results/f5relang/ru/judged_raw.json @@ -0,0 +1,676 @@ +{ + "config": { + "outdir": "experiments/4A_unit_of_return/results/f5relang/ru", + "n": 16, + "ids": "", + "arms": "B", + "trials": 5, + "judge_repeats": 1, + "parallel": 4, + "reader_model": "opencode-go/longcat-2.0", + "judge_model": "opencode-go/qwen3.7-plus", + "timeout": 400, + "seed": 42, + "reader_instr_file": null, + "queries_file": null, + "dry_run": false + }, + "records": [ + { + "id": "F5S-01", + "population": "code", + "gold_file": "src/core/rate_limiter.py", + "reference": "Защита от race conditions: threading.Lock (НЕ asyncio.Lock) —", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-02", + "population": "code", + "gold_file": "src/core/redact.py", + "reference": "Ловит только ИЗВЕСТНЫЕ ФОРМЫ по структурному префиксу.", + "arms": { + "B": { + "answers": [ + "Модуль не является надёжной границей безопасности — он ловит только известные формы по структурному префиксу, а новый формат, секрет разбитый по строкам или записанный словами, проходит. Принцип — prefix-anchored (якоря на префикс, а не на энтропию/слово «key»), чтобы не резать легитимные hex-дайджесты, хеши коммитов и пути. Снижает цену случайности, но не делает канал безопасным для секретов.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Модуль ловит только известные формы по структурному префиксу (prefix-anchored, не по энтропии), снижая цену случайности, но не являясь границей безопасности — новый формат, секрет по строкам или записанный словами, пройдёт.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Модуль ненадёжен как граница безопасности — он ловит только известные формы по структурному префиксу (prefix-anchored), а новый формат или секрет, записанный иначе, пройдёт. Принцип: якорение на префикс (AKIA, sk-ant-, `Bearer` и т.п.), а не на энтропию или слово «key», чтобы не резать легитимные hex-дайджесты, хеши и пути. Для присваивания (`KEY=value`) — режет только если значение похоже на кред (длинное, не путь/число/плейсхолдер).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Модуль ловит credential-строки **только по структурному префиксу** (anchor на известные префиксы вроде `AKIA`, `sk-ant-`, `-----BEGIN...PRIVATE KEY-----`), а не по энтропии или слову «key» — чтобы не резать легитимные hex-дайджесты и пути. **Это не граница безопасности:** новый формат, секрет, разбитый по строкам или записанный словами — проходит. Модуль снижает цену случайности, но не делает канал доставки безопасным для секретов (строки 11–14).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Модуль ловит секреты по **структурному префиксу** (prefix-anchored), а не по энтропии — он не является границей безопасности и вырезает только известные форматы; новый формат, секрет по строкам или записанный словами — пройдёт.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-03", + "population": "code", + "gold_file": "src/core/gitignore_parser.py", + "reference": "# Используем 'gitignore' вместо 'gitwildmatch' (deprecated)", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-04", + "population": "code", + "gold_file": "src/core/embedder_lease.py", + "reference": "Fix: an indexing process creates a short-lived *lease* file refreshed on every", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-05", + "population": "code", + "gold_file": "src/core/reindex_ledger.py", + "reference": "- MUST NOT raise, ever (a recording channel cannot break indexing).", + "arms": { + "B": { + "answers": [ + "Модуль **никогда не выбрасывает исключения** — все ошибки проглотатываются по дизайну (`except Exception: pass`), потому что канал записи не может ломать индексацию.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Модуль **никогда не должен выбрасывать исключения** — он поглощает все ошибки по дизайну (строка 12: \"MUST NOT raise, ever\"), потому что канал записи не может ломать индексацию. Функция `record` перехватывает любое исключение и ничего не пробрасывает наружу.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Модуль **никогда не должен выбрасывать исключения** — все ошибки проглатываются намеренно (swallowed by design), потому что канал записи не может ломать индексацию.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Модуль **никогда не должен выбрасывать исключение** — он поглощает все ошибки по дизайну (строка 12: \"MUST NOT raise, ever\"), потому что канал записи не может ломать индексацию.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Модуль **никогда не должен выбрасывать исключения** — он перехватывает и подавляет все ошибки по дизайну (строка 12: \"MUST NOT raise, ever\"; строка 65-66: `except Exception: pass`). Это гарантирует, что канал записи не сломает индексацию при любом сбое.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"correct\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-06", + "population": "code", + "gold_file": "src/core/quiet_break_gate.py", + "reference": "(fail-open, причина логируется): инфраструктурный сбой не должен блокировать коммит.", + "arms": { + "B": { + "answers": [ + "При сбое git или графа гейт возвращает статус `unavailable` (или `empty` при отсутствии изменений), findings остаются пустыми, и гейт **не срабатывает** по принципу fail-open — причина логируется, но инфраструктурный сбой не блокирует коммит.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "При инфраструктурном сбое (git недоступен, граф отсутствует или пуст) гейт ведёт себя **fail-open**: возвращает статус `\"unavailable\"`, пустой список `findings` и **не блокирует коммит** — причина логируется, поскольку инфраструктурный отказ не должен останавливать разработку.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "При инфраструктурном сбое (git недоступен, граф не найден или изменений нет) гейт возвращает статус `\"unavailable\"` или `\"empty\"` с пустым списком findings и **не срабатывает** — это fail-open поведение: инфраструктурный сбой не должен блокировать коммит, а причина логируется.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "Gate ведёт себя **fail-open**: при недоступности git или графа возвращает статус `\"unavailable\"` с пустым списком findings и **не блокирует коммит**. Причина сбоя логируется (строка 17-18, 232-241).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "При инфраструктурном сбое тихий гейт ведёт себя как **fail-open**: возвращает статус `\"unavailable\"` или `\"empty\"` с пустым `findings` и **не блокирует коммиты**. Причина сбоя логируется, но gate не срабатывает — принцип «сбой инфраструктуры не должен останавливать коммит».\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-07", + "population": "code", + "gold_file": "src/providers/reranker/reranker_scoring.py", + "reference": "score: конечное число. NaN/Inf отбрасываются", + "arms": { + "B": { + "answers": [ + "`validate_scores` отбрасывает элементы с `NaN`Inf-скорами, пропуская их через `continue`, — это трактуется как отсутствие оценки (а не «высокий скор»).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`validate_scores` отбрасывает элементы с NaN/Inf (`math.isfinite` → `continue`), так как это рассматривается как отсутствие оценки, а не низкий скор — иначе clamp молча превратил бы NaN в 1.0 и неоценённый чанк получил бы максимальный скор.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`validate_scores` отбрасывает элементы с NaN/Inf score (не добавляет их в validated-список), потому что это отсутствие оценки, а не максимальный скор.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`validate_scores` отбрасывает элементы с NaN/Inf (`if not math.isfinite(score): continue`) — такие чанки не получают скор вообще, а не clamp'ятся к 1.0 или 0.0. Это fail-safe: раньше NaN молча превращался в максимальный скор 1.0, что давало неоценённому чанку незаслуженное преимущество при сортировке.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`validate_scores` отбрасывает элементы с `NaN`/`Inf` (строка 70–71: `if not math.isfinite(score): continue`), не включая их в результат — это fail-safe, чтобы неоценённый чанк не получил максимальный скор 1.0 после clamp.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"correct\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-08", + "population": "code", + "gold_file": "src/core/error_envelope.py", + "reference": "``next_action`` hint so an agent can branch and recover without guessing", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\n\nКонтекст содержит только скрипт `check_lsp_health.py` для диагностики LSP-сервера — в нём нет информации об MCP-инструментах, конверте ошибок или поле `next_action`.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\n \"verdict\": \"incorrect\"\n}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-09", + "population": "prose", + "gold_file": "docs/en/GRACEFUL_DEGRADATION.md", + "reference": "it **degrades gracefully** through 6 levels,", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-10", + "population": "prose", + "gold_file": "docs/en/SEARCH_PIPELINE.md", + "reference": "Ranks are computed **separately** for each channel, starting from 1.", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-11", + "population": "prose", + "gold_file": "docs/TRUST_BOUNDARY.md", + "reference": "**Принцип: содержимое репозитория — это ДАННЫЕ, а не код.**", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-12", + "population": "prose", + "gold_file": "docs/en/TELEMETRY.md", + "reference": "Metrics persist across MCP server restarts — `load_metrics()` reads the saved JSON on startup.", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-13", + "population": "prose", + "gold_file": "docs/adr/0003-verify-on-read.md", + "reference": "вердикт **INCONCLUSIVE, никогда VERIFIED** (fail-closed:", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "**INCONCLUSIVE**. Согласно контексту, если шаг не удалось выполнить (нет окружения/таймаут/нет доступа), вердикт — **INCONCLUSIVE**, а не REFUTED — недостаточно наблюдений ≠ отрицательный результат (`src/core/action_receipt.py:12-13`).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "correct", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"correct\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-14", + "population": "prose", + "gold_file": "docs/en/HANDFOFF.md", + "reference": "**Solution:** read Zed's SQLite database directly:", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-15", + "population": "prose", + "gold_file": "docs/en/ZED_WINDOWS_QUIRKS.md", + "reference": "**Status:** ⚠️ Always `` on Windows (Zed bug #36019)", + "arms": { + "B": { + "answers": [ + "На Windows переменная `ZED_WORKTREE_ROOT` всегда не установлена (``) из-за известного бага Zed #36019, закрытого без исправления. Полагаться на неё в `settings.json` нельзя, так как она вернёт `None`, в отличие от Linux/macOS, где она работает корректно. Из-за этого MCP-сервер не может определить проект по этой переменной и использует другие механизмы (SQLite DB, `.env`).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "На Windows переменная `ZED_WORKTREE_ROOT` **всегда не установлена** (``) — это известный баг Zed #36019, закрытый без исправления. Любая попытка использовать `$ZED_WORKTREE_ROOT` в `current_dir` или `env` даст `None`. На Linux/macOS переменная устанавливается корректно, на Windows MSCodeBase компенсирует это цепочкой fallback'ов (SQLite DB, PROJECT_PATH из `.env`).\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "На Windows переменная `ZED_WORKTREE_ROOT` **всегда ``** — это известный баг Zed (issue #36019), закрытый без исправления. Из-за этого в `settings.json` нельзя использовать `$ZED_WORKTREE_ROOT` в `current_dir` или `env` — любая попытка обратиться к ней вернёт `None`. MSCodeBase обходит это через fallback-цепочку (SQLite DB → PROJECT_PATH → CWD), которая работает без этой переменной.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`ZED_WORKTREE_ROOT` на Windows **всегда ``** — это известный баг Zed #36019, закрытый без исправления. Из-за этого в `settings.json` нельзя использовать `$ZED_WORKTREE_ROOT` в `current_dir` или `env`. MSCodeBase компенсирует это через цепочку fallback'ов: SQLite DB → `PROJECT_PATH` из `.env` → CWD.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "`ZED_WORKTREE_ROOT` на Windows всегда `` — это известный баг Zed #36019, закрытый без исправления. Переменная окружения не устанавливается, поэтому полагаться на неё нельзя (`$ZED_WORKTREE_ROOT` в `settings.json` даст `None`). MSCodeBase компенсирует это через цепочку fallback'ов (SQLite DB → PROJECT_PATH → CWD).\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "correct", + "correct", + "correct", + "correct", + "correct" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"correct\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"correct\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + }, + { + "id": "F5S-16", + "population": "prose", + "gold_file": "docs/en/SYSTEM_REQUIREMENTS.md", + "reference": "Moving to llama.cpp GGUF reduced MCP memory by 88% and total by 60%.", + "arms": { + "B": { + "answers": [ + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n", + "I don't know.\r\n\n\r\n> build · longcat-2.0\r\n\r\n" + ], + "verdicts": [ + "incorrect", + "incorrect", + "incorrect", + "incorrect", + "incorrect" + ], + "judge_texts": [ + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "```json\n{\"verdict\": \"incorrect\"}\n```\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ], + [ + "{\"verdict\": \"incorrect\"}\r\n\n\r\n> build · qwen3.7-plus\r\n\r\n" + ] + ] + } + } + } + ] +} \ No newline at end of file diff --git a/scripts/f5_judged_run.py b/scripts/f5_judged_run.py index 9977648c..281e96bc 100644 --- a/scripts/f5_judged_run.py +++ b/scripts/f5_judged_run.py @@ -123,8 +123,9 @@ def _norm(p: str) -> str: return (p or "").replace("\\", "/").lstrip("./") -def _load_queries() -> list[dict]: - lines = FROZEN.read_text(encoding="utf-8").splitlines() +def _load_queries(path: Path | None = None) -> list[dict]: + src = path or FROZEN + lines = Path(src).read_text(encoding="utf-8").splitlines() return [json.loads(line) for line in lines if line.strip()] @@ -279,6 +280,9 @@ def main() -> int: ap.add_argument("--seed", type=int, default=42) ap.add_argument("--reader-instr-file", default=None, help="path to file with reader instruction verbatim (default: built-in READER_INSTR)") + ap.add_argument("--queries-file", default=None, + help="path to queries jsonl (default: frozen/f5/queries.jsonl). " + "Use for language/population variants — never write into frozen/.") ap.add_argument("--dry-run", action="store_true") args = ap.parse_args() @@ -288,7 +292,7 @@ def main() -> int: reader_instr = READER_INSTR arms = [a.strip().upper() for a in args.arms.split(",") if a.strip()] - queries = _load_queries() + queries = _load_queries(Path(args.queries_file) if args.queries_file else None) if args.ids: want = {i.strip() for i in args.ids.split(",") if i.strip()} queries = [q for q in queries if q["id"] in want] diff --git a/scripts/purge_experiment_outputs.py b/scripts/purge_experiment_outputs.py new file mode 100644 index 00000000..037abd72 --- /dev/null +++ b/scripts/purge_experiment_outputs.py @@ -0,0 +1,89 @@ +#!/usr/bin/env python3 +"""One-time purge: удалить выводы экспериментов из индекса (2026-09-28). + +Замер: 828 файлов / 3127 чанков (20.3%) под experiments/**/results|work. +Штатный prune_deleted_files отказывает (safety-guard >50%), т.к. файлы +НА диске — их newly-excludes FileGuard. Поэтому точечное удаление по +тому же delete-механизму, guard безопасности не трогаем. + +Usage: + python scripts/purge_experiment_outputs.py # dry-run + python scripts/purge_experiment_outputs.py --apply # удалить + compaction +""" +from __future__ import annotations + +import argparse +import sys +import traceback +from pathlib import Path + +if sys.stdout is not None: + try: + sys.stdout.reconfigure(encoding="utf-8") + except Exception: + pass + +ROOT = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(ROOT)) + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("--apply", action="store_true") + args = ap.parse_args() + + from src.core.artifact_paths import get_db_path + from src.core.di_container import create_service_collection + from src.core.indexing.file_guard import FileGuard + from src.core.indexing.indexer import Indexer + from src.core.indexing.parser import CodeParser + from src.core.indexing.symbol_index import SymbolIndex + from src.core.system_artifacts import SystemArtifacts + from src.providers.embedder.remote_embedder import RemoteEmbedder + + services = create_service_collection(ROOT) + embedder = services.resolve(RemoteEmbedder) + indexer = Indexer(db_path=get_db_path(ROOT), embedder=embedder, file_guard=FileGuard(ROOT), + project_path=ROOT, parser=CodeParser(), symbol_index=SymbolIndex()) + table = indexer.table + rows = table.search().limit(100000).to_list() + files = {r["file_path"] for r in rows if "file_path" in r} + garbage = sorted(f for f in files if SystemArtifacts.is_experiment_output(Path(f))) + gset = set(garbage) + n_chunks = sum(1 for r in rows if r.get("file_path") in gset) + print(f"files in db: {len(files)}, garbage files: {len(garbage)}, garbage chunks: {n_chunks}") + if not args.apply: + print("dry-run: nothing deleted (use --apply)") + return 0 + + tbl = indexer.indexer_table if hasattr(indexer, "indexer_table") else indexer + ok, fail = 0, 0 + for i, fp in enumerate(garbage, 1): + try: + if tbl.delete_file(fp): + ok += 1 + else: + fail += 1 + pg = getattr(getattr(tbl, "_symbol_index", None), "graph", None) + if pg: + pg.remove_file(str(fp).replace("\\", "/")) + except Exception as e: # noqa: BLE001 - one bad file must not stop purge + print(f" FAIL {fp}: {e}") + fail += 1 + if i % 200 == 0: + print(f" ...{i}/{len(garbage)}") + try: + table.compact_files() + print("compaction done") + except Exception as e: # noqa: BLE001 + print(f"compaction skipped: {e}") + print(f"purged files: {ok}, failed: {fail}") + return 0 if fail == 0 else 1 + + +if __name__ == "__main__": + try: + raise SystemExit(main()) + except Exception: + traceback.print_exc() + raise SystemExit(1) diff --git a/src/core/indexing/file_guard.py b/src/core/indexing/file_guard.py index 22044ea5..02cf3c53 100644 --- a/src/core/indexing/file_guard.py +++ b/src/core/indexing/file_guard.py @@ -138,6 +138,12 @@ def is_safe_to_index(self, file_path: Path) -> bool: logger.debug(f"[FILEGUARD SKIP] System directory: {file_path}") return False + # Выводы экспериментов (experiments/**/results|work) — в индекс не берём. + # Git-трекинг не трогаем: frozen/results обязаны жить в репо (§17). + if SystemArtifacts.is_experiment_output(file_path): + logger.debug(f"[FILEGUARD SKIP] Experiment output: {file_path}") + return False + # Проверка .gitignore (Требует POSIX путей) if self._gitignore_patterns: try: diff --git a/src/core/search/engine.py b/src/core/search/engine.py index 2323cbc7..0cccf132 100644 --- a/src/core/search/engine.py +++ b/src/core/search/engine.py @@ -962,10 +962,12 @@ async def hybrid_search_async( except Exception as e: logger.warning(f"Не удалось выполнить dense поиск: {e}") - # FTS5 (full-text) — параллельно, с защитой от таймаута. - # _fts5_search делает lazy build (to_pandas на весь индекс, ~0.5s на - # первом вызове). Чтобы не усугублять 15s-лимит search_code, оборачиваем - # в wait_for(2s): при превышении — degraded ([]), основной поиск жив. + # FTS5 (full-text) — build вне таймаута, поиск под wait_for(2s). + # Замер 2026-09-28: холодный build (to_pandas всего индекса) = 2.47s > + # 2.0s — первый поиск в свежем процессе МОЛЧА терял FTS-тир + # (flaky A/B: пилот 18/20 vs 8/20 на тех же запросах). Build идемпотентен + # (double-checked lock в _build_fts5_index), поиск — быстрый (~0.05s). + await asyncio.to_thread(self._build_fts5_index) try: fts5_raw = await asyncio.wait_for( self._fts5_search_async(query, limit=raw_limit * 2), diff --git a/src/core/system_artifacts.py b/src/core/system_artifacts.py index bfd97809..da6e0093 100644 --- a/src/core/system_artifacts.py +++ b/src/core/system_artifacts.py @@ -101,9 +101,20 @@ # ══════════════════════════════════════════════════════════════ -# Public API +# Layer 4: Experiment Output Guard — выводы экспериментов # ══════════════════════════════════════════════════════════════ +# Замер 2026-09-28: 3127/15426 чанков индекса (20.3%) — мусор из +# experiments/**/results|work (ctx-дампы по 329 чанков, judged_raw.json — +# 249). Душит лексику (кейс P2) и раздувает холодный FTS-билд (2.47s). +# Git-трекинг НЕ трогаем (§17: frozen/results обязаны жить в репо) — +# исключаем только из ИНДЕКСА. Исходники экспериментов (*.py) индексируются. +_EXPERIMENT_ROOT = "experiments" +_EXPERIMENT_OUTPUT_DIRS = frozenset({"results", "work"}) + + +# ─── Public API ───────────────────────────────────────────── + class SystemArtifacts: """Единый источник правды о системных файлах проекта. @@ -197,7 +208,24 @@ def is_feedback_risk(cls, path: Path) -> bool: name = path.name.lower() return name in _FEEDBACK_PATTERNS - # ─── Layer 4: Unified Check ───────────────────────────── + # ─── Layer 4: Experiment Output Guard ──────────────── + + @classmethod + def is_experiment_output(cls, path: Path) -> bool: + """Проверяет, является ли файл выводом эксперимента. + + experiments/**/results/** и experiments/**/work/** — сырьё прогонов + (ctx-дампы, judged_raw.json, work-файлы). В индекс не берём; + frozen/-входы и *.py-исходники — берём. + """ + parts = [p.lower() for p in Path(path).parts] + try: + i = parts.index(_EXPERIMENT_ROOT) + except ValueError: + return False + return any(p in _EXPERIMENT_OUTPUT_DIRS for p in parts[i + 1 :]) + + # ─── Layer 5: Unified Check ───────────────────────────── @classmethod def is_system_path(cls, path: Path) -> bool: diff --git a/src/providers/reranker/llama_install.py b/src/providers/reranker/llama_install.py index d39db3a3..55e6ddb3 100644 --- a/src/providers/reranker/llama_install.py +++ b/src/providers/reranker/llama_install.py @@ -307,8 +307,11 @@ def _get_ext_dir() -> Path: p = Path(sys.executable).resolve().parent.parent.parent if (p / "src" / "main.py").exists() and (p / "__mscodebase_ext__.marker").exists(): return p - # Режим разработки (исходники с маркером расширения в корне) - p = Path(__file__).resolve().parent.parent.parent + # Режим разработки (исходники с маркером расширения в корне). + # __file__ = /src/providers/reranker/llama_install.py → 4 уровня + # до корня (было 3 — указывало на src/, ветка была мёртвой, и модели + # резолвились в пустой data_root; найдено 2026-09-28 по отсутствию GGUF). + p = Path(__file__).resolve().parent.parent.parent.parent if (p / "src" / "main.py").exists() and (p / "__mscodebase_ext__.marker").exists(): return p # Установленный пакет (pip/uvx) или неизвестный контекст: единый data root diff --git a/tests/test_experiment_output_guard.py b/tests/test_experiment_output_guard.py new file mode 100644 index 00000000..69a21346 --- /dev/null +++ b/tests/test_experiment_output_guard.py @@ -0,0 +1,46 @@ +"""Guard: выводы экспериментов не индексируются (experiments/**/results|work). + +Замер 2026-09-28: 3127/15426 чанков (20.3%) — мусор (ctx-дампы, judged_raw). +Git-трекинг не трогаем (§17) — только индекс. +""" +from __future__ import annotations + +from pathlib import Path + +from src.core.indexing.file_guard import FileGuard +from src.core.system_artifacts import SystemArtifacts + + +def test_results_and_work_excluded(): + assert SystemArtifacts.is_experiment_output( + Path("experiments/4A_unit_of_return/results/f5judged/judged_raw.json") + ) + assert SystemArtifacts.is_experiment_output( + Path("experiments/4A_unit_of_return/results/f5judged/work/ctx_F5S-01_A.txt") + ) + assert SystemArtifacts.is_experiment_output( + Path("D:/Project/MSCodeBase/experiments/noderag/results/r.json") + ) + + +def test_sources_and_frozen_kept(): + assert not SystemArtifacts.is_experiment_output( + Path("experiments/4A_unit_of_return/run_experiment.py") + ) + assert not SystemArtifacts.is_experiment_output( + Path("experiments/4A_unit_of_return/frozen/f5/queries.jsonl") + ) + assert not SystemArtifacts.is_experiment_output(Path("src/core/search/engine.py")) + assert not SystemArtifacts.is_experiment_output(Path("docs/en/SEARCH_PIPELINE.md")) + + +def test_fileguard_skips_experiment_output(tmp_path): + guard = FileGuard(tmp_path) + out = tmp_path / "experiments" / "x" / "results" / "r.json" + out.parent.mkdir(parents=True) + out.write_text('{"a": 1}', encoding="utf-8") + assert guard.is_safe_to_index(out) is False + + src = tmp_path / "experiments" / "x" / "run.py" + src.write_text("x = 1\n", encoding="utf-8") + assert guard.is_safe_to_index(src) is True