Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions AGENT_DIARY.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,17 @@
- **Чёрные окна CMD (2026-08-14):** MCP запускался как `venv\Scripts\python.exe` (console-подсистема) → каждое окно Zed = своё чёрное окно; фикс: `pythonw.exe` в extension.toml + CREATE_NO_WINDOW во ВСЕХ runtime subprocess (13 файлов) — с pythonw (нет консоли) незакрытые git/wmic/netstat мигали бы окнами
- **FA=0.00 ≠ качество guardrail (2026-08-15):** Exp 1-L Day 3 — qwen3.6/3.7 (zero-shot VOR) достигают FA=0.00 ценой recall(real)=0.08–0.20 (code_first: 2/25 правды принято, 7/25 активно отвергнуто) — fail-closed политика, а не «фильтрация лжи»; выбор LLM для verify-on-read = выбор политики (fail-closed qwen vs max-coverage glm), recall(real) обязан быть в метриках. CoT (V3/Part 5) НЕ окупается: только qwen3.6 recall 0.08→0.20 при цене ×30–65

## [2026-09-25] Reindex deadlock + concurrent indexers + embedder lease — Fixed (live)
**Status:** ✅ Fixed + live-verified (full reindex completed; index collapse confirmed).
**Root Cause:** `_bounded_link` (timeout fix 2026-09-25) ran `bulk_write` on a NEW daemon thread, while `run()` held the global write RLock on the caller thread for the whole reindex. `bulk_write` acquires the SAME RLock → permanent deadlock (py-spy: "bounded-bulk_write" idle at `db_writer.py:336`; job stuck 52% "running", 0 CPU). Timeout only fired at 300s → job failed. Same class for prune/verify (recreate_table_physical).
**Fix:** `_bounded_link` now wraps the bounded call in `_suspend_write_lock()` (same remedy as `_safe_ivf_index`). One place, covers all links. Guard: `tests/test_bounded_link_deadlock.py` (deterministic, negative control fails-fast at bound).
**Superseded (same day):** the `_suspend_write_lock` remedy *released* the global RLock and exposed a 2nd bug — auto-index + manual trigger ran TWO indexers concurrently (py-spy: two threads in `index_project`). Proper fix: **split the locks** — `run()` holds `begin_run()` (separate non-reentrant Lock) for run exclusion; `_table_write_lock` is acquired per DB op; `_bounded_link` no longer suspends. Guards: `test_run_singleflight.py`, updated `test_bounded_link_deadlock.py`.
**Live result:** full reindex `c09c2e22` → **completed 858.5s**; index **19653 → 10103**; path-duplication **668 → 0**; dup(file_path,chunk_index) **144 → 0**.
**Also:** `src/core/reindex_ledger.py` — durable `<data_root>/logs/reindex_ledger.jsonl` recording start/phase/error+traceback/zombie/end; `layer.py` guarantees terminal status + watchdog (`MSCODEBASE_REINDEX_STALL_SEC=900`). First real row captured the deadlock traceback → no more guessing.
**Second incident (same session):** 22:09 a *devbase* MCP server spawned and started its own llama-server on fixed :8080/:8081; our MSCodeBase server died hard mid-parse (ledger: start, no end; driver got `ClosedResourceError`). Confirms fixed-port / multi-window conflict class (see KNOWN_ISSUES).
**Files:** `src/core/indexing/index_project_runner.py`, `src/core/reindex_ledger.py`, `src/core/intelligence/layer.py`, tests.
**verified_from_clean_state:** ⚠️ не проверено (локальный pytest: 8 passed; live reindex прерван крахом сервера).

## [2026-09-22] E17 v7 — CORRECTION: AST proxies do not predict kill-rate (v6 not replicated)

**Status:** Measured (v6 REFUTED on replication).
Expand Down
18 changes: 18 additions & 0 deletions EXPERIMENTS_LOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,23 @@
# EXPERIMENTS_LOG.md — Audit Verification (2026-07-22)

## [2026-09-25] — E18: graph write throughput — per-entity transaction vs batched (CONFIRMED)

**Гипотеза:** `PropertyGraph.add_node/add_edge` открывают ОДНУ SQLite-транзакцию + named-mutex на каждый узел/ребро (`graph.py:526-544, 816-851`) → graph-сборка сериализована и I/O-bound; это root cause «CPU 5% + диск ~3 МБ/с» на фазе parsing.
**Команда:** `<ext-venv-python> experiments/misc_probes/exp_graph_write_throughput.py` (4 руки, temp-БД, одинаковый INSERT SQL).
**Сырой вывод (после фикса, arm D = реальный `PropertyGraph.batch()`):**
```
N_NODES=2000 N_EDGES=4000
A per-call PropertyGraph (mutex+txn) : 20.58s 194 entities/s
B raw, ONE transaction+reused : 0.03s 140875 entities/s
C raw, per-node transaction : 16.89s 237 entities/s
D PropertyGraph.batch() [the fix] : 0.15s 26826 entities/s
speedup D/A (the fix) = 138.0x ; speedup B/A = 724.8x
negative control counts equal: True -> [(2000, 2000)]
```
**Вердикт:** CONFIRMED. Per-row commit/fsync — доминанта; named-mutex добавляет ~35%. Отрицательный контроль — counts идентичны (нет порчи). Внешний baseline совпадает: SQLite 429 rows/s отдельными транзакциями vs 2.457M rows/s (одна транзакция+reused) — voidstar.tech.
**Фикс внедрён (2026-09-25):** `PropertyGraph.batch()` (одна tx на файл) + `indexer._parse_file_only` оборачивает graph-update файла; измерено **138× на реальном API**. Guard: `tests/test_graph_batch.py` (counts/атомарность/вложенность/персистентность).
**Live-контроль (job 268ac11b, 2026-09-26):** первая попытка на `batch()` **упала на finalize** (`cannot start a transaction within a transaction` в `GraphSymbolResolver.resolve_all`): `batch()` хранил состояние глобально, а parse идёт в 4 потока — файлы делили одну транзакцию. Фикс: владелец батча по thread-id (`_batch_owner`); чужой поток не присоединяется, а блокируется и получает свой батч (проверка и в add_node/add_edge/delete_node). Guard: `test_concurrent_batches_do_not_share_a_transaction` (Barrier, 2 потока, без OperationalError, counts). Повторный прогон → **completed 603с, 10106/10106**, path-dup 0.

## [2026-09-20] — E13 / поискочное качество: 6 исследовательских задач (план)

**Статус:** Plan (код не тронут; задачи в ISSUE.md KI-R1..R6)
Expand Down
Loading
Loading