Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
50 commits
Select commit Hold shift + click to select a range
167172a
chore(privacy): remove personal absolute paths from docs and plugin
Sep 26, 2026
5f158ec
docs(experiments): pre-register 4A unit-of-return with agent-confound…
Sep 26, 2026
4196b1b
docs(experiments): add 4A freeze template and F3 research deltas
Sep 26, 2026
7e43a56
docs(experiments): verify F3 numbers against primary sources
Sep 26, 2026
c404118
test(experiments): guard frozen inputs tracked + recover E7 handout
Sep 26, 2026
b300c4c
feat(experiments): operationalize G6 overlap gate and judge-confound …
Sep 26, 2026
39aa25b
docs(experiments): F4 regression on recovered E7 list (qualitative re…
Sep 26, 2026
c94e2c6
feat(experiments): pin reasoning variant; harden blind-runner (F4 rerun)
Sep 26, 2026
5d307de
docs(experiments): confirm #11 control and draft fresh F4b candidates
Sep 26, 2026
8d63e71
docs(experiments): finalize diverse F4b candidate list after semantic…
Sep 26, 2026
2fb583c
fix(experiments): gate counts only first contiguous numbered block
Sep 26, 2026
8d5e2cf
docs(experiments): record design-level Red Team as explicit pre-run a…
Sep 26, 2026
6ce70d1
docs(experiments): freeze current crystal catalogue snapshot (15 entr…
Sep 26, 2026
ecf3398
docs(experiments): F4b targets both indexes; add pre-freeze twin check
Sep 26, 2026
709ebb5
feat(experiments): freeze F4b handouts (symptom+arrival) and manifest
Sep 26, 2026
b3b55f2
fix(experiments): raise blind-run timeout to 900s; add F4b orchestrator
Sep 26, 2026
13955c6
fix(experiments): replace false F4b must-NONE control #14, add validator
Sep 26, 2026
9ea6869
feat(experiments): F4b held-out run results, aggregator, manifest
Sep 26, 2026
60fb3b4
docs(experiments): F4b F6 red team on the numbers
Sep 26, 2026
3094cd2
docs: log F4b held-out result and G6 paraphrase-twin gap
Sep 26, 2026
0cd493a
fix(experiments): make the G6 novelty gate catch paraphrase twins
Sep 26, 2026
02b0781
feat(experiments): F5 4-arm retrieval harness + in-sample pilot
Sep 26, 2026
4af6e82
feat(experiments): freeze fresh F5 query set and run objective arms
Sep 26, 2026
f662b98
feat(experiments): F5 judged-arm harness + measured compute blocker
Sep 26, 2026
60431e7
fix(experiments): reject silent opencode model fallback in F5 judge run
Sep 26, 2026
277077e
feat(experiments): F5 judged reader arm — full pilot result
Sep 26, 2026
9376b16
docs: record F5 judged-arm pilot result
Sep 26, 2026
0a1c212
feat(experiments): F5 judged full run (trials=10) + artifact manifest
Sep 26, 2026
49e21a2
docs: record F5 judged full run (trials=10) and data completeness
Sep 26, 2026
386c8e7
chore(experiments): F0c — strip owner username from tracked files
Sep 26, 2026
55f8e00
data(experiments): refresh 4A artifact manifest hashes
Sep 26, 2026
4dea746
fix(experiments): opaque filenames in F5 judge harness (defense-in-de…
Sep 26, 2026
06e7637
docs(experiments): full red-team coverage and global artifact inventory
Sep 26, 2026
46cbe9a
fix(experiments): real failure markers + guard the judge call
Sep 26, 2026
c85c4dc
feat(experiments): closure-walk on personal-path guard — 4699 leaks o…
Sep 26, 2026
ea71590
fix(repo): normalize 17 production path leaks after closure-walk
Sep 26, 2026
b153ab3
docs: record closure-walk fix — 17 production path leaks normalized
Sep 26, 2026
153cdc6
feat(experiments): NodeRAG vs chunked retrieval on long docs — REFUTE…
Sep 26, 2026
2b8cd83
feat(tests): planted-break gate — 6 controls (3 guards x pos/neg)
Sep 26, 2026
adee4fb
docs: record planted-break gate results
Sep 26, 2026
694ba99
docs: record planted-break gate experiment
Sep 26, 2026
9491e4c
feat(scripts): redact.py — scrub personal paths with planted-key test…
Sep 26, 2026
21db238
docs: record redact.py experiment
Sep 26, 2026
fd505e8
feat(memory): stale_after + discriminator for memory notes — 17/17 te…
Sep 26, 2026
3f745f2
docs: record stale_after + discriminator results
Sep 26, 2026
5e7fd2d
feat(experiments): token reduction v1 — pooling 30% saves 62% tokens,…
Sep 26, 2026
5ab7265
feat(experiments): token reduction v2 — full-repo protocol, pooling 5…
Sep 26, 2026
54f0c57
docs(4a): correct judged majority rule + invalid count; changelog F5/…
Sep 27, 2026
088ebdf
Merge remote-tracking branch 'origin/main' into experiment/4a-unit-of…
Sep 27, 2026
0cec4a6
style: fix 4 ruff errors failing CI lint
Sep 27, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
47 changes: 47 additions & 0 deletions AGENT_DIARY.md

Large diffs are not rendered by default.

11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,17 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased] - 2026-09-27

### Added
- **F5 4-arm judged run** (4A unit-of-return): 16 frozen queries (8 code + 8 prose), 10 trials/arm, reader `opencode-go/longcat-2.0`, judge `opencode-go/qwen3.7-plus` (blind, disjoint from reader).
- Overall correct: A top-k chunks 16.3% (26/160), B top-1 full-doc 34.4% (55/160), C oracle 97.5% (156/160), D closed-book 0%.
- Code split: B 50.0% (40/80) vs A 6.3% (5/80), non-overlapping CIs; unit of return affects the reader, not gold-file hit.
- Prose split: A 26.3% (21/80) vs B 18.8% (15/80), overlapping CIs — fragile, no claim.
- Majority (strict >50% per query-arm): B 5/16 (F5S-13/B 5/10 tie counted out); guard: 0 `invalid` in trials=10 run (1 in t5 pilot, F5S-03/B).
- NodeRAG refuted on the same bench: TF-IDF baseline 80% vs graph BFS 70% — graph adds cost without gain here.
- Caveats: n=16 pilot scale; index snapshot not hard-frozen; raw answers path-normalized before commit, verdicts unchanged.

## [3.5.0] - 2026-09-22

### Added
Expand Down
171 changes: 171 additions & 0 deletions EXPERIMENTS_LOG.md

Large diffs are not rendered by default.

412 changes: 160 additions & 252 deletions KNOWN_ISSUES.md

Large diffs are not rendered by default.

689 changes: 608 additions & 81 deletions docs/archive/KNOWN_ISSUES_2026_09.md

Large diffs are not rendered by default.

11 changes: 11 additions & 0 deletions docs/en/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,17 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased] - 2026-09-27

### Added
- **F5 4-arm judged run** (4A unit-of-return): 16 frozen queries (8 code + 8 prose), 10 trials/arm, reader `opencode-go/longcat-2.0`, judge `opencode-go/qwen3.7-plus` (blind, disjoint from reader).
- Overall correct: A top-k chunks 16.3% (26/160), B top-1 full-doc 34.4% (55/160), C oracle 97.5% (156/160), D closed-book 0%.
- Code split: B 50.0% (40/80) vs A 6.3% (5/80), non-overlapping CIs; unit of return affects the reader, not gold-file hit.
- Prose split: A 26.3% (21/80) vs B 18.8% (15/80), overlapping CIs — fragile, no claim.
- Majority (strict >50% per query-arm): B 5/16 (F5S-13/B 5/10 tie counted out); guard: 0 `invalid` in trials=10 run (1 in t5 pilot, F5S-03/B).
- NodeRAG refuted on the same bench: TF-IDF baseline 80% vs graph BFS 70% — graph adds cost without gain here.
- Caveats: n=16 pilot scale; index snapshot not hard-frozen; raw answers path-normalized before commit, verdicts unchanged.

## [3.5.0] - 2026-09-22

### Added
Expand Down
11 changes: 11 additions & 0 deletions docs/ru/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,17 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased] - 2026-09-27

### Added
- **F5 4-arm judged run** (4A unit-of-return): 16 frozen queries (8 code + 8 prose), 10 trials/arm, reader `opencode-go/longcat-2.0`, judge `opencode-go/qwen3.7-plus` (blind, disjoint from reader).
- Overall correct: A top-k chunks 16.3% (26/160), B top-1 full-doc 34.4% (55/160), C oracle 97.5% (156/160), D closed-book 0%.
- Code split: B 50.0% (40/80) vs A 6.3% (5/80), non-overlapping CIs; unit of return affects the reader, not gold-file hit.
- Prose split: A 26.3% (21/80) vs B 18.8% (15/80), overlapping CIs — fragile, no claim.
- Majority (strict >50% per query-arm): B 5/16 (F5S-13/B 5/10 tie counted out); guard: 0 `invalid` in trials=10 run (1 in t5 pilot, F5S-03/B).
- NodeRAG refuted on the same bench: TF-IDF baseline 80% vs graph BFS 70% — graph adds cost without gain here.
- Caveats: n=16 pilot scale; index snapshot not hard-frozen; raw answers path-normalized before commit, verdicts unchanged.

## [3.5.0] - 2026-09-22

### Added
Expand Down
11 changes: 11 additions & 0 deletions docs/zh/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,17 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased] - 2026-09-27

### Added
- **F5 4-arm judged run** (4A unit-of-return): 16 frozen queries (8 code + 8 prose), 10 trials/arm, reader `opencode-go/longcat-2.0`, judge `opencode-go/qwen3.7-plus` (blind, disjoint from reader).
- Overall correct: A top-k chunks 16.3% (26/160), B top-1 full-doc 34.4% (55/160), C oracle 97.5% (156/160), D closed-book 0%.
- Code split: B 50.0% (40/80) vs A 6.3% (5/80), non-overlapping CIs; unit of return affects the reader, not gold-file hit.
- Prose split: A 26.3% (21/80) vs B 18.8% (15/80), overlapping CIs — fragile, no claim.
- Majority (strict >50% per query-arm): B 5/16 (F5S-13/B 5/10 tie counted out); guard: 0 `invalid` in trials=10 run (1 in t5 pilot, F5S-03/B).
- NodeRAG refuted on the same bench: TF-IDF baseline 80% vs graph BFS 70% — graph adds cost without gain here.
- Caveats: n=16 pilot scale; index snapshot not hard-frozen; raw answers path-normalized before commit, verdicts unchanged.

## [3.5.0] - 2026-09-22

### Added
Expand Down
2 changes: 1 addition & 1 deletion experiments/1V_memory_contamination/burst_rename_audit.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@
sys.stdout.reconfigure(encoding="utf-8")

ROOT = Path(r"D:\Project\MSCodeBase")
MEM = Path(r"C:\Users\misha\AppData\Local\mscodebase\projects\bfe9644b\intelligence\project_memory.json")
MEM = Path(r"<user>AppData\Local\mscodebase\projects\bfe9644b\intelligence\project_memory.json")

_CREATE_NO_WINDOW = 0x08000000 if sys.platform == "win32" else 0

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@
"date": "2026-08-11",
"scan_ms_total": 91.4,
"isolation": {
"store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\e8d56bb9\\intelligence",
"real_project_store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"store_dir": "<user>AppData\\Local\\mscodebase\\projects\\e8d56bb9\\intelligence",
"real_project_store_dir": "<user>AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"isolated": true
}
},
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@
"date": "2026-08-11",
"scan_ms_total": 86.4,
"isolation": {
"store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\f3bbb8a9\\intelligence",
"real_project_store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"store_dir": "<user>AppData\\Local\\mscodebase\\projects\\f3bbb8a9\\intelligence",
"real_project_store_dir": "<user>AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"isolated": true
}
},
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@
"replication": false,
"control_group": "v3 (add-only) / 1-R (retraction)",
"isolation": {
"store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\ace00824\\intelligence",
"real": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"store_dir": "<user>AppData\\Local\\mscodebase\\projects\\ace00824\\intelligence",
"real": "<user>AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"isolated": true
},
"verify_stats": {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,8 @@
"control_group": "v3 (memory_contamination_facts_v3_generated.json, same 50 facts)",
"parity_with_v3_A_code_first_adoption": "OK",
"isolation": {
"store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\107b52a6\\intelligence",
"real_project_store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"store_dir": "<user>AppData\\Local\\mscodebase\\projects\\107b52a6\\intelligence",
"real_project_store_dir": "<user>AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"isolated": true
},
"validation": {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,8 @@
"date": "2026-08-11",
"control_group": "v3 (add-only) / 1-R (retraction)",
"isolation": {
"store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\5517d5cd\\intelligence",
"real": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"store_dir": "<user>AppData\\Local\\mscodebase\\projects\\5517d5cd\\intelligence",
"real": "<user>AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"isolated": true
},
"verify_stats": {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@
"replication": true,
"control_group": "1-V (facts v3)",
"isolation": {
"store_dir": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\ab604b45\\intelligence",
"real": "C:\\Users\\misha\\AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"store_dir": "<user>AppData\\Local\\mscodebase\\projects\\ab604b45\\intelligence",
"real": "<user>AppData\\Local\\mscodebase\\projects\\bfe9644b\\intelligence",
"isolated": true
},
"verify_stats": {
Expand Down
2 changes: 1 addition & 1 deletion experiments/1V_memory_contamination/redteam_burst.py
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@
sys.stdout.reconfigure(encoding="utf-8")

ROOT = Path(r"D:\Project\MSCodeBase")
MEM = Path(r"C:\Users\misha\AppData\Local\mscodebase\projects\bfe9644b\intelligence\project_memory.json")
MEM = Path(r"<user>AppData\Local\mscodebase\projects\bfe9644b\intelligence\project_memory.json")
_CREATE_NO_WINDOW = 0x08000000 if sys.platform == "win32" else 0


Expand Down
51 changes: 51 additions & 0 deletions experiments/4A_unit_of_return/BASELINE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# 4A — BASELINE (freeze template + system snapshot)

> **Правило:** НИ ОДИН прогон 4A не запускается без заполненного раздела **Freeze** ниже.
> «После» — только **дописывается** (`## Run N`), никогда не перезаписывается (Tom: удаление =
> тюнинг под уже увиденное). Каждое число в отчёте несёт referent (популяция / corpus / candidate
> set / n / judge / noise).
>
> **Frozen inputs живут ТОЛЬКО в репо:** `experiments/4A_unit_of_return/frozen/` (git-tracked).
> Хранить замороженный список в `%TEMP%`/`/tmp` **запрещено** — прецедент 2026-09-26: список E7/E11
> лежал в `%TEMP%/opencode/e11` и был удалён, verbatim-регрессия стала невозможной. Guard:
> `tests/test_frozen_inputs_tracked.py`.

## System baseline — 2026-09-26 (MSCodeBase)

| Поле | Значение |
|---|---|
| Branch | `chore/privacy-paths` (PR #49) → база для эксперимента — до мерджа ветвиться от неё |
| HEAD | `5f158ece` (после PR #49) |
| MCP RUN_ID | `45b0a471f48e` (PID 15624, state READY) |
| Index | **10 106 chunks / 716 files / 14 099 symbols** |
| Embedder / Reranker | llama.cpp 🟢 / BGE-M3 🟢 |
| Tests | `pytest tests/`: **1866 passed, 5 skipped, 0 failed** (313s) |
| Privacy guard | `tests/test_no_personal_paths.py` ✅ |

## Freeze (заполнить ДО прогона, hash-фиксация)

| Поле | Значение |
|---|---|
| Experiment | 4A unit-of-return |
| git HEAD | `<sha>` |
| Index stats до | `chunks=___ files=___ symbols=___` |
| queries file | `<path>` sha256=`<hash>` |
| populations file | `<path>` sha256=`<hash>` |
| retriever / top-k | `<заморожено>` |
| judge model + budget | `<модель> / <reasoning>` |
| n per arm | `___` |
| noise floor | измеренный повторным прогоном: `___` |
| **Prediction (Tom):** | проза: whole-doc ≤ closed book; код: нет. Если обе популяции движутся одинаково → двухпопуляционная история неверна на нашем индексе. |

## Run 1 — `<дата>` (append, не редактировать прошлое)

| Arm | reader gets | hit@1 | hit@3 | top-1-doc-is-gold | context tokens | referent |
|---|---|---|---|---|---|---|
| A | top-k chunks | | | | | |
| B | top-1 whole document | | | | | |
| C | oracle file (chunked) | | | | | |
| D | nothing (closed book) | | | | | |

- Controls passed: `__/__`
- Verdict: `<…>` (raw output referenced)
- Index stats after: `chunks=___ files=___ symbols=___` (must equal "до", иначе прогон невалиден)
Loading
Loading