Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 46 additions & 22 deletions docs/blog/bootstrap-pipeline.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ description: "Part 4 of MSCodeBase Intelligence — Field Notes. Full source-mat
tags: search, rag, codearchitecture, testing
---

> **Disclaimer & Status:** draft (source-material for the article). This is not a "feature advertisement", but an honest engineering story: figures are reproducible, weak points are named, and unaddressed risks are listed in the "What Could Go Wrong" section.
> **Disclaimer & Status:** This is an ongoing research investigation, not a feature announcement. We explore what happens when execution-derived test-to-code traceability becomes repository evidence for AI coding agents. Figures are reproducible, limitations are named, and the central question — whether this evidence helps LLMs — remains open.

---

Expand Down Expand Up @@ -168,9 +168,9 @@ Linked % on clean (non-mocked) third-party projects proved **higher** than on ou

---

## And Now: Edges Met the Consumer (E17)
## And Now: Runtime Evidence Becomes Repository Evidence (E17)

Everything prior built `test ──TESTS──> function` edges inside PropertyGraph without leveraging them during search. E17 closed the loop:
Everything prior built `test ──TESTS──> function` edges inside PropertyGraph. The question: what happens when we expose this execution-backed evidence to the retrieval system?

- **Data:** 1,727 tests → **16,172 TESTS edges**, 1,595 Test nodes, 1,132 covered functions.
- **Implementation:** `SymbolIndexAdapter.get_tests_for_symbol()` (incoming `TESTS` edges) + `Searcher._append_tests_signal()`: appends up to 3 tests per function (capped at `min(len, 6)` per query), `graph_score = 0.4` vs 1.0 for definitions, sentinel `chunk_index = -(20_000_000 + line)` avoiding collisions with code chunks in RRF ranking.
Expand Down Expand Up @@ -199,7 +199,11 @@ TESTS-signal: 34/35 queries received new covering tests (97.1%)
graph_stage avg dt: off=6.52ms, on=7.53ms (overhead +15.3%)
```

**Critical finding:** TESTS-signal **does not improve hit@1** (off=on). It only **adds context** (tests) to already-found results: 97.1% of queries received new covering tests. This means TESTS-signal is **context for LLM**, not a search improvement. If LLM doesn't use tests, the signal is useless.
**Critical finding:** TESTS-signal **does not improve retrieval ranking** (hit@1 off=on at 94.3%, MRR unchanged at 0.957). It adds execution-backed evidence to the context (97.1% of queries received new covering tests), but this evidence does not change which function is found first.

This means the value of TESTS-signal, if any, lies **not in retrieval** but potentially in **LLM understanding**: does seeing the actual tests that exercise a function help the model understand behavior, identify edge cases, or propose safer changes?

This remains an open experiment.

### Language Coverage (Critical Limitation)

Expand All @@ -225,44 +229,36 @@ We tested TESTS-signal against 5 attack vectors:

## What Could Go Wrong

A transparent list of risks and open validation items:
### Fundamental question
The central risk is not technical but conceptual: **execution-derived test evidence may not help LLMs at all**. If models already infer behavior from code structure, naming, and docstrings, adding explicit test links may provide no additional signal. This can only be answered by measuring LLM performance with and without TESTS evidence on tasks like behavior understanding, edge case detection, and change planning.

### Technical limitations
1. **Evaluation scope.** While expanded to a 35-query panel, evaluation is still performed on a single primary codebase without deep reranker interaction.

2. **A/B did not improve hit@1.** Wide panel (35 queries):
2. **A/B did not improve retrieval.** Wide panel (35 queries):
- hit@1 off=33/35 on=33/35 (94.3%)
- hit@3 off=34/35 on=34/35 (97.1%)
- MRR(function) off=0.957 on=0.957

TESTS-signal **does not help find the function** (hit@1 did not improve). It only **adds context** (tests) to already-found results: 34/35 queries received new covering tests (97.1%).

→ **Conclusion:** TESTS-signal is **context for LLM**, not a search improvement.
→ **Risk:** if LLM doesn't use tests, the signal is useless.
→ **Fix:** verify on real LLM pipeline (not in this experiment).
TESTS-signal adds evidence but does not change retrieval ranking. The next experiment must measure LLM-level outcomes.

3. **Overhead +15.3% for wide panel.** Average graph_stage time:
- off: 6.52ms
- on: 7.53ms
- overhead: +15.3%

For bootstrap (one-time run) this is acceptable. For prod search — may be critical with many queries.

→ **Risk:** at 1000 queries/sec, overhead may be noticeable.
→ **Fix:** cache TESTS-signal (not done).

4. **Red team: 5/5 attacks repelled.** Tested:
- ✅ **Concurrency:** 10 threads × 100 calls = 1,000 calls in 17.2s, 0 errors
- ✅ **Boundaries:** function with 234 tests = 16.11ms (acceptable)
- ✅ **Abuse:** query for nonexistent function = 0 results (graceful degradation)
- ✅ **TOCTOU:** graph closed between calls = graceful degradation
- ✅ **Dependency failure:** PropertyGraph with nonexistent path = 0 results (graceful degradation)

→ **Conclusion:** TESTS-signal is resilient to concurrency, boundaries, abuse, TOCTOU, and dependency failures.
- ✅ Concurrency: 10 threads × 100 calls = 1,000 calls in 17.2s, 0 errors
- ✅ Boundaries: function with 234 tests = 16.11ms (acceptable)
- ✅ Abuse: query for nonexistent function = 0 results (graceful degradation)
- ✅ TOCTOU: graph closed between calls = graceful degradation
- ✅ Dependency failure: PropertyGraph with nonexistent path = 0 results (graceful degradation)

5. **Language limitation.** Dynamic trace is currently Python-only (~34% function coverage in Python, 0% in JS/TS/Go).

6. **`graph_score = 0.4` is an empirical constant.** Chosen to stay strictly below function definitions, but unverified against BM25/reranker weight interactions.
→ Verify on full pipeline; constant may become a parameter.

7. **Pointer `:0`.** Test nodes from dynamic trace lack line numbers (`line=0`). Indexers must resolve test decorator line positions before enabling in prod.

Expand All @@ -280,6 +276,34 @@ A transparent list of risks and open validation items:

---

## Related Work

Test-to-code traceability is not a new problem. TCTracer (White & Krinke, 2022) and PyTCTracer already use dynamic execution traces to establish test→code links. Chen et al. (2025) replicated this on Python projects and found that many classical techniques work worse on Python than on Java.

What differs in our approach is the **downstream application**:

```
Classical test-to-code traceability:
test → code (for maintenance, refactoring, impact analysis)

Our approach:
test → runtime execution → persistent graph → repository retrieval → LLM context
```

Recent work is moving in similar directions:
- **TDAD** (2026): graph-based impact analysis for coding agents, but uses static AST
- **RepoGraph**: runtime overlay on static graph, but static-first
- **TICoder** (2026): tests as behavioral context for LLM, but without runtime trace
- **Agent Retrieval Bench** (2026): benchmark with code2test/trace2code tasks

We have not found published work that explicitly builds the full chain from runtime execution to persistent graph to LLM context. This does not mean it doesn't exist — only that in our search we did not find it.

Our research question is therefore: **what happens when execution-derived test-to-code traceability becomes repository evidence for an AI coding agent?**

We do not yet know the answer. Current results show that TESTS-signal does not improve retrieval ranking (hit@1 unchanged at 94.3%). The next experiment is whether this evidence helps LLM understand behavior, identify edge cases, or propose safer changes.

---

## Acknowledgments

A huge thank you to everyone who engages with these posts in the comments. Your feedback, real-world observations, counter-examples, and benchmark numbers directly shape these experiments. This kind of open technical critique is what keeps engineering honest.
Expand Down
19 changes: 19 additions & 0 deletions experiments/bootstrap/build_gemma_tests.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# -*- coding: utf-8 -*-
"""Build TESTS edges for gemma_agent from trace_result.json."""
import sys
from pathlib import Path

sys.path.insert(0, r"D:\Project\MSCodeBase")
sys.stdout.reconfigure(encoding="utf-8")

from src.core.bootstrap_tests import build_from_trace_file

project_root = Path("D:/Project/gemma_agent")
trace_file = project_root / "trace_result.json"

print("=" * 80)
print("Building TESTS edges for gemma_agent")
print("=" * 80)

result = build_from_trace_file(trace_file, project_root)
print(f"\n✅ Result: {result}")
33 changes: 33 additions & 0 deletions experiments/bootstrap/build_gemma_tests_v2.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# -*- coding: utf-8 -*-
"""Build TESTS edges for gemma_agent with correct src_dir."""
import sys
from pathlib import Path

sys.path.insert(0, r"D:\Project\MSCodeBase")
sys.stdout.reconfigure(encoding="utf-8")

import json

from src.core.artifact_paths import get_graph_db_path
from src.core.bootstrap_tests import build_tests_edges
from src.core.graph import PropertyGraph

project_root = Path("D:/Project/gemma_agent")
trace_file = project_root / "trace_result.json"
graph_path = get_graph_db_path(project_root)

print("=" * 80)
print("Building TESTS edges for gemma_agent (with src_dir=core)")
print("=" * 80)

trace = json.loads(trace_file.read_text(encoding="utf-8"))
graph_db = PropertyGraph(graph_path)

# Try with src_dir=core
src_dir = project_root / "core"
result = build_tests_edges(trace, project_root, graph_db, src_dir=src_dir)

print(f"\nResult: {result.as_dict()}")
print(f"Missing (first 10): {result.missing[:10]}")

graph_db.close()
40 changes: 40 additions & 0 deletions experiments/bootstrap/check_graph_tests.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
# -*- coding: utf-8 -*-
import sys
from pathlib import Path

sys.path.insert(0, str(Path(__file__).resolve().parents[2]))

import sqlite3

from src.core.artifact_paths import get_graph_db_path

db_path = get_graph_db_path(Path('.'))
conn = sqlite3.connect(str(db_path))

# Проверим TESTS-рёбра
print('=== TESTS edges in graph ===')
result = conn.execute("SELECT COUNT(*) FROM edges WHERE type='TESTS'").fetchone()
print(f'Total TESTS edges: {result[0]}')

# Проверим safe_mkdir
print('\n=== safe_mkdir node ===')
result = conn.execute("SELECT id, name, file_path FROM nodes WHERE name LIKE '%safe_mkdir%' LIMIT 5").fetchall()
for r in result:
print(f' id={r[0]}, name={r[1]}, file={r[2]}')

# Проверим TESTS-рёбра для safe_mkdir
if result:
node_id = result[0][0]
print(f'\n=== TESTS edges TO safe_mkdir (id={node_id}) ===')
edges = conn.execute(f"SELECT COUNT(*) FROM edges WHERE type='TESTS' AND target_id={node_id}").fetchone()
print(f'Incoming TESTS edges: {edges[0]}')

# Покажем несколько тестов
tests = conn.execute(f"SELECT source_id FROM edges WHERE type='TESTS' AND target_id={node_id} LIMIT 5").fetchall()
print('\n=== Sample test nodes ===')
for t in tests:
test_node = conn.execute(f"SELECT name, file_path FROM nodes WHERE id={t[0]}").fetchone()
if test_node:
print(f' {test_node[0]} @ {test_node[1]}')

conn.close()
35 changes: 35 additions & 0 deletions experiments/bootstrap/debug_gemma_paths.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# -*- coding: utf-8 -*-
"""Debug: check how functions are stored in gemma_agent graph."""
import sys
from pathlib import Path

sys.path.insert(0, r"D:\Project\MSCodeBase")
sys.stdout.reconfigure(encoding="utf-8")

from src.core.artifact_paths import get_graph_db_path
from src.core.graph import NodeLabel, PropertyGraph

project_root = Path("D:/Project/gemma_agent")
graph_path = get_graph_db_path(project_root)
graph_db = PropertyGraph(graph_path)

# Check sample function nodes
print("=== Sample Function nodes in graph ===")
funcs = graph_db.find_nodes(label=NodeLabel.FUNCTION, limit=10)
for f in funcs:
print(f" name={f.name}, file_path={f.file_path}")

# Check trace paths
print("\n=== Sample trace entries ===")
import json

trace_file = project_root / "trace_result.json"
trace = json.loads(trace_file.read_text(encoding="utf-8"))
for i, (nodeid, entries) in enumerate(trace.items()):
if i >= 3:
break
print(f" {nodeid}")
for e in entries[:3]:
print(f" {e}")

graph_db.close()
60 changes: 60 additions & 0 deletions experiments/bootstrap/debug_get_tests.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# -*- coding: utf-8 -*-
"""Debug get_tests_for_symbol step by step."""
import sys
from pathlib import Path

sys.path.insert(0, str(Path(__file__).resolve().parents[2]))

from src.core.artifact_paths import get_graph_db_path
from src.core.graph import EdgeType, NodeLabel, PropertyGraph
from src.core.search.graph_adapter import SymbolIndexAdapter

ROOT = Path(__file__).resolve().parents[2]
pg = PropertyGraph(get_graph_db_path(ROOT))
adapter = SymbolIndexAdapter(pg, mode=SymbolIndexAdapter.MODE_PURE)

symbol = "safe_mkdir"
file_path = "src/core/artifact_paths.py"

print("=== Debug get_tests_for_symbol ===")
print(f"symbol: {symbol}")
print(f"file_path: {file_path}")

# Step 1: find_nodes
print(f"\n[Step 1] find_nodes(label=FUNCTION, name_pattern=%{symbol}%, file_path={file_path})")
candidates = pg.find_nodes(
label=NodeLabel.FUNCTION,
name_pattern=f"%{symbol}%",
file_path=file_path,
limit=5,
)
print(f" Found {len(candidates)} candidates")
for c in candidates:
print(f" {c.name} @ {c.file_path} (qname={c.qualified_name})")

# Step 2: Попробуем без file_path
print("\n[Step 2] find_nodes без file_path")
candidates2 = pg.find_nodes(
label=NodeLabel.FUNCTION,
name_pattern=f"%{symbol}%",
limit=5,
)
print(f" Found {len(candidates2)} candidates")
for c in candidates2:
print(f" {c.name} @ {c.file_path}")

# Step 3: Проверим get_neighbors для первого кандидата
if candidates2:
node = candidates2[0]
print(f"\n[Step 3] get_neighbors({node.qualified_name}, TESTS, incoming)")
neighbors = pg.get_neighbors(
node.qualified_name,
edge_type=EdgeType.TESTS,
direction="incoming",
max_nodes=10,
)
print(f" Found {len(neighbors)} neighbors")
for n, e, d in neighbors[:5]:
print(f" {n.name} @ {n.file_path} (label={n.label})")

pg.close()
Loading
Loading