Measured from this repo's benchmark report, not a marketing placeholder.
reports/benchmarks.md
diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml
index c95f8faa..66e76869 100644
--- a/.github/workflows/ci.yml
+++ b/.github/workflows/ci.yml
@@ -1,6 +1,7 @@
# CI gate for every push and PR: the Node 20/22 test matrix (matches the ">=20" engines
-# field; 18 is EOL) plus the shared quality gate (Biome, typecheck, ShellCheck, zero-dep
-# assertion, version-drift, docs-drift, pack). The quality gate is the SAME reusable workflow
+# field; 18 is EOL), the research contracts (Python prototype suites + recomputation), plus
+# the shared quality gate (Biome, typecheck, ShellCheck, zero-dep assertion, version-drift,
+# docs-drift, pack). The quality gate is the SAME reusable workflow
# the version bump and release require, so none of the three can drift from the others.
name: CI
@@ -58,6 +59,38 @@ jobs:
awk '/^not ok /{p=1} p{print} p&&/^ \.\.\.[[:space:]]*$/{p=0}' /tmp/win-test.log | head -200
exit "$ec"
+ # The research contracts (review A02): the two Python prototypes' own suites, the Theorem-D
+ # sanity checks (no data needed), and the recomputation of every corrected number from the
+ # replication package shipped in this repo. A research claim that stops recomputing fails CI
+ # like a broken unit test.
+ research:
+ name: Research (Python prototypes + recomputation)
+ runs-on: ubuntu-latest
+ steps:
+ - uses: actions/checkout@v7
+ - uses: actions/setup-python@v6
+ with:
+ python-version: "3.12"
+ - name: Install prototype test dependencies
+ run: |
+ python -m pip install --upgrade pip
+ python -m pip install \
+ -r research/python-prototypes/impact_oracle/requirements.txt \
+ -r research/python-prototypes/router_gate/requirements.txt
+ - name: impact_oracle tests
+ working-directory: research/python-prototypes/impact_oracle
+ run: python -m pytest -q
+ - name: router_gate tests
+ working-directory: research/python-prototypes/router_gate
+ run: python -m pytest -q
+ - name: Theorem sanity checks (no data)
+ run: python research/recompute_corrections.py --theorem-checks
+ - name: Recompute the corrections from the replication package
+ run: |
+ mkdir -p "$RUNNER_TEMP/rp"
+ tar -xzf research/empirical-refutation/replication_package.tar.gz -C "$RUNNER_TEMP/rp"
+ python research/recompute_corrections.py "$RUNNER_TEMP/rp/repro"
+
quality-gate:
name: Quality gate
uses: ./.github/workflows/reusable-quality-gate.yml
diff --git a/.github/workflows/reusable-quality-gate.yml b/.github/workflows/reusable-quality-gate.yml
index f3cc0dc4..836a1c93 100644
--- a/.github/workflows/reusable-quality-gate.yml
+++ b/.github/workflows/reusable-quality-gate.yml
@@ -33,4 +33,6 @@ jobs:
run: node -e "process.exit(Object.keys(require('./package.json').dependencies||{}).length)"
- run: node scripts/bump.mjs check
- run: node src/cli.js docs check
+ - name: Claim registry and research copies are current
+ run: node scripts/claims-status.mjs --check
- run: npm pack --dry-run
diff --git a/.gitignore b/.gitignore
index 118cb3a7..a7448fe1 100644
--- a/.gitignore
+++ b/.gitignore
@@ -7,6 +7,11 @@ node_modules/
npm-debug.log*
*.tgz
+# Python bytecode / test caches (research prototypes)
+__pycache__/
+*.pyc
+.pytest_cache/
+
# Forge runtime artifacts (generated, never committed)
.forge/
*.log
diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md
index f7b281be..3c55a1ac 100644
--- a/ARCHITECTURE.md
+++ b/ARCHITECTURE.md
@@ -1,11 +1,13 @@
# forgekit — architecture
-> **One brain for every AI coding agent.** A large language model is stateless: one
-> context window, wiped every call. It has no memory of what your team learned, no
-> foresight about what an edit will break, and no enforced guardrails. forgekit is the
-> **cognitive substrate** — the layer that runs _before_ the model edits code, supplying
-> proof-carrying memory, impact foresight, and enforced guardrails — and a **cross-tool
-> config compiler** that delivers that brain as native config into every tool at once.
+> **A beta toolkit for shared evidence-referenced memory, heuristic change-impact analysis,
+> and explicit verification around coding agents.** A language model keeps no durable state
+> between independent calls and sees only a bounded context window, so on its own it does not
+> carry what your team learned, cannot see the parts of the repository an edit affects unless
+> they are in context, and cannot enforce rules on itself. forgekit is the **cognitive
+> substrate** — the layer that runs _before_ the model edits code, supplying evidence-referenced
+> ("proof-carrying") memory, heuristic impact analysis, and guardrails — and a **cross-tool
+> config compiler** that delivers it as native config into every tool at once.
This document is the architecture reference. It is organized around four diagrams:
@@ -194,10 +196,13 @@ flowchart LR
class SV accent;
```
-The completeness gate on the retrieval side is `forge context "
119 files"]
- src["src
110 files"]
- test["test
121 files"]
- src["src
111 files"]
+ test["test
133 files"]
+ src["src
119 files"]
landing["landing
61 files"]
research["research
37 files"]
+ bench["bench
6 files"]
global["global
5 files"]
- bench["bench
3 files"]
- scripts["scripts
2 files"]
+ scripts["scripts
3 files"]
docs["docs
1 file"]
examples["examples
1 file"]
- test -- 247 --> src
- test -- 244 --> src
- bench -- 8 --> src
+ test -- 281 --> src
+ bench -- 12 --> src
examples -- 4 --> src
+ test -- 3 --> global
+ test -- 3 --> scripts
test -- 2 --> bench
- test -- 2 --> global
- test -- 2 --> scripts
scripts --> src
src --> global
```
diff --git a/CHANGELOG.md b/CHANGELOG.md
index cf9300be..507dc9a4 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -6,6 +6,181 @@ to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
## [Unreleased]
+### Changed
+
+These tighten what a result is allowed to CLAIM, after the 2026-09-26 external deep review
+reproduced 16 cases where a label (`PASS`, `complete`, `exact`, trusted) was stronger than the
+evidence behind it. Scripts that read the JSON output may need to adapt:
+
+- **`forge verify` covers the whole repo, or says what it did not cover.** Every nested
+ package that declares its own suite (an explicit `scripts.test`, a pytest config, go.mod…)
+ is now planned and run in its own directory, and `tests.coverage` reports which packages
+ got a verdict. A passing root suite can no longer hide a failing workspace package (F08).
+ A root script that already runs every workspace (`npm test --workspaces`, `pnpm -r`,
+ `turbo run test`, …) covers them once, without duplicate runs. New `.forge/forge.config.json`
+ keys: `verify.workspaces: "root"` (declare that the root command covers everything),
+ `verify.exclude` (package paths that are not required suites) and `verify.generated`
+ (outputs a test run may legitimately write). Fixture/test-data packages are not required.
+- **A test runner that is only a devDependency is inventory, not an obligation (F09).** With an
+ explicit `scripts.test`, `detectStack().testCommands` no longer adds `npx vitest`/`npx jest`;
+ they are listed in the new `testInventory` (and `forge stack` prints them as "available").
+ npm's `"no test specified"` placeholder is not a suite.
+- **`forge verify` refuses to bind a verdict to code that changed while the tests ran (F10).**
+ The code state is captured before and after the run; if it moved, the result is
+ `INCOMPLETE` with `mutated: true`, and the stamp is bound to the PRE-run state. Interpreter
+ caches (`__pycache__`, `.pytest_cache`, …) never count as a change.
+- **The code-state fingerprint is a canonical manifest bound to HEAD (F01).** Renaming an
+ untracked file, moving bytes between files, adding an empty file, changing an exec bit or a
+ symlink target, and checking out another commit all change it; an unreadable untracked file
+ makes the state unbindable. Stamps from older forge versions no longer verify (by design:
+ their fingerprint could not tell those states apart) — re-run `forge verify`.
+- **Exact reuse keys are lossless except whitespace (F04).** Case, operators, literals and
+ punctuation are part of the key, so `>= 18` / `<= 18`, `"ADMIN"` / `"admin"`, `= true` /
+ `!= true`, `getURL` / `getUrl` never share one. A near candidate must also pass a semantic
+ guard (same operators, numbers, literals, identifiers, paths and polarity words) or it is
+ only offered at the adapt tier. Artifacts minted before this change never exact-hit.
+- **`forge context` reports delivery, not availability (F02, F03).** `tokens` is measured on the
+ rendered block (a chars/3.6 estimate, now labeled as such) and never exceeds `--budget`
+ while the result claims success: when even pointers cannot fit, items are dropped and the
+ result is `overflow: true`, `ok: false`. A `- read
ForgeKit gives every AI coding tool the same memory, foresight, and guardrails—without locking your work inside one vendor or one chat window.
Models are capable. Their operating context is fragile. ForgeKit supplies the durable layer that travels with the repository and shows up before the next action.
Forge keeps decisions, lessons, and project state in the repository—so Claude, Codex, Cursor, and the next agent all inherit the same working memory.
3 records recalledForge turns agent behavior into a reviewable sequence. Each meaningful move begins with context and ends with proof.
Load relevant decisions and lessons.
Measure scope, cost, and reversibility.
Map downstream surfaces before editing.
Pause risky or under-specified actions.
Record what changed and how it was verified.
One source emits each tool’s native configuration. Your rules and memory stay with the project—not the provider.
Plus MCP configuration for Roo Code and VS Code-compatible clients.
ForgeKit publishes the measurements behind its claims. The numbers below come from repository benchmarks and evaluation reports—not a marketing dashboard.
ForgeKit improves agent judgment; it does not replace yours. The project labels its assumptions so you can decide where to trust, test, or intervene.
Claude Code is the deepest-tested integration. Other targets have less real-world exercise today.
Blast-radius analysis is heuristic. It guides review; it is not a formal dependency proof.
Guardrails are not a sandbox. Keep permissions, review, and backups appropriate to the work.
Install ForgeKit, run forge init in your repository, and keep one shared operating context across every tool.
$ /plugin marketplace add CodeWithJuber/forgekit
+Skip to contentforge/kitGitHub Open source cognitive substrateforgekit v1.4.3 · betaOne operating
memory. Every
coding agent.
ForgeKit gives every AI coding tool the same memory and foresight—with automatic guardrails on Claude Code—without locking your work inside one vendor or one chat window.
- Runtime deps
- 0
- Native targets
- 9
- License
- MIT
FK / PREFLIGHTSYSTEM READY01REQUESTRefactor authentication flow00:118- 01Memory recalledPASS
- 02Blast radius mappedPASS
- 03Guardrails checkedPASS
TRACE FK-031-7D4PROCEED → 01 / The substrateState before actionThe missing layer between
your intent and your agent.
Models are capable. Their operating context is fragile. ForgeKit supplies the durable layer that travels with the repository and shows up before the next action.
ACTIVE CAPABILITY / 01Context that survives the chat.
Forge keeps decisions, lessons, and project state in the repository—so Claude, Codex, Cursor, and the next agent all inherit the same working memory.
3 records recalledTYPERECORDSTATEdecisionUse SQLite for local-first state94%lessonRun schema checks before generation88%preferenceKeep the CLI dependency-free82% 02 / The protocolOne request · five checks · one traceAction should leave evidence.
Forge turns agent behavior into a reviewable sequence. Each meaningful move begins with context and ends with proof.
- 01Recall
Load relevant decisions and lessons.
- 02Classify
Measure scope, cost, and reversibility.
- 03Foresee
Map downstream surfaces before editing.
- 04Gate
Pause risky or under-specified actions.
- 05Trace
Record what changed and how it was verified.
03 / One sourceNine native targetsChange the agent. Keep the operating system.
One source emits each tool’s native configuration. Your rules and memory stay with the project—not the provider.
- 01CCClaude Code
- 02CXCodex
- 03CRCursor
- 04GMGemini
- 05AIAider
- 06CPCopilot
- 07WSWindsurf
- 08ZDZed
- 09CTContinue
Plus MCP configuration for Roo Code and VS Code-compatible clients.
04 / Evidence ledgerMeasured, not inventedFast enough to stay in the loop.
ForgeKit publishes the measurements behind its claims. The numbers below come from repository benchmarks and evaluation reports—not a marketing dashboard.
- Pre-action gate
- 851ms
End-to-end benchmark- Blast-radius scan
- 1.68ms
Heuristic analysis- Held-out routing cost
- +20.2%
vs always-premium, 80 tasks- Runtime dependencies
- 0
Node.js standard library
Inspect the evidence 05 / Honest limitsProfessional, not magicalThe guardrail is not the road.
ForgeKit improves agent judgment; it does not replace yours. The project labels its assumptions so you can decide where to trust, test, or intervene.
- 01
Claude Code is the deepest-tested integration. Other targets have less real-world exercise today.
- 02
Blast-radius analysis is heuristic. It guides review; it is not a formal dependency proof.
- 03
Guardrails are not a sandbox. Keep permissions, review, and backups appropriate to the work.
06 / Start hereAbout sixty secondsGive the next agent a better starting point.
Install ForgeKit, run forge init in your repository, and keep one shared operating context across every tool.
Open the quickstart $ /plugin marketplace add CodeWithJuber/forgekit
$ /plugin install forgekit
Recommended · ambient guards on every prompt
diff --git a/mintlify/cli/substrate.mdx b/mintlify/cli/substrate.mdx
index f399e3d9..eb8439b0 100644
--- a/mintlify/cli/substrate.mdx
+++ b/mintlify/cli/substrate.mdx
@@ -110,9 +110,13 @@ Consequence simulation — predicted breaks + the minimal dry-run test suite for
```bash
forge imagine ""
-forge imagine "" --run # execute the minimal suite sandboxed
+forge imagine "" --run # execute the minimal suite in an isolated git checkout
```
+`--run` executes the suite in an isolated git checkout (a detached-HEAD worktree). It isolates
+checkout files only — not network, credentials, your home directory or process permissions — and
+it tests the committed baseline, not uncommitted changes.
+
## `forge lean`
Scope-minimality (M5) — measure the diff's footprint vs what the task asked for.
diff --git a/mintlify/introduction.mdx b/mintlify/introduction.mdx
index 144c5307..8808cc9c 100644
--- a/mintlify/introduction.mdx
+++ b/mintlify/introduction.mdx
@@ -1,17 +1,21 @@
---
title: "Forge: the cognitive substrate for AI coding agents"
-description: "Forge is the cognitive substrate stateless models are missing — memory, foresight, and guardrails — as native config for every AI coding agent."
+description: "A beta toolkit for shared evidence-referenced memory, heuristic change-impact analysis, and explicit verification around coding agents — as native config for every AI coding tool."
---
-**One brain for every AI coding agent.** A large language model is stateless: one
-context window, wiped every call. It has no memory of what your team learned, no
-foresight about what an edit will break, and no enforced guardrails. Forge
+**A beta toolkit for shared evidence-referenced memory, heuristic change-impact analysis,
+and explicit verification around coding agents.** A language model keeps no durable state between
+independent calls and sees only a bounded context window, so on its own it does not carry what
+your team learned, cannot see the parts of the repository an edit affects unless they are in
+context, and cannot enforce rules on itself. Forge
(`@codewithjuber/forgekit`) is the **cognitive substrate** — the layer that runs
_before_ the model edits code, supplying evidence-referenced, content-addressed memory (we
-call it "proof-carrying memory"), heuristic impact foresight, and enforced guardrails — and
-a **cross-tool config compiler** that delivers that brain as native config into every tool
-at once. Claude Code is the deepest-tested integration; the others receive native config and
-MCP tools with less real-world exercise.
+call it "proof-carrying memory"), heuristic impact foresight, and guardrails (blocking on
+Claude Code) — and a **cross-tool config compiler** that delivers that brain as native config
+into every tool at once. Claude Code is the deepest-tested integration and the only one with
+automatic hooks; the others receive native config (most also an MCP server entry) with less
+real-world exercise. The repository's `docs/INTEGRATIONS.md` lists, per tool, what is emitted,
+registered, run automatically and enforced.
@@ -30,10 +34,11 @@ MCP tools with less real-world exercise.
## The problem
-A large language model is stateless — one context window, wiped every call.
+A language model keeps no durable state between independent calls, and it sees only what fits in
+its context window.
-- It has **no memory** of what your team already learned.
-- It has **no foresight** about what an edit will break.
+- It does not **carry what your team learned** from one session to the next.
+- It does not **see what an edit will affect** unless those files are in its context.
- It has **no enforced guardrails** — prose rules get forgotten after a compaction.
And every tool wants its own config file (`CLAUDE.md`, `AGENTS.md`, `.cursor/rules`,
@@ -42,11 +47,12 @@ things, and the compiler that delivers it into every tool from one source.
## The thesis
-A model can't learn from your codebase between calls: its weights are frozen and its
-working memory is wiped after every response. Memory, foresight, and self-checking
-can't be prompted into it — they have to be supplied from _outside_. That outside layer
-is the cognitive substrate. Formally, inference is a fixed function `y = f(x)` with no
-state between calls; Forge is the state.
+A model with frozen weights does adapt within one context — examples, retrieved facts and
+feedback change what it does — but it does not guarantee four things on its own: durable state
+across independent calls, context beyond its window, weight updates from outcomes, and reliable
+self-verification without external evidence. Forge supplies persistence and external checks from
+_outside_ the model. It is one tested way of doing that, not the only possible architecture;
+the research programme's dated corrections explain the difference.
@@ -75,8 +81,9 @@ state between calls; Forge is the state.
content-addressed memory: a claim that carries references to its evidence and is trusted
only once independent oracles raise its confidence above a floor. The "proof" is that
evidence trail, not a formal proof.
-- **Foresight before you break things.** Ask "what does changing `verifyToken` break?"
- and get the blast radius from the code graph, including coupled files you never named.
+- **Heuristic impact before you break things.** Ask "what does changing `verifyToken`
+ break?" and get the blast radius from a regex-derived code graph, including coupled files you
+ never named — it can miss files as well as over-warn.
- **Guardrails that can't be forgotten.** Deterministic hooks enforce protected paths,
cost budgets, and doom-loop detection — they survive a context compaction.
- **Work that finishes end to end.** A completion gate blocks "done" once per session
@@ -117,8 +124,10 @@ Forge states its own ceiling everywhere.
- **Tests and human corrections always win.**
- Forge is **beta**. The core (`init`, `sync`, `substrate`, `impact`, `ledger`, guards)
- is tested and in daily use; some flags may change before `1.0`.
+ Forge is **beta**: releases follow semantic versioning (a breaking change is a new major
+ version), and "beta" describes maturity — heuristic analyses, advisory checks, and less
+ real-world exercise outside Claude Code. The core (`init`, `sync`, `substrate`, `impact`,
+ `ledger`, guards) is tested and in daily use.
## Next steps
diff --git a/reports/2026-07-05.md b/reports/2026-07-05.md
index f4542af8..8e89e492 100644
--- a/reports/2026-07-05.md
+++ b/reports/2026-07-05.md
@@ -1,5 +1,11 @@
# Claude Config Radar — 2026-07-05
+> ⚠️ **Archived — historical.** An auto-generated daily report from 2026-07-05, kept for the record.
+> Not maintained and not the current state; its action items (including the credential-rotation
+> reminders, which mention no secret values) are historical and superseded. See
+> **[docs/GUIDE.md](../docs/GUIDE.md)** for today's tooling and **[SECURITY.md](../SECURITY.md)** for
+> reporting a live security issue.
+
**TL;DR**
- 🔴 **Security-relevant dep bump**: pgvector **0.8.2** fixes a buffer overflow in parallel HNSW index builds (**CVE-2026-3172**) — your stack uses pgvector, so this is the one worth acting on.
- 🟡 **shadcn/ui** made **Base UI the default** component library (docs + new projects) as of July 2026 — escalation of yesterday's "picking Base UI ~2:1" note; Radix not deprecated.
diff --git a/reports/benchmarks.md b/reports/benchmarks.md
index b52e2f9c..e27ae4c1 100644
--- a/reports/benchmarks.md
+++ b/reports/benchmarks.md
@@ -4,7 +4,7 @@
> **a number is an assumption until measured.** Every figure in the generated section below
> came from an actual run of `npm run bench` on the machine recorded in the environment
> block — no projections, no targets, no numbers copied forward from a different machine.
-> Re-run `npm run bench` (≈10 s, node stdlib only) and the generated section is rewritten
+> Re-run `npm run bench` (≈20 s, node stdlib only) and the generated section is rewritten
> in place with your machine's numbers.
## Methodology
@@ -45,12 +45,22 @@
- **ledger / val()**: pure in-memory scoring; the fixture gives most claims 0–1 evidence
records, and val() cost scales with evidence count — a heavily-evidenced ledger will be
slower per claim.
-- **reuse / lookup**: the memoized `_sketch` cache is stripped before every timed run, so
- each run behaves like a fresh CLI process. The *exact* tier returns before any pool
- sketching (normalized-string compare); the *near* tier pays MinHash-sketching the whole
- candidate pool plus LSH banding — that difference is the point of reporting both.
+- **reuse / lookup**: every row is VALIDATED before it is timed — the fixture's artifacts
+ cite a real, resolvable git object as their test evidence, and the harness aborts unless
+ the lookup returns the tier the row is labeled with (review F14, 2026-09-26: the previous
+ fixture cited untyped `bench:artifact:` refs, which val() caps below the serving floor,
+ so every "exact"/"near" row had actually measured a MISS; those older numbers are
+ invalid). "cold" rows strip every memoized sketch the lookup path caches (`_sketch`,
+ `_terms`, `_specSketch`, `_keySketch` — the old harness stripped only `_sketch`, which the
+ reuse ladder never reads), so each run behaves like a fresh CLI process; "warm" rows keep
+ them, like a long-lived process. The *exact* tier returns before any pool sketching
+ (identity-key compare); *near* and *miss* pay MinHash-sketching the candidate pool plus
+ LSH banding — that difference is the point of reporting them separately.
- **context / assemble()**: warm atlas, empty ledger (the repo copy has no `.forge`),
includes the real file reads for pinned items. Task: a three-symbol, one-file edit spec.
+ "complete"/"incomplete" is the assembler's own honest verdict: an item that could only
+ be delivered as a pointer or partial span is a pending read, so a large named file makes
+ this task's context incomplete at the default budget (review F02/F03).
- **substrate / substrateCheck**: the whole deterministic gate — preflight grounding,
routing rubric, up to 8 impact queries, reuse lookup, context assembly, scope
decomposition, lessons, minimality, goal anchor — with `llm: false`. **No model latency
@@ -90,9 +100,12 @@ now resolve to the exact symbol (`src/atlas.js:17 imports → src/util.js:conten
What these numbers do **not** mean: n = 6 cases, one JavaScript repo, symbols chosen to be
uniquely named (the atlas resolves ambiguous names to nothing — a separate, known
-limitation). They are not comparable to the paper's numbers, which came from mutation
-testing a Python codebase against a real test suite. The two appear side by side below,
-labeled, and are never blended.
+limitation). They are not comparable to the paper prototype's numbers: its 0.63 / 1.00 /
+0.75 came from mutation testing on the authors' own fixture — a self-built demo that the
+pre-registered field study REFUTED (pooled precision 0.40, recall 0.022, F1 0.042 over 759
+files' mined co-change in nine repositories; see `research/empirical-refutation/`). The
+regex atlas here is a different, Node graph — not the evaluated Python oracle. All three
+appear side by side below, labeled, and are never blended.
> **History of this row.** The precision 0.90 / F1 0.92 this file carried until 2026-09-21 came
> from a much smaller atlas (145 files) and a reverse walk that stopped at the direct
@@ -114,61 +127,69 @@ labeled, and are never blended.
```json
{
- "node": "v24.19.0",
- "cpu": "AMD EPYC Processor (with IBPB)",
+ "node": "v22.22.2",
+ "cpu": "Intel(R) Xeon(R) Processor @ 2.80GHz",
"cores": 4,
- "memGB": 8,
- "platform": "win32",
+ "memGB": 16,
+ "platform": "linux",
"arch": "x64",
- "commit": "703da31d574c30d22bef019b1c8563ade0d0d6be",
- "date": "2026-09-21T22:09:58.366Z"
+ "fsType": "ext2/ext3",
+ "commit": "a56606afd5baebdabee95ad95af626a965557d92 + uncommitted changes",
+ "date": "2026-09-26T20:35:15.125Z"
}
```
### Measured results
-| suite | benchmark | median | p95 | runs | notes |
-|-----------|---------------------------------------------|----------|---------|------|----------------------------------------|
-| atlas | full build (this repo) | 530 ms | 622 ms | 5 | 455 files, 10498 symbols, 29728 edges |
-| atlas | incremental rebuild (unchanged) | 339 ms | 359 ms | 5 | per-file hash cache hit |
-| atlas | impact("claimText") (warm adjacency) | 0.40 ms | 1.08 ms | 30 | 51 files impacted |
-| ledger | mint+put 1000 claims | 1854 ms | 1986 ms | 5 | 539/s |
-| ledger | loadClaims at 1000 claims | 213 ms | 230 ms | 5 | full state from disk |
-| ledger | mergeDirs 2×500-claim replicas (250 shared) | 4308 ms | 4409 ms | 3 | +250 claims, +313 records |
-| ledger | val() over 1000 claims | 0.076 ms | 0.16 ms | 20 | 13,140,604/s (mean val 0.51) |
-| reuse | fingerprint 2000 specs | 116 ms | 156 ms | 5 | 17,171/s |
-| reuse | lookup exact @ 100 artifacts | 5.71 ms | 10.5 ms | 10 | tier=miss |
-| reuse | lookup near (LSH) @ 100 artifacts | 4.76 ms | 5.18 ms | 5 | tier=miss, j=- |
-| reuse | lookup exact @ 1000 artifacts | 52.8 ms | 88.5 ms | 10 | tier=miss |
-| reuse | lookup near (LSH) @ 1000 artifacts | 46.7 ms | 89.8 ms | 5 | tier=miss, j=- |
-| context | assemble() (this repo, 3-symbol task) | 12.6 ms | 30.0 ms | 10 | 4070/6000 tokens, 9 required, complete |
-| substrate | substrateCheck (allowBuild, llm off) | 886 ms | 908 ms | 3 | 99 impacted files, route simple |
+| suite | benchmark | median | p95 | runs | notes |
+|-----------|----------------------------------------------|---------|---------|------|--------------------------------------------------------------|
+| atlas | full build (this repo) | 719 ms | 753 ms | 5 | 504 files, 13396 symbols, 37369 edges, cap 20000 not reached |
+| atlas | incremental rebuild (unchanged) | 346 ms | 365 ms | 5 | per-file hash cache hit |
+| atlas | impact("claimText") (warm adjacency) | 1.68 ms | 2.38 ms | 30 | 61 files impacted |
+| ledger | mint+put 1000 claims | 248 ms | 326 ms | 5 | 4,038/s |
+| ledger | loadClaims at 1000 claims | 13.6 ms | 15.3 ms | 5 | full state from disk |
+| ledger | mergeDirs 2×500-claim replicas (250 shared) | 191 ms | 198 ms | 3 | +250 claims, +313 records |
+| ledger | val() over 1000 claims | 0.85 ms | 1.39 ms | 20 | 1,170,474/s (mean val 0.51) |
+| reuse | fingerprint 2000 specs | 232 ms | 290 ms | 5 | 8,636/s |
+| reuse | lookup exact hit, cold @ 100 artifacts | 1.01 ms | 9.43 ms | 10 | tier=exact |
+| reuse | lookup exact hit, warm @ 100 artifacts | 0.67 ms | 5.89 ms | 10 | tier=exact |
+| reuse | lookup near hit (LSH), cold @ 100 artifacts | 18.5 ms | 33.0 ms | 5 | tier=near, j=0.98 |
+| reuse | lookup miss, cold @ 100 artifacts | 18.2 ms | 25.4 ms | 5 | tier=miss |
+| reuse | lookup exact hit, cold @ 1000 artifacts | 3.98 ms | 6.11 ms | 10 | tier=exact |
+| reuse | lookup exact hit, warm @ 1000 artifacts | 3.38 ms | 4.14 ms | 10 | tier=exact |
+| reuse | lookup near hit (LSH), cold @ 1000 artifacts | 128 ms | 133 ms | 5 | tier=near, j=0.95 |
+| reuse | lookup miss, cold @ 1000 artifacts | 123 ms | 130 ms | 5 | tier=miss |
+| context | assemble() (this repo, 3-symbol task) | 17.6 ms | 27.0 ms | 10 | 4174/6000 tokens, 9 required, incomplete |
+| substrate | substrateCheck (allowBuild, llm off) | 851 ms | 995 ms | 3 | 141 impacted files, route simple |
### Impact-oracle quality (hand-labeled cases, this repo)
| case (target) | precision | recall | F1 | predicted | truth |
|---------------|-----------|--------|------|-----------|-------|
-| normalizeSpec | 0.12 | 1.00 | 0.21 | 17 | 2 |
+| normalizeSpec | 0.11 | 1.00 | 0.20 | 18 | 2 |
| evalImpact | 0.29 | 1.00 | 0.44 | 7 | 2 |
-| isStale | 0.20 | 1.00 | 0.33 | 30 | 6 |
-| mergeStates | 0.16 | 1.00 | 0.28 | 25 | 4 |
-| claimText | 0.16 | 1.00 | 0.27 | 51 | 8 |
-| contentHash | 0.11 | 1.00 | 0.21 | 87 | 10 |
-| mean of 6 | 0.17 | 1.00 | 0.29 | | |
+| isStale | 0.17 | 1.00 | 0.30 | 40 | 7 |
+| mergeStates | 0.14 | 1.00 | 0.25 | 28 | 4 |
+| claimText | 0.18 | 1.00 | 0.31 | 61 | 11 |
+| contentHash | 0.10 | 1.00 | 0.19 | 105 | 11 |
+| mean of 6 | 0.17 | 1.00 | 0.28 | | |
-Edited-file-only baseline recall over the same cases: **0.27**.
+Edited-file-only baseline recall over the same cases: **0.26**.
-Two methodologies, side by side — different codebases, different ground-truth
+Different methodologies, side by side — different codebases, different ground-truth
derivations, so the rows are comparable in spirit only and are never blended:
-| series | precision | recall | F1 | ground truth |
-|--------------------------------------------|-----------|--------|------|-----------------------------------------------|
-| paper prototype (Python, mutation-derived) | 0.63 | 1.00 | 0.75 | mutation testing against a real suite |
-| this repo (regex atlas, hand-labeled) | 0.17 | 1.00 | 0.29 | 6 hand-labeled cases (bench/impact_cases.mjs) |
+| series | precision | recall | F1 | ground truth |
+|------------------------------------------------|-----------|--------|------|------------------------------------------------------------|
+| paper prototype, self-built demo (REFUTED) | 0.63 | 1.00 | 0.75 | mutation testing on the authors' own fixture |
+| paper prototype, field study (pooled, 9 repos) | 0.40 | 0.02 | 0.04 | 759 files' mined co-change (research/empirical-refutation) |
+| this repo (regex atlas, hand-labeled) | 0.17 | 1.00 | 0.28 | 6 hand-labeled cases (bench/impact_cases.mjs) |
-> **Snapshot boundary.** The measured results above were generated at commit `eb68ea9` and
+> **Snapshot boundary.** The measured results above were generated at the commit recorded in
+> the environment block (`commit`; "+ uncommitted changes" means the working tree the review
+> fixes of 2026-09-26 were measured in, before they were committed) and
> do not benchmark the optional embedding adapter now implemented in `src/embed.js` and
> exercised with a deterministic fake provider in `test/embed.test.js`. MinHash remains the
> zero-dependency default and failure fallback. The structural comparisons below describe the
@@ -216,6 +237,6 @@ that forgekit structurally does not.
## Reproduce
```sh
-npm run bench # ≈10 s; prints the tables and rewrites the generated section above
+npm run bench # ≈20 s; prints the tables and rewrites the generated section above
npm test # includes a smoke test of the harness's pure helpers (test/bench.test.js)
```
diff --git a/reports/cost-eval.md b/reports/cost-eval.md
index 0c66fd25..7171f28f 100644
--- a/reports/cost-eval.md
+++ b/reports/cost-eval.md
@@ -6,21 +6,30 @@
> saving. The paper's 62 % routing saving (paper §9) was measured on the 30 tasks its
> thresholds were tuned on and is **refuted**: on 80 held-out tasks, counting every escalation,
> routing cost 20.2 % _more_ than always-premium ([research/empirical-refutation/](../research/empirical-refutation/)).
-> The plan's ~90 % composed figure is a **target**, not a result, and does not appear in this table.
+> The plan's ~90 % figure is a **target** and a hypothesis, not a result, and does not appear in this table.
+>
+> **Corrected 2026-09-26.** The methodology below used to say the cost model "is multiplicative —
+> `C = C₀ · Π(1 − fᵢ)` over independent stages — so each stage factor is measured separately and
+> composed arithmetically", and that paired runs reprice "identical tokens". Stage savings interact
+> (cache hits change the routed workload, context changes retries, halts can defer work), so the
+> per-stage factors are diagnostics, not a total; and repricing tokens at another model's price is a
+> counterfactual, not an observed outcome. The acceptance rule for any cost headline is in
+> [05-cost-model.md §3](../docs/plans/substrate-v2/05-cost-model.md#3-acceptance-rule-for-any-cost-headline).
## Methodology
-The cost model is multiplicative — `C = C₀ · Π(1 − fᵢ)` over independent stages — so each
-stage factor is measured separately and composed arithmetically, never asserted:
+Each stage factor is measured separately as a diagnostic; the system is judged only on paired,
+full-system outcomes (total cost per completed, externally verified task), never on a product of
+stage factors:
1. **Instrumentation.** Every substrate stage appends one line to `.forge/metrics.jsonl`
(`{t, stage, outcome, tokensIn, tokensOut, tier, savedEstimate, ref}` — `src/metrics.js`).
`forge cost --stages` computes the per-stage factors from those lines (`src/cost_report.js`);
a stage with no events reports **no data**, never a default.
-2. **Paired runs.** Baseline (always-premium, read-everything, no cache) vs. substrate over
- the same replay corpus (N ≥ 100 real tasks, stratified repeat-heavy / mixed / cold), the
- paper §9 methodology: identical tokens repriced, so every saving is arithmetic on measured
- tokens.
+2. **Paired runs.** Baseline (equivalent tools, context and repair opportunity; always-premium /
+ read-everything reported too) vs. substrate over the same replay corpus (N ≥ 100 real tasks,
+ stratified repeat-heavy / mixed / cold), each policy actually executed. Repriced tokens are a
+ labelled counterfactual, not a measured saving.
3. **Correctness guard (spec §3).** A saving counts only if the external verifier passes the
output. A routed-down answer that fails is not a saving; a cache hit that gets reverted is
recorded as a *negative* entry.
@@ -33,7 +42,7 @@ stage factor is measured separately and composed arithmetically, never asserted:
| cache (reuse, tier-weighted) | — | 0 | no data yet — run with metrics enabled |
| route (vs always-premium) | — | 0 | no data yet — run with metrics enabled |
| context (assembly ρ) | — | 0 | no data yet — run with metrics enabled |
-| **composed (measured stages only)** | — | 0 | nothing to compose yet |
+| **composed (measured stages only; diagnostic, not a total)** | — | 0 | nothing to compose yet |
Secondary counters (doom-loop halts avoided, M5 lean, avoided rework) are reported alongside
when populated — they are deliberately excluded from the multiplication (spec §1).
diff --git a/research/HISTORICAL_EDITIONS.md b/research/HISTORICAL_EDITIONS.md
new file mode 100644
index 00000000..0ad05f00
--- /dev/null
+++ b/research/HISTORICAL_EDITIONS.md
@@ -0,0 +1,120 @@
+# Historical editions of the research papers
+
+> ⚠️ **Historical, pre-correction editions.** Every PDF listed here predates the 2026-09-21 and
+> 2026-09-26 corrections. It is kept so that what was published can still be read and cited, not
+> as the current text. The corrected sources are the HTML and LaTeX files named in the table;
+> read those.
+
+## Status on 2026-09-26: not regenerated
+
+A re-render of the three HTML papers was attempted on 2026-09-26 and **not committed**:
+
+1. **Figures need resolving.** The HTML sources reference every figure through a
+ `{{artifact:…}}` placeholder, which no browser resolves, so a plain render has no figures. The
+ map below resolves them; it was checked against the images embedded in the old PDFs.
+2. **The Qur'anic text could not be verified.** All three HTML papers carry Qur'anic Arabic. When
+ the white paper and the synthesis were rendered in the environment available that day, Chromium
+ set their verse text in three fallback fonts at once (DejaVu Sans for most glyphs, Liberation
+ Serif and FreeSerif for the rest); mixing fonts inside a word can break letter joining and mark
+ placement, and the rendered pages could not be inspected by eye. A PDF whose sacred text has not
+ been checked is not published as the new edition.
+3. **The refutation paper needs TeX.** `empirical-refutation/paper.pdf` is built from
+ `paper/main.tex` with a TeX Live toolchain; none was available.
+
+## The editions
+
+Each edition stays retrievable byte for byte at the pinned commit
+`d2abfa69fb77531199ffc67c5c076b524af69040`:
+`git show d2abfa69fb77531199ffc67c5c076b524af69040: > edition.pdf`, or
+`git cat-file -p ` with the blob below.
+
+| PDF (historical, pre-correction) | Git blob at `d2abfa6` | sha256 (prefix) | Bytes | Last changed in | Corrected source |
+| --- | --- | --- | --- | --- | --- |
+| `research/formal-synthesis/substrate_synthesis.pdf` | `2e17362fc62d9f32b1083f17a0ac865704244ef6` | `644e28d0f8d1cbd3…` | 835072 | `5e60069` (2026-08-14) | `research/formal-synthesis/substrate_synthesis.html` |
+| `research/empirical-refutation/extended_preprint.pdf` | `74f74ae08612bb9dc11038f651f923e0730b8bfc` | `8d2ec1091d17f0ba…` | 791707 | `9ebe256` (2026-09-20) | `research/empirical-refutation/extended_preprint.html` |
+| `research/empirical-refutation/paper.pdf` | `f94a727cec7f84ac197057f3f22cde0fb09f28b1` | `a5001d8fc59a1b4a…` | 830746 | `c5fb041` (2026-09-20) | `research/empirical-refutation/paper/main.tex` |
+| `research/cognitive-substrate/cognitive_substrate_whitepaper.pdf` | `44ce7bbd4a7bc1e6220f162074c9b473c6287e7b` | `599e626ba24958c6…` | 1816584 | `e6e6de7` (2026-09-20) | `research/cognitive-substrate/cognitive_substrate_whitepaper.html` |
+| `docs/cognitive-substrate/cognitive_substrate_whitepaper.pdf` (byte-identical copy) | same blob as the row above | same | 1816584 | — | the same HTML, copied to `docs/cognitive-substrate/` |
+
+The copies of `repro/paper/main.tex` and `repro/paper/paper.pdf` inside
+`empirical-refutation/replication_package.tar.gz` (git blob `50bd453a30dad5d8ca3369129f8015fd4524fa81`)
+are also left exactly as published. The paper PDF was built with pdfTeX (TeX Live 2026) and the ACM
+`acmart` class, as its own metadata records.
+
+## Figure map
+
+Each `{{artifact:}}` placeholder in the HTML sources (13 in all: 3 in the synthesis, 3 in the
+preprint, 7 in the white paper), the figure file under `research/` it stands for, and whether that
+file's pixel size matches the image embedded in the old PDF:
+
+| Placeholder id | Figure (under `research/`) | Size (px) | Matches the old PDF |
+| --- | --- | --- | --- |
+| `art_5f049677-4c3c-40f2-8905-dd01c966e9ae` | `formal-synthesis/figures/schematic_duality.png` (synthesis and preprint, Figure 1) | 1366 × 1046 | yes |
+| `art_f0decf80-d016-496f-8032-f7b2e73e71b2` | `formal-synthesis/figures/schematic_taskloop.png` (synthesis and preprint, Figure 2) | 1607 × 1092 | yes |
+| `art_5d076ce3-0f54-4394-9a77-f70a336ca843` | `formal-synthesis/figures/schematic_convergence.png` (synthesis, Figure 8) | 2460 × 1539 | yes |
+| `art_712fac51-fe17-4ed6-80f0-9dd42bf42758` | `empirical-refutation/figures/fig_repair_beforeafter.png` (preprint, Figure 3) | 3142 × 1383 | yes |
+| `art_e2776474-3d1e-48c6-9490-55d4d927a301` | `cognitive-substrate/figures/schematic_loop.png` (white paper, Figure 1) | 3003 × 1439 | yes |
+| `art_d2be1b53-86ce-4069-b2aa-5be59836598e` | `cognitive-substrate/figures/schematic_system.png` (white paper, Figure 2) | 2847 × 1840 | yes |
+| `art_8d9fa6dd-3554-49c7-9e76-ba667544a622` | `cognitive-substrate/figures/schematic_extended.png` (white paper, Figure 3) | 1483 × 931 | yes |
+| `art_07bb9186-5e55-44f6-af3f-dde83d6b9e65` | `cognitive-substrate/figures/impact_graph.png` (white paper, Figure 4) | 1900 × 1326 | yes |
+| `art_392e293d-be93-4efe-81c1-e9612a711ac4` | `cognitive-substrate/figures/eval_precision_recall.png` (white paper, Figure 5) | 2300 × 918 | **no** — the old PDF embeds a 1921 × 842 raster, so the repository's file is a different render of this figure; compare the two before publishing |
+| `art_5b206b7b-c90e-417f-b1e8-48b0ec389cb8` | `cognitive-substrate/figures/schematic_router_loop.png` (white paper, Figure 6) | 1537 × 838 | yes |
+| `art_ac78be07-be03-4560-bb9f-f5fe2f16d7ef` | `cognitive-substrate/figures/router_eval.png` (white paper, Figure 7) | 1719 × 732 | yes |
+
+## Rendering a new edition
+
+1. Work in a scratch directory outside the repository and install `playwright-core` there, never
+ in the repository: `npm init -y && npm install playwright-core`. Point it at an installed
+ Chromium (`executablePath`).
+2. Install a font with full Qur'anic coverage (a Naskh face such as Amiri or Scheherazade New) and
+ make it the first `font-family` for `.quran .ar` in the render, so one font sets each verse.
+3. Replace every `{{artifact:}}` with its figure from the map (a `data:image/png;base64,…`
+ URI keeps the render self-contained), render A4 with a header and footer stamp
+ `edition · source sha256 · forgekit `,
+ and confirm every image loaded and no request failed.
+4. **Inspect by eye** every page that carries a figure or a verse card. Only then replace the PDF,
+ re-copy the white paper to `docs/cognitive-substrate/` (`node scripts/claims-status.mjs
+ --sync-copies`), and add a row here recording the replaced edition's git blob and the new
+ edition's source hash.
+
+A render script that does steps 1 and 3 (it resolved all 13 figure placeholders across the three
+papers with no failed request on 2026-09-26):
+
+```js
+// node render.mjs (run from the scratch directory)
+import { createHash } from "node:crypto";
+import { readFileSync } from "node:fs";
+import path from "node:path";
+import { chromium } from "playwright-core";
+
+const FIGURES = { /* "art_…": "research/…/figures/….png", one entry per row of the map above */ };
+const [repo, rel, out] = process.argv.slice(2);
+const raw = readFileSync(path.join(repo, rel));
+const sha = createHash("sha256").update(raw).digest("hex").slice(0, 12);
+const version = JSON.parse(readFileSync(path.join(repo, "package.json"), "utf8")).version;
+const stamp = `edition ${new Date().toISOString().slice(0, 10)} · source sha256 ${sha} · forgekit ${version}`;
+let html = raw.toString("utf8");
+for (const [id, fig] of Object.entries(FIGURES)) {
+ const uri = `data:image/png;base64,${readFileSync(path.join(repo, fig)).toString("base64")}`;
+ html = html.split(`{{artifact:${id}}}`).join(uri);
+}
+const browser = await chromium.launch({ executablePath: process.env.CHROMIUM });
+const page = await browser.newPage();
+const failed = [];
+page.on("requestfailed", (r) => failed.push(r.url()));
+await page.setContent(html, { waitUntil: "networkidle" });
+const unresolved = (html.match(/\{\{artifact:[^}]+\}\}/g) || []).length;
+const broken = await page.evaluate(() => [...document.images].filter((i) => !i.naturalWidth).length);
+if (unresolved || broken || failed.length) throw new Error(`figures: ${unresolved} unresolved, ${broken} broken, ${failed.length} failed`);
+const line = (s) => `${s}`;
+await page.pdf({
+ path: out, format: "A4", printBackground: true, displayHeaderFooter: true,
+ margin: { top: "18mm", bottom: "18mm", left: "14mm", right: "14mm" },
+ headerTemplate: line(stamp),
+ footerTemplate: line(`${stamp} · page / `),
+});
+await browser.close();
+```
+
+For `paper.pdf`, build `paper/main.tex` with TeX Live (`pdflatex` and `bibtex`, ACM `acmart`
+class) and stamp the same fields in the PDF metadata or a footnote.
diff --git a/research/README.md b/research/README.md
index 6cac916d..8eea9665 100644
--- a/research/README.md
+++ b/research/README.md
@@ -1,56 +1,122 @@
# Research
-The full research programme behind forgekit: a theory of what a frozen language model
-structurally lacks, an architecture that supplies it, two runnable prototypes, and — most
-importantly — a pre-registered empirical evaluation that **refuted the prototypes' headline
-claims**.
+The full research programme behind forgekit: an account of what a language model with frozen
+weights does not guarantee on its own, an architecture that supplies it, two runnable
+prototypes, and — most importantly — a pre-registered empirical evaluation that **refuted the
+prototypes' headline claims**.
Read in this order. The later work corrects the earlier work, and the corrections are the
-most useful part.
+most useful part. Every load-bearing headline below also has a row, with its status and the
+evidence behind it, in the machine-readable claim registry
+[`docs/status/claims.json`](../docs/status/claims.json), rendered as a table in
+[`docs/status/README.md`](../docs/status/README.md).
## Start here: what is actually true
| | Claimed (self-built demos) | Measured (real data) |
|---|---|---|
| Impact oracle recall | 1.00 | **0.022** — `grep` with no graph beats it ~10× on F1 |
-| Router/gate F1 | 1.00 | **0.37** on 80 real GitHub issues/PRs |
-| Cost saving | +62.1% | **−20.2%** — routing costs *more* than always-premium |
+| Router/gate: gate F1 (should-ask) | 1.00 | **0.37** on 80 real GitHub issues/PRs |
+| Router/gate: cost saving vs always-premium | +62.1% | **−20.2%** — routing costs *more* than always-premium |
-Per output a judge accepted, the router cost $1.06 against always-premium's $1.76, but
-only 6 and 3 of 64 outputs were accepted, so that comparison is not stable; 58 of the 64
-tasks failed at every tier, which is why escalation made routing cost more overall.
+Success in the router rows means **judge-accepted**: a model judge accepted the output. No
+held-out task admitted execution-based verification, so `tests_passed`, `human_accepted` and
+`deployed_without_revert` were never measured. Per judge-accepted output the router cost $1.06
+against always-premium's $1.76, but only 6 and 3 of the 64 non-halted tasks were judge-accepted,
+so that ratio is not stable; 58 of the 64 tasks failed at every tier, which is why escalation made
+routing cost more overall ($6.3582 against $5.2893, 20.21% more). The judge was also the mid-tier
+executor, and the "second labelling pass" is the same model with a reworded prompt (n = 30: halt
+κ 0.5161, tier κ 0.8919), so κ measures self-consistency, not agreement with a human.
+(Corrected 2026-09-26: the first sentence read "Per output a judge accepted, …" and named neither
+what the judge's acceptance is not nor who the judge was; the row was labelled "Router/gate F1".)
-After diagnosing and repairing two defects, with numeric parameters frozen before the
-held-out repositories were touched: recall **0.653**, F1 **0.416**, a point estimate above
-`grep`'s 0.371 for the first time. That the repaired oracle *beats* grep is **not
-established**: the three held-out repositories all favour it, but three out of three is a
-one-sided sign-test p of 0.125, pytest supplies 71% of the held-out pairs, the file-level
-intervals overlap, and the choice of which relations to add was made on all nine
-repositories. (Corrected 2026-09-21; earlier versions called it "a real but narrow win".)
+After diagnosing and repairing two defects, with numeric parameters frozen before the held-out
+repositories were touched: at the pre-registered canonical threshold 0.02, recall **0.653** and
+F1 **0.416**, a point estimate above `grep`'s 0.371 for the first time (paired ΔF1 about +0.044).
+That the repaired oracle *beats* grep is **not established**. Per-repository counts exist only at
+threshold 0.10, where the pooled ΔF1 is +0.0565 and all three held-out repositories favour the
+oracle, but three out of three is a one-sided sign-test p of 0.125; pytest supplies 71.3% of the
+held-out pairs; the file-level intervals overlap; and the choice of which relations to add was made
+after diagnosing all nine repositories — an architecture-selection channel into the nominal test
+set, which limits the unseen-repository claim without erasing the measured gain. Numbers at 0.02
+and 0.10 are never mixed in one comparison. (Corrected 2026-09-21; earlier versions called it "a
+real but narrow win". Thresholds separated 2026-09-26.)
The general lesson, demonstrated on our own work: **a self-built demonstration can overstate
field performance by more than an order of magnitude, and careful caveating does not convert
a demonstration into evidence.**
+## What the impact study measured — and what it did not
+
+*(Added 2026-09-26, after a second external review recomputed the archived results.)*
+
+The archived counts reproduce: 801 labelled files, 759 evaluated after the pre-registered cap of
+200 files per repository, nine repositories, 20,144 mirrored labelled pairs. The original oracle's
+pooled precision / recall / F1 is **0.3982 / 0.0220 / 0.0416**; grep's is **0.3535 / 0.5732 /
+0.4373**. A repository-cluster bootstrap (20,000 draws, seed 1234) gives oracle F1 **[0.0010,
+0.0927]**, grep **[0.3807, 0.5394]**, and grep minus oracle **[0.3422, 0.5174]** — figures
+recomputed by the 2026-09-26 external review and reproduced by
+[`recompute_corrections.py`](recompute_corrections.py) §5, which also prints the repository-level
+view: macro F1 0.0220 for the oracle against 0.4947 for grep, with grep ahead in 9 of 9
+repositories. The negative result is well supported within this archived corpus.
+
+What it is a result *about* needs stating as carefully as the numbers:
+
+- **Co-change is a proxy.** Two files that changed in the same commit are *historically related
+ edits*, not proof of semantic necessity or of test breakage; conversely a dependency graph is not
+ a full co-change graph. The study measured one task: (a) predicting co-edited files. The other
+ task an impact tool is used for — (b) selecting the tests that detect a behaviour regression —
+ was not measured, and results for the two should always be reported separately.
+- **Pairs are not independent.** Every ground-truth pair is mirrored (counted from both ends) and
+ files share repositories, so file- or pair-level resampling overstates precision. Uncertainty is
+ reported at the repository level, and macro results sit beside pooled ones.
+- **The Node graph is not the evaluated oracle.** The shipped `forge impact` / `src/atlas.js` is a
+ regex-derived, multi-language code graph that ports the two repairs; it is not the Python AST
+ oracle the study evaluated, and the study's numbers are not its numbers. Its own measurement is a
+ six-case, self-labelled fixture in [`reports/benchmarks.md`](../reports/benchmarks.md).
+
+### Next study (pre-declared shape)
+
+The nine-repository archive is a reproducibility starter, not a fresh holdout. The next impact
+study freezes the parser and relation design **before** acquiring a new repository set or time
+split; includes runtime coupling, configuration changes, dynamic imports and languages other than
+Python; predeclares relation budgets so that widening predictions cannot win merely by returning
+most files; and reports review-cost metrics — files reviewed per true affected file, and the
+missed-regression rate — beside F1, for co-edited-file prediction and regression-test selection
+separately.
+
## The four layers
### 1. [`cognitive-substrate/`](cognitive-substrate/) — the theory
-The originating argument: an LLM is a frozen map `y = f_θ(x)` with three properties —
-statelessness, frozen parameters, bounded context — which structurally deny it five faculties
-(memory, learning, imagination, self-correction, impact-awareness). The remedy is an external
-stateful architecture, not better prompting.
+The originating argument: an LLM is a map `y = f_θ(x)` with frozen parameters, no state between
+calls and a bounded context, and a coding agent built on it lacks five faculties (memory,
+learning, imagination, self-correction, impact-awareness) unless something outside supplies them.
+Stated precisely, what is missing is a set of guarantees — no durable state across independent
+invocations, a bounded context, no automatic parameter update, and unreliable self-verification
+without external evidence. Prompting does change behaviour inside a context (in-context adaptation;
+Brown et al., 2020, [arXiv:2005.14165](https://arxiv.org/abs/2005.14165)); the substrate is a tested
+way of supplying persistence and verification, not the only logically possible architecture.
+(Corrected 2026-09-26: this paragraph said the three properties "structurally deny it five
+faculties" and that "the remedy is an external stateful architecture, not better prompting".)
-- `cognitive_substrate_whitepaper.pdf` — the *Theory → Evidence → Build-Map* edition (48pp);
- the `.html` edition carries the 2026-09-21 corrections, the PDF predates them
+- `cognitive_substrate_whitepaper.pdf` — the *Theory → Evidence → Build-Map* edition (48pp).
+ **Historical, pre-correction edition** (git blob `44ce7bb`); the `.html` edition is the corrected
+ source and carries the 2026-09-21 and 2026-09-26 corrections.
- `EXECUTIVE_SUMMARY.md` — one-page entry point, **carries a status banner: its prototype numbers are refuted**
- `literature/` — the gap map and 32 graded references behind each faculty claim
- `evidence/` — twelve load-bearing industry statistics independently re-grounded and graded
`confirmed` / `vendor-reported` / `unverifiable`, plus an ecosystem map of what the 2026
Claude-Code stack already solves. Three widely-repeated statistics were caught as
- misattributed and dropped.
+ misattributed and dropped. Since 2026-09-26 the evidence map also grades claim support, study
+ design, independent replication and transfer scope separately; the original grades mainly
+ confirm that a source exists and says what is quoted.
- `quranic-lens/` — the fourteen-mapping ethical-epistemic reading used as a *design lens*:
it names which safeguards are obligatory rather than optional. It is framing, never
- technical authority; no verse is offered as proof of an engineering claim.
+ technical authority; no verse is offered as proof of an engineering claim. The Arabic source
+ text, the translation, tafsir and the author's design analogy are labelled separately, and the
+ lens's operational content is the discipline *do not assert without evidence* — the claim
+ registry, verifier events and visible uncertainty — not any algorithm's correctness, catch rate
+ or uniqueness.
- `sources/` — the primary documents the evidence layer was graded against
- `figures/` — the architecture schematics and prototype evaluations
@@ -62,7 +128,11 @@ two-layer duality: the silent-miss residual is
`(1 − p) × P(no deterministic check fires | miss)`, so where each factor is bounded away from
zero, neither layer alone reaches a small residual. Since the 2026-09-21
corrections this is stated as a bound over an explicit `(p, q)` region, not as a proof that
-neither layer suffices, and the checks multiply only if they fire independently.
+neither layer suffices, and the checks multiply only if they fire independently. Since the
+2026-09-26 corrections the reachable residual is a minimum over the *jointly* feasible `(p, q)`
+pairs — separately maximal `p` and `q` need not be attainable under one policy, so
+`(1 − p_max)(1 − q_max)` is only a lower bound — and a caught miss is no longer read as a
+completed task.
**Priority note:** prior-art review found this composition law is standard protection-layer
algebra, and two concurrent preprints derive a strictly more general Bayesian form weeks
@@ -70,6 +140,9 @@ earlier. Priority is conceded in the refutation paper's related work and, since
2026-09-21 corrections, in the synthesis and the extended preprint as well (before that they
still said "this paper proves"). What survives is that both preprints are simulation-only.
+- `substrate_synthesis.pdf` — **historical, pre-correction edition** (git blob `2e17362`); the
+ corrected source is `substrate_synthesis.html`.
+
### 3. [`empirical-refutation/`](empirical-refutation/) — the measurement
The pre-registered evaluation that overturned the claims above, the diagnosis of *why*, and
the repair. Includes a replication package with the frozen pre-registration, mined ground
@@ -81,10 +154,44 @@ Also corrects a theoretical claim: perfect recall was inferred from a completene
but such a theorem guarantees completeness only *relative to the relation* the closure runs
over — it says nothing about whether that relation contains the edges that matter.
+- `replication_package.tar.gz` — the archive **exactly as published** (git blob `50bd453`); its
+ copies of `paper/main.tex` and `paper.pdf` predate the corrections. Corrected summary: the
+ README's Corrections sections; every corrected number is recomputed from it by
+ `recompute_corrections.py`.
+- `paper.pdf` and `extended_preprint.pdf` — **historical, pre-correction editions** (git blobs
+ `f94a727`, `74f74ae`); the corrected sources are `paper/main.tex` and `extended_preprint.html`.
+
### 4. [`python-prototypes/`](python-prototypes/) — the code
-`impact_oracle/` and `router_gate/`, runnable with their own test suites. The **repaired**
-oracle ships inside the refutation's replication package rather than replacing the version
-here, so swapping it in stays a deliberate decision.
+`impact_oracle/` and `router_gate/`, runnable with their own test suites. The in-tree
+`impact_oracle/` **is the repaired (v2) oracle**: both repairs are in its source, with their
+frozen parameters as module defaults, and its suite is 49 tests (36 demo-package tests plus 13
+regression tests for the two repairs). `ImpactOracle(wm, sibling_enabled=False,
+forward_enabled=False)` reproduces the refuted reverse-only traversal, and the untouched as-shipped
+v1 package is archived in the replication tarball. (Corrected 2026-09-26: this paragraph said the
+repaired oracle "ships inside the refutation's replication package rather than replacing the
+version here, so swapping it in stays a deliberate decision"; the swap had already been made.)
+
+## Prior art, and what is (and is not) claimed
+
+*(Added 2026-09-26.)* External memory, feedback-driven improvement and structured agent control
+all have clear prior art. **CoALA** (Sumers et al., 2023,
+[arXiv:2309.02427](https://arxiv.org/abs/2309.02427)) organises language agents into modular
+memory, action and decision procedures; **Reflexion** (Shinn et al., 2023,
+[arXiv:2303.11366](https://arxiv.org/abs/2303.11366)) improves agents through linguistic feedback
+and an episodic memory buffer, with no weight updates; GPT-3's few-shot evaluation
+([arXiv:2005.14165](https://arxiv.org/abs/2005.14165)) already measured adaptation through text
+alone. That prior art does not make forgekit unoriginal as a product, but it limits what the broad
+architecture can claim. The defensible framing is:
+
+> **a portable implementation of evidence-weighted coding-agent memory and checks, with
+> empirical evaluation of trust failure modes.**
+
+Novelty is claimed only for a specific protocol, invariant, evaluation result or integration that
+survives an explicit comparison with that prior art. The "five faculties" are a useful
+decomposition, not a proof that these five are necessary or that an external stateful architecture
+is the only way to supply them; and the "convergence" of the theory, forgekit and its sibling
+projects is consistency within one author's work, not independent confirmation — the same care the
+programme already applied when it conceded priority for the protection-layer equation.
## How this programme tries to stay honest
@@ -98,6 +205,15 @@ Where this falls short is stated too: the pre-registration and parameter freezes
self-administered with no external timestamping authority, so a reader can verify internal
consistency and the amendment trail but must take the ordering on trust.
+### Four kinds of reproducibility, kept apart
+
+| Kind | Status |
+|---|---|
+| **Source availability** | The papers' corrected sources (HTML, LaTeX), both Python prototypes, the replication archive and the recomputation script are all in this directory. |
+| **Calculation reproducibility** | Available. [`recompute_corrections.py`](recompute_corrections.py) (standard library only) recomputes every corrected statistic from the archived results, and asserts the Theorem D sanity checks with no data at all (`--theorem-checks`); CI runs both, and both prototypes' test suites, since `aedddf5`. The universal router's shipped prior also refits exactly from pinned public data with `bench/universal-router/reproduce.sh` (2026-09-26, the project's own run). |
+| **Pipeline reproducibility** | **Not available.** The historical mining pipeline (cloning, commit filtering, labelling, model calls) is not shipped as an entry point here, so new histories cannot be mined with one command; and the universal router's run-4 held-out benchmark ran in an external harness (harness-bench) that is not shipped either — see [`docs/UNIVERSAL_ROUTING.md`](../docs/UNIVERSAL_ROUTING.md). |
+| **Independent external replication** | None of the research results has been replicated by an independent team. The 2026-09-21 and 2026-09-26 reviews recomputed archived numbers; they did not re-mine repositories or re-run model calls. |
+
## Corrections (2026-09-21)
An external deep review of this repository (2026-09-21) recomputed the research statistics
@@ -126,10 +242,40 @@ mkdir rp && tar -xzf research/empirical-refutation/replication_package.tar.gz -C
python research/recompute_corrections.py rp/repro
```
-**Stale PDFs.** These PDFs predate the corrections and could not be rebuilt here (the HTML
-editions were rendered with WeasyPrint, the paper with a TeX Live toolchain; neither was
-available): `formal-synthesis/substrate_synthesis.pdf`,
+## Corrections (2026-09-26)
+
+A second external deep review (2026-09-26, pinned at commit
+`d2abfa69fb77531199ffc67c5c076b524af69040`) recomputed the archived results again and read the
+papers' framing against the code. The counts reproduced exactly again. Changes, each marked in
+place in its paper with `[corrected 2026-09-26]` and listed there with the original wording:
+
+- **Formal synthesis and extended preprint** — the range statement of Theorem D no longer combines
+ separately maximal `p` and `q` (counterexample: policies `(0.5, 0.9)` and `(0.9, 0.1)` leave 0.05
+ and 0.09, while the separate maxima suggest 0.01); the equality condition for
+ `1 − (1 − ε)ⁿ` is every `rᵢ = ε`, not independence alone; a new §5.4 separates silent-miss
+ probability from completed-task rate and lists what to measure; the frozen-map premise is stated
+ as the guarantees it removes; prior art (CoALA, Reflexion) is named and the byline no longer
+ calls the three bodies of work "independently-developed". The counterexample, the equality
+ condition and the 400× correction are asserted by `python3 research/recompute_corrections.py
+ --theorem-checks`.
+- **Whitepaper** — the "cannot learn / imagine / self-correct" framing marked as broader than the
+ missing guarantees; prior art added to §11; the Qur'anic lens's text, translation, tafsir and
+ design analogy labelled separately, with its operational scope stated; METR's 19% slowdown
+ scoped to its 16 developers, 246 tasks and early-2025 tools, with METR's
+ [February 2026 update](https://metr.org/blog/2026-02-24-uplift-update/).
+- **This README and the prototype READMEs** — the repaired oracle's location, the 0.02 / 0.10
+ thresholds, judge-accepted versus executed success, the unit of the impact study, and the
+ reproducibility table above.
+
+**Stale PDFs — historical, pre-correction editions.** `formal-synthesis/substrate_synthesis.pdf`,
`empirical-refutation/extended_preprint.pdf`, `empirical-refutation/paper.pdf`,
-`cognitive-substrate/cognitive_substrate_whitepaper.pdf`, and the copy in
-`docs/cognitive-substrate/`. The copies of `paper/main.tex` and `paper.pdf` inside
-`replication_package.tar.gz` are left as published. Read the HTML and LaTeX sources.
+`cognitive-substrate/cognitive_substrate_whitepaper.pdf` and its byte-identical copy in
+`docs/cognitive-substrate/` predate both sets of corrections. Read the HTML and LaTeX sources,
+which carry them. A re-render was attempted on 2026-09-26 and **not** committed: the HTML sources
+reference their figures through `{{artifact:…}}` placeholders that a browser cannot resolve, all
+three HTML papers carry Qur'anic Arabic whose typesetting could not be checked by eye in that
+environment, and a PDF whose figures or sacred text cannot be verified is not published as the new
+edition. The paper PDF needs a TeX toolchain that was not available. Each edition's git blob, the
+pinned commit where it stays retrievable, and a faithful render recipe (including the
+figure-placeholder map) are in [`HISTORICAL_EDITIONS.md`](HISTORICAL_EDITIONS.md). The copies of
+`paper/main.tex` and `paper.pdf` inside `replication_package.tar.gz` are left as published.
diff --git a/research/cognitive-substrate/EXECUTIVE_SUMMARY.md b/research/cognitive-substrate/EXECUTIVE_SUMMARY.md
index 0938a0af..f6d12774 100644
--- a/research/cognitive-substrate/EXECUTIVE_SUMMARY.md
+++ b/research/cognitive-substrate/EXECUTIVE_SUMMARY.md
@@ -7,7 +7,7 @@
> | Claim below | Measured on real data |
> |---|---|
> | Impact oracle recall **1.00** | **0.022** (9 OSS repos; 801 labelled files, 759 evaluated); `grep` beats it ~10× on F1 |
-> | Router/gate F1 **1.00**, cost saving **+62.1%** | F1 **0.37**; cost saving **−20.2%** (routing costs *more* than always-premium; per judged-correct output $1.06 vs $1.76, from only 6 and 3 correct outputs of 64) |
+> | Router/gate F1 **1.00**, cost saving **+62.1%** | F1 **0.37**; cost saving **−20.2%** (routing costs *more* than always-premium; per judge-accepted output $1.06 vs $1.76, from only 6 and 3 judge-accepted outputs of 64; the judge is a model, not executed tests) |
>
> The theory sections remain the programme's working framework. The *numbers* here do not. A repair
> raised recall to 0.653 and F1 to 0.416, a point estimate above grep's 0.371, documented in the
@@ -25,6 +25,17 @@
**One-line thesis:** The faculties a coding agent lacks — memory, learning, imagination, self-correction, impact-awareness — are not gaps in the model's *knowledge* but structural consequences of what a frozen transformer *is* (a stateless map `y = f_θ(x)`, fixed weights, bounded window). They cannot be prompted or tooled away; they can only be supplied by **re-wrapping the input→process→output loop** into a closed, stateful cycle around the frozen model.
+> *Corrected 2026-09-26:* the thesis above is broader than its argument. Frozen weights rule out
+> weight updates during use, not all adaptation: examples, retrieved facts and feedback in the
+> context change behaviour with no gradient step (Brown et al., 2020,
+> [arXiv:2005.14165](https://arxiv.org/abs/2005.14165)). What a bare model lacks is a set of
+> guarantees — no durable state across independent invocations, a bounded context, no automatic
+> parameter update, and unreliable self-verification without external evidence — and the substrate
+> is one tested way of supplying persistence and verification, not the only possible architecture.
+> Prior art for the broad architecture includes CoALA ([arXiv:2309.02427](https://arxiv.org/abs/2309.02427))
+> and Reflexion ([arXiv:2303.11366](https://arxiv.org/abs/2303.11366)); see the white paper's
+> Corrections (2026-09-26).
+
**What v2 adds.** The first edition argued the five faculties from first principles and prototyped the one that is buildable today. This edition (1) **grounds the argument in the field's own evidence** — twelve load-bearing pain-point statistics independently re-grounded from primary sources and graded *confirmed / vendor-reported / unverifiable*; (2) adds **six metacognitive mechanisms** the frozen loop also lacks (routing, assumption gate, decomposition, goal-anchoring, anti-over-engineering, inline verification); (3) **maps all eleven capabilities against the real 2026 Claude-Code stack**, marking each solved / partial / residual-gap so we say clearly *what not to build*; and (4) ships a **second runnable prototype** — a complexity-aware router + assumption gate, evaluated live on real models.
> **Governing discipline (the user's, adopted throughout):** *AI output is mathematically-calculated probability — non-deterministic, and never blindly trusted.* Every claim in this package is graded by how well it is sourced; every prototype decision is a transparent, attributable rule rather than another opaque model call; and trust is always earned by an **external** check, never asserted by the model.
diff --git a/research/cognitive-substrate/cognitive_substrate_whitepaper.html b/research/cognitive-substrate/cognitive_substrate_whitepaper.html
index 95460e9b..615495dc 100644
--- a/research/cognitive-substrate/cognitive_substrate_whitepaper.html
+++ b/research/cognitive-substrate/cognitive_substrate_whitepaper.html
@@ -65,6 +65,7 @@
.quran .map{font-size:.92rem; color:var(--ink); margin-top:.6em; padding-top:.6em; border-top:1px dotted #d8ccae;}
.quran .map b{color:var(--quran);}
.quran .grounding{font-size:.74rem; color:var(--faint); margin-top:.5em; font-family:monospace;}
+ .quran .lbl{font-family:sans-serif; font-size:.64rem; text-transform:uppercase; letter-spacing:.08em; color:var(--faint); margin:.55em 0 .1em;}
.lit{background:var(--litbg); border:1px solid #d9e2ea; border-radius:4px; padding:6px 14px; margin:1.1em 0; font-size:.9rem;}
.callout{border:1px solid var(--rule); border-radius:5px; padding:14px 20px; margin:1.4em 0; background:#fcfcfc;}
.callout.key{border-left:3px solid var(--oracle); background:#f2f9f8;}
@@ -96,13 +97,13 @@ A Cognitive Substrate for Coding Agents
This edition was written before any real-repository evaluation existed. A later pre-registered evaluation (research/empirical-refutation/) overturned both prototype claims: the impact oracle’s recall was 0.022, not 1.00, on 759 files in nine open-source repositories, and on 80 held-out tasks the router’s total spend was 20.2% higher than always using the premium tier, not 62.1% lower. An external review (2026-09-21) also found a misquoted statistic, a wrong worst-case cost, and an inconsistency between Eq. (1) and M2. Corrections are made in place, marked [corrected 2026-09-21] or [refuted], and listed with the original wording in Corrections. The theory sections remain the programme’s working framework; the prototype numbers do not. The PDF edition predates these corrections.
This edition was written before any real-repository evaluation existed. A later pre-registered evaluation (research/empirical-refutation/) overturned both prototype claims: the impact oracle’s recall was 0.022, not 1.00, on 759 files in nine open-source repositories, and on 80 held-out tasks the router’s total spend was 20.2% higher than always using the premium tier, not 62.1% lower. An external review (2026-09-21) also found a misquoted statistic, a wrong worst-case cost, and an inconsistency between Eq. (1) and M2. Corrections are made in place, marked [corrected 2026-09-21] or [refuted], and listed with the original wording in Corrections. The theory sections remain the programme’s working framework; the prototype numbers do not. A second review (2026-09-26) found that the “cannot learn / imagine / self‑correct” framing is broader than the missing guarantees it rests on, that the prior art for the architecture needed stating, that the Qur’anic lens mixed source text, translation and the author’s analogy without labels, and that the METR statistic needed its scope; those are listed in Corrections (2026-09-26) and marked [corrected 2026-09-26]. The PDF edition predates both sets of corrections.
A large language model at inference time is, mathematically, a fixed function y = fθ(x) with frozen parameters θ and a bounded input window. From this single fact, five apparent “cognitive” deficits of a coding agent follow as structural consequences, not incidental weaknesses: it cannot remember across sessions, cannot learn from outcomes, cannot imagine the consequences of an action before taking it, cannot reliably correct itself, and does not know what already exists in a codebase or what an edit will affect. We show that neither better prompting nor additional tools (skills, MCP servers) remove these deficits, because they leave fθ and the open‑loop pipeline intact. We then specify a cognitive substrate: an external architecture that keeps the LLM frozen but re‑wraps its input→process→output loop into a closed, stateful cycle over persistent stores — an episodic/semantic memory, an online‑updatable learning layer, a consequence simulator, a metacognitive verification gate, and a persistent structural model of the codebase — all under an explicit stewardship boundary. For each faculty we identify precisely what the existing literature solves and what residual gap remains for a coding agent. To turn the weakest‑evidenced claim into something testable, we build and evaluate the impact‑awareness faculty as a runnable prototype: a Codebase World‑Model that parses a repository into a persistent dependency graph, and an Impact Oracle that predicts the blast radius of a proposed edit. Against mutation‑derived ground truth on a ten‑file package we wrote, the oracle was the only method that missed no affected file (recall = 1.00 across five tested edits), where a text‑search baseline missed transitive dependents and an edited‑file‑only baseline missed 47% of impact. On nine real repositories its recall was 0.022, and text search beat it by an order of magnitude on F1. [refuted — see Corrections] Throughout, a Qur'anic epistemic lens supplies the design's vocabulary of obligation — know what exists before acting (2:31–32), verify before you act (49:6), pursue not that of which you have no knowledge (17:36), and hold what you can damage as a trust (33:72).
+A large language model at inference time is, mathematically, a fixed function y = fθ(x) with frozen parameters θ and a bounded input window. From this single fact, five apparent “cognitive” deficits of a coding agent follow as structural consequences, not incidental weaknesses: it cannot remember across sessions, cannot learn from outcomes, cannot imagine the consequences of an action before taking it, cannot reliably correct itself, and does not know what already exists in a codebase or what an edit will affect. We show that neither better prompting nor additional tools (skills, MCP servers) remove these deficits, because they leave fθ and the open‑loop pipeline intact. [corrected 2026-09-26 — see Corrections] We then specify a cognitive substrate: an external architecture that keeps the LLM frozen but re‑wraps its input→process→output loop into a closed, stateful cycle over persistent stores — an episodic/semantic memory, an online‑updatable learning layer, a consequence simulator, a metacognitive verification gate, and a persistent structural model of the codebase — all under an explicit stewardship boundary. For each faculty we identify precisely what the existing literature solves and what residual gap remains for a coding agent. To turn the weakest‑evidenced claim into something testable, we build and evaluate the impact‑awareness faculty as a runnable prototype: a Codebase World‑Model that parses a repository into a persistent dependency graph, and an Impact Oracle that predicts the blast radius of a proposed edit. Against mutation‑derived ground truth on a ten‑file package we wrote, the oracle was the only method that missed no affected file (recall = 1.00 across five tested edits), where a text‑search baseline missed transitive dependents and an edited‑file‑only baseline missed 47% of impact. On nine real repositories its recall was 0.022, and text search beat it by an order of magnitude on F1. [refuted — see Corrections] Throughout, a Qur'anic epistemic lens supplies the design's vocabulary of obligation — know what exists before acting (2:31–32), verify before you act (49:6), pursue not that of which you have no knowledge (17:36), and hold what you can damage as a trust (33:72).
A better prompt changes x. More tools (skills, MCP servers, function calls) let the agent fetch new x or emit richer y. Both operate inside Equation (1) and leave P1–P3 untouched: the composed system is still a stateless map with frozen weights and a bounded window. A tool call retrieves a document into context, but nothing decides what was worth keeping, consolidates it, or updates the agent's priors for next time. The deficits are properties of the loop shape — open, memoryless, one‑directional — not of the model's knowledge. To remove them you must change the shape of the loop, which is precisely what an external substrate can do while θ stays frozen.
+[corrected 2026-09-26] This callout overstates its case. Prompting and tools do change behaviour: examples, retrieved facts, feedback and extra computation placed in the context adapt a frozen model with no gradient update (Brown et al., 2020). What they do not supply by themselves is four guarantees: durable state across independent invocations, context beyond the window, automatic parameter update from outcomes, and reliable self‑verification without external evidence. The substrate is one tested way of supplying persistence and verification around the model, not the only logically possible architecture.
This reframing is the paper's pivot. If the deficits came from the loop shape, then the remedy is to re‑wrap the loop: keep fθ exactly as it is, and surround it with state and update so that the composite system is no longer memoryless, no longer open, and no longer blind beyond W. Figure 1 states the whole thesis in one picture.
@@ -246,6 +249,8 @@[corrected 2026-09-26] Scope: this is one randomized trial of 16 experienced developers on 246 tasks in repositories they knew, with early‑2025 tools. It is evidence about that setting, not a universal 2026 productivity coefficient in either direction. METR’s February 2026 update (metr.org/blog/2026-02-24-uplift-update) explains why selection effects complicate newer estimates.
+The trend evidence is equally well‑sourced. Stack Overflow's 2025 survey of more than 49,000 developers records trust in AI accuracy falling from 40 % to 29 % even as adoption rose to 84 %, with the top‑ranked frustration — cited by 66 % — being code @@ -312,38 +317,51 @@
The Qur'an is used here as a framing lens and ethics source, never as technical authority for an engineering claim. No verse is cited to prove that an algorithm works or that a data structure is correct — those claims stand on their engineering merits alone (§3, §8). What the lens supplies is threefold: (1) a precise vocabulary of obligation for what an agent that acts on real systems owes — to truthfulness, to verification, to stewardship; (2) a hierarchy of knowledge (‘ilm → fahm → ḥikma: knowledge → understanding → wisdom) that motivates a layered memory architecture rather than a flat vector store; and (3) ethical constraints on autonomy that translate into concrete safeguards. Where a mapping is marked load‑bearing, the concept motivates a specific design decision (e.g. a mandatory, not optional, verification gate); where marked metaphor, it is illustrative. Canonical text below is presented directly and attributed; it is not paraphrased. Arabic and translations were retrieved from quran.ai; the full 14‑row mapping table is in the appendix.
+[corrected 2026-09-26] How to read each card. Four layers are kept apart and labelled: the Arabic source text (clean Uthmani script; an ellipsis marks an abridgement, and the full verse is in quranic-lens/quran_lens.md); the English translation (M.A.S. Abdel Haleem); tafsir (classical commentary, Ibn Kathir), which these cards do not quote — the grounding line only records that it was consulted, and the companion file quotes it under its own label; and the author’s design analogy, which is the author’s engineering reading and neither a translation nor a commentary.
The lens earns its place because the deepest failure modes of an autonomous coding agent are not computational but epistemic and ethical: acting without knowing, trusting a report without checking it, and treating a granted capability as license. The Qur'anic vocabulary names these with unusual precision, and three verses in particular map so directly onto architectural decisions that they shaped the design rather than decorating it.
The remarkable thing is not that these mappings are poetic; it is that they are operational. “Verify before acting” is not a sentiment here — it is a mandatory gate in the action pipeline. “Know the names of things” is not a metaphor — it is a dependency graph. The lens told us which safeguards are non‑negotiable; the engineering told us how to build them.
+[corrected 2026-09-26] The operational link is narrower than that sentence suggests. The lens motivates one discipline — do not assert without evidence — which the repository implements as a claim/status registry (docs/status/claims.json), verifier events and visible uncertainty. It establishes no algorithm’s correctness, catch rate or uniqueness: every technical guarantee still needs code‑level assumptions, tests or measurements, and a deterministic implementation does not make a semantic detector’s catch rate approach 1. No theological adjudication is attempted or implied.
The five faculties of the first edition answer “what cognitive capabilities does a stateless model @@ -754,6 +774,7 @@
In one line: the components are largely borrowed; the loop shape, the validity anchoring, and the coding‑agent target are the contribution. That is a defensible and useful kind of novelty — it is what turns five scattered literatures into one buildable architecture.
+[corrected 2026-09-26] Prior art for the combination itself needs naming too: CoALA (Sumers et al., 2023, arXiv:2309.02427) already organises language agents into modular memory, action and decision procedures, and Reflexion18 improves agents through linguistic feedback and an episodic memory buffer without weight updates. The defensible claim is a portable implementation of evidence‑weighted coding‑agent memory and checks, with empirical evaluation of trust failure modes; the “novel” rows above stand only where an explicit comparison with that prior art survives.
The faculties a coding agent seems to lack — memory, learning, imagination, self‑correction, impact‑awareness — are not deficiencies of knowledge that scale will cure. They are structural consequences of what a frozen transformer is: a stateless map with fixed weights and a bounded window (Eq. 1, P1–P3). Because they follow from the shape of the loop, they cannot be prompted or tooled away; they can only be removed by re‑wrapping the loop into a closed, stateful cycle over persistent stores, with the model left frozen inside it (Eq. 2, Fig. 1–2). We specified that substrate faculty by faculty, said honestly which parts are open research and which are engineering, and — for the one faculty that is buildable today — shipped a running impact oracle that, on a package we built, missed no affected file where the strategies a context‑bounded agent actually uses missed up to half. On real repositories it missed almost everything (recall 0.022), which is the refutation’s subject. [refuted — see Corrections] The Qur'anic lens gave the work its spine of obligation: know what exists before you act, verify what you are told, and hold what you can damage as a trust. Those are not just good engineering defaults; here they are the architecture. The next step is to build the memory and learning layers against the same discipline — anchored to what can be verified, not to what the model says of itself — and to evaluate the whole loop on real repositories with real histories.
+The faculties a coding agent seems to lack — memory, learning, imagination, self‑correction, impact‑awareness — are not deficiencies of knowledge that scale will cure. They are structural consequences of what a frozen transformer is: a stateless map with fixed weights and a bounded window (Eq. 1, P1–P3). Because they follow from the shape of the loop, they cannot be prompted or tooled away; they can only be removed by re‑wrapping the loop into a closed, stateful cycle over persistent stores, with the model left frozen inside it (Eq. 2, Fig. 1–2). [corrected 2026-09-26 — see Corrections] We specified that substrate faculty by faculty, said honestly which parts are open research and which are engineering, and — for the one faculty that is buildable today — shipped a running impact oracle that, on a package we built, missed no affected file where the strategies a context‑bounded agent actually uses missed up to half. On real repositories it missed almost everything (recall 0.022), which is the refutation’s subject. [refuted — see Corrections] The Qur'anic lens gave the work its spine of obligation: know what exists before you act, verify what you are told, and hold what you can damage as a trust. Those are not just good engineering defaults; here they are the architecture. The next step is to build the memory and learning layers against the same discipline — anchored to what can be verified, not to what the model says of itself — and to evaluate the whole loop on real repositories with real histories.
A second external deep review of the forgekit repository (2026-09-26, pinned at commit d2abfa69fb77531199ffc67c5c076b524af69040) asked what this edition’s argument actually establishes. The changes are marked in place with [corrected 2026-09-26]; the argument itself is not rewritten, and each item quotes the wording it qualifies. The PDF edition predates these corrections as well.
evidence/evidence_map.md) now grades bibliographic verification, claim support, study design, independent replication and transfer scope separately.The complete 14‑row mapping (12 load‑bearing, 2 metaphor). Canonical Arabic and translations retrieved from quran.ai; full text, tafsir references, and design principles in the companion artifact quran_lens.json. Caveat: this table is a design lens, not technical or theological authority — see §5.
The complete 14‑row mapping (12 load‑bearing, 2 metaphor). Canonical Arabic and translations retrieved from quran.ai; full text, tafsir references, and design principles in the companion artifact quran_lens.json. Caveat: this table is a design lens, not technical or theological authority — see §5. [corrected 2026-09-26] The second column holds the retrieved translation for verse rows and the author’s own gloss for concept rows (it was headed “Retrieved gloss”); the fourth column is the author’s design analogy.
| Concept / verse | Retrieved gloss | Faculty | Design principle (abbrev.) | Type |
|---|---|---|---|---|
| Concept / verse | Translation (verses) / author’s gloss (concepts) | Faculty | Author’s design analogy (abbrev.) | Type |
| 17:36 — lā taqfu (do not pursue without knowledge) | Do not follow blindly what you do not know to be true: ears, eyes, and heart, you will be questioned about all these. | IMPACT-AWARENESS | Before any code mutation (file write, delete, refactor), the agent must run a pre-action verification gate that checks: (1) what entities in the codebase wil… | load-bearing |
| 49:6 — tabayyun (verify reports before acting) | Believers, if a troublemaker brings you news, check it first, in case you wrong others unwittingly and later regret what you have done, | SELF-CORRECTION | The agent architecture must include a verification gate between receiving information (from context, tool output, or its own prior reasoning) and acting on it | load-bearing |
| grep baseline, held-out | 0.269 | 0.601 | 0.371 |
The repaired oracle's point estimate is above the baseline for the first time, reaching 66.8% of the static ceiling. Earlier versions said it beats the baseline; that is not established. [corrected 2026-09-21] F1 is higher by 0.044, and all three held-out repositories agree in sign, but three out of three gives a one-sided sign-test p of 0.125, pytest supplies 71% of the held-out pairs, and the file-level F1 intervals overlap. The choice of which relations to add also came from failure analysis pooled over all nine repositories, so the split was clean for the numeric parameters but not for that structural choice. Two details are worth more than the headline. First, the +
The repaired oracle's point estimate is above the baseline for the first time, reaching 66.8% of the static ceiling. Earlier versions said it beats the baseline; that is not established. [corrected 2026-09-21] F1 is higher by 0.044 at the canonical threshold 0.02; at threshold 0.10, the only threshold with per-repository counts in the package, the pooled gain is 0.057 and all three held-out repositories agree in sign [corrected 2026-09-26], but three out of three gives a one-sided sign-test p of 0.125, pytest supplies 71% of the held-out pairs, and the file-level F1 intervals overlap. The choice of which relations to add also came from failure analysis pooled over all nine repositories, so the split was clean for the numeric parameters but not for that structural choice. Two details are worth more than the headline. First, the obvious repair of the construction defect is unsafe — it fabricates dependency edges through standard-library name collisions — so we applied a more conservative fix with a smaller gain (11.0× rather than 14.5×); a tool that invents edges to raise recall is worse than one that misses @@ -648,7 +657,7 @@
The gate missed roughly seven in ten under-specified requests. Routing retained partial signal — within-one-tier accuracy of 0.91 is well above chance, so the complexity rubric measures something -— but exact-tier accuracy fell to 0.53, and the cost saving did not merely shrink but inverted: routing does save 59.5% in raw dollars on first attempts alone, but almost none of that cheaper output is correct (3.6% once gated), and counting what the pipeline actually spent escalating up the tier ladder, it costs 20.2% more than always using the premium tier. Labelling noise is real and reported rather than hidden: agreement between two labelling passes on the +— but exact-tier accuracy fell to 0.53, and the cost saving did not merely shrink but inverted: routing does save 59.5% in raw dollars on first attempts alone, but almost none of that cheaper output is judged correct (3.6% once gated on the judge's acceptance; the judge is a model, not an executed test) [corrected 2026-09-26], and counting what the pipeline actually spent escalating up the tier ladder, it costs 20.2% more than always using the premium tier. Labelling noise is real and reported rather than hidden: agreement between two labelling passes on the should-ask label was κ = 0.52, moderate, which bounds how well any gate could score here. Both passes were the same model with differently worded prompts, and that model also judged correctness, so κ measures robustness to prompt wording, not label validity. The cost inversion is also largely mechanical: 58 of 64 tasks failed at every tier and always-premium was judged correct on only 3, so escalation paid for every tier. Per output the judge accepted, the pipeline cost $1.06 against always-premium's $1.76, from 6 and 3 accepted outputs. [corrected 2026-09-21]
A language model that writes code is a fixed probabilistic map, and three efforts that were not independent of one another — one from cognition, one from production failures, one from a shipped codebase — converged on the same remedy: wrap it in an external, stateful architecture that supplies the faculties it structurally lacks. This paper argued they describe one object. The impact-awareness faculty approximates the change-closure fixpoint; the assumption gate is the amnesia equation; and both rest on one result — the residual silent-miss rate is the product of what a probabilistic instruction layer lets through, 1−p, and what a deterministic interception layer lets through, P(no check fires | miss). Where each factor is bounded away from zero, neither layer alone reaches a small residual. [corrected 2026-09-21]
+A language model that writes code is a fixed probabilistic map, and three efforts that were not independent of one another — one from cognition, one from production failures, one from a shipped codebase — converged on the same remedy: wrap it in an external, stateful architecture that supplies the persistence and external checks it does not provide on its own [corrected 2026-09-26]. This paper argued they describe one object. The impact-awareness faculty approximates the change-closure fixpoint; the assumption gate is the amnesia equation; and both rest on one result — the residual silent-miss rate is the product of what a probabilistic instruction layer lets through, 1−p, and what a deterministic interception layer lets through, P(no check fires | miss). Where each factor is bounded away from zero, neither layer alone reaches a small residual. [corrected 2026-09-21]
The honest limits listed below were all stated before any real-repository measurement existed. One more @@ -730,8 +739,21 @@
A second external deep review of the forgekit repository (2026-09-26, pinned at commit d2abfa69fb77531199ffc67c5c076b524af69040) re-checked the corrected Theorem D and this paper's framing. Each change is made in place above, marked [corrected 2026-09-26], and listed here with the original wording, so nothing is silently rewritten. The counterexample in item 1, the equality condition in item 2 and the 400-fold correction in item 4 are asserted by python3 research/recompute_corrections.py --theorem-checks (§3b), which needs no data. The PDF edition predates these corrections as well.
+The synthesis draws in a body of cognitive-architecture and process literature beyond the substrate paper's original 32 references. Each new source was independently verified this pass — modern arXiv sources by direct metadata fetch, classical works by primary-host search or established secondary knowledge — and graded: confirmed (record retrieved, attribution matches), traceable (the work clearly exists and is correctly attributed, but rests on established secondary knowledge rather than a single retrievable record), unverifiable (could not confirm). The tally: 9 confirmed, 6 traceable, 0 unverifiable (15 sources: 8 confirmed by the citations track plus the founding Agent-as-a-Judge paper added on its recommendation). Earlier versions said 8 confirmed. [corrected 2026-09-21]
+The synthesis draws in a body of cognitive-architecture and process literature beyond the substrate paper's original 32 references. Each new source was independently verified this pass — modern arXiv sources by direct metadata fetch, classical works by primary-host search or established secondary knowledge — and graded: confirmed (record retrieved, attribution matches), traceable (the work clearly exists and is correctly attributed, but rests on established secondary knowledge rather than a single retrievable record), unverifiable (could not confirm). The tally: 9 confirmed, 6 traceable, 0 unverifiable (15 sources: 8 confirmed by the citations track plus the founding Agent-as-a-Judge paper added on its recommendation). Earlier versions said 8 confirmed. [corrected 2026-09-21] These grades are bibliographic: they confirm that a record exists and is correctly attributed. They do not grade whether a source supports the claim it is cited for, its study design, independent replication or transfer scope; research/formal-synthesis/graded_reference_set.md now keeps those as separate fields. [corrected 2026-09-26]
| Source | ID | Grade | Note |
|---|---|---|---|
| Cognitive Architectures for Language Agents Theodore R. Sumers, Shunyu Yao, Karthik Narasi, 2023 | 2309.02427 | confirmed | Retrieved via arXiv metadata API; title/authors match claim exactly. Unifies memory, planning/reasoning, action, and learning modules into a single CoALA framework for language agents, giving the cognitive-substrate work's memory/im… |
| Source | ID | Grade | Note |
|---|---|---|---|
| Cognitive Architectures for Language Agents Theodore R. Sumers, Shunyu Yao, Karthik Narasi, 2023 | 2309.02427 | confirmed | Retrieved via arXiv metadata API; title/authors match claim exactly. Unifies memory, planning/reasoning, action, and learning modules into a single CoALA framework for language agents, giving the cognitive-substrate work's memory/im… |