fix: close the 2026-09-27 follow-up review findings (N01–N08) - #169
Merged
Merged
Conversation
N03 - a root test script covers the workspaces only when it is a recognized, unfiltered recursive run of every member's test script whose failure reaches the exit status; node -r, filters and masked runs cover nothing. Coverage is labelled measured, inferred or declared. N01 - exact reuse is byte-exact (key v3: a digest of the spec's code units). The near tier's semantic guard compares code layout and no longer folds Unicode; inline code the ledger would rewrite is refused at mint. N08 - a dependency contract is its whole declaration minus the body (the atlas records definition extents, v5), bound to the module the artifact imports. N02 - consolidation and ledger compaction merge or archive exact duplicates only; near-duplicates are proposed, and only an exact refuted claim drops a lesson. N06/N07 - verifier events carry a derived authenticated flag (none without a key; the v2 MAC covers every field), and outcome provenance is re-derived on every read from authenticated events with a matching verdict. N04 - a context span delivers a definition only whole; cut ones stay pending and are listed under `partial`. N05 - nested repositories and submodules are bound by their own code state (manifest-v3); unbound code turns a PASS into INCOMPLETE unless declared in verify.external. Also: maxPossibleCost is renamed estimatedCostIfAllAttemptsRun; the claim registry records assessed_release and review counterexamples, and bump.mjs stamps unreleased claims; seeded property tests along the semantic boundaries; linear scans replace two backtracking regexes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVVG2VDETWsDxMu6MBWPz2
Re-assessed against 2699aa0 (assessed_release "unreleased" until a release ships it): reuse-exact-identity and phase-p3-reuse-cache (N01, N08), verify-pass-binding (N03, N05), context-completeness and phase-p4-context-assembly (N04), route-outcome-provenance (N06, N07), and the budget contract's renamed cost field. New claim similar-rules-never-merged records the N02 guarantee. Each links its review counterexamples. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVVG2VDETWsDxMu6MBWPz2
…N01-N08 fixes Three adversarial reviews of the N01-N08 fixes found further ways to get a false PASS, a wrong reuse hit, a merged rule or a borrowed verified label. Each is fixed conservatively: when something cannot be established, the result is INCOMPLETE, unknown, proposed or self-reported, never PASS/valid. N03, workspace coverage: - per-tool option allowlist; `--` refused - stateful shell builtins and npm_config_* assignments refused; a hop that passes arguments is not followed - .npmrc/env, lerna, nx and turbo config that narrows the run honoured - turbo needs --force, nx/lerna --skip-nx-cache - package-manager-aware lists; conservative negations; yarn members need a version; nx project config or masked member scripts run on their own - `bun run test` replaces `bun test` - a non-shell script-shell is INCOMPLETE - refused runs reported with the reason N05, nested repositories: - gitlinks from the index are bound; .gitmodules is read by git itself - missing registered paths are unbound - nested repos read under the outer config, their committed .gitignore only - fsmonitor never runs N01, reuse near tier: - keys must be current and stored verbatim - the guard compares symbols with operands, typographic/backtick literals, document order, all whitespace, spelling and format characters N08, dependency contracts: - every import form resolved through the shared resolver (aliases, Python, barrels); whole-module digests (moduleDeps) - export aliases and alias consts followed; a dropped export is a change - decorators, split keywords and regex literals read; lexing on masked text - stale dependency files unknown N04, definition extents: - correct for generics, return-type literals, overloads, Go result types, templates, Ruby `end`, Python strings/continuations, JSX apostrophes - untrusted extents (unbalanced, preprocessor branches) unknown - a stale atlas locates nothing N02, rule identity: - statementKey folds only edge whitespace and one sentence period - fact names, lesson triggers and scope are identity - drop only when every exact claim is refuted - small ledgers still report near-duplicates N06/N07, provenance: - events record their checkout; outcomes record their code state - a run backs an attempt only in its checkout, on the same code - attempt keys cannot alias - non-finite numbers never authenticate Docs (GUIDE, Mintlify, routing, plans), CHANGELOG and the rendered changelog page updated in the same change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVVG2VDETWsDxMu6MBWPz2
The internal adversarial re-review of 2699aa0 found counterexamples against seven claims (reuse identity, the reuse cache, PASS binding, context completeness and assembly, outcome provenance, rule merging); a6be83c repairs them with regression tests. Each claim is re-assessed against a6be83c, its notes name what the re-review found and the scope that still applies, and the generated table is refreshed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVVG2VDETWsDxMu6MBWPz2
…l-clock bound The Windows runner failed "the literal scanner, symbols and spelling are linear on hostile input": 4.0 s against a 4 s bound, about 3× the local time. The bound was machine-dependent, and the run exposed real super-linear work: - literalSpans re-scanned the rest of an unbroken line (`indexOf`) from every ASCII quote; each line's end is now found once; - an unmatched backtick run scanned to the end once per distinct run length; runs are now indexed by length with a forward-only cursor; - semanticConflicts compared differing feature lists with a pairwise `includes` (quadratic in the feature count) and masked each text twice; it now uses Set lookups and analyzes each text once. The test now measures scaling (4× the input must take well under the 16× a quadratic scan would), which holds on a fast machine and a slow CI runner alike, and adds a case with thousands of differing identifiers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVVG2VDETWsDxMu6MBWPz2
CodeWithJuber
marked this pull request as ready for review
September 27, 2026 04:16
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What & why
This fixes the eight findings of the 2026-09-27 follow-up review of v1.7.1 (N01–N08) and applies its suggestions. Three adversarial reviews of those fixes then found further counterexamples in every area; the third commit closes them. The review package's acceptance scripts both pass on this branch:
node reproduce_remaining.mjs <repo> --assert-fixedexits 0: all of N01–N08 are fixed.node original-reproduce.mjs <repo> --assert-fixedstill exits 0: the original F01–F16 stay fixed.Fixes, in the review's priority order:
testscript covered the workspaces whenever-ror--workspacesappeared anywhere, sonode -r ./setup.cjs --testhid a failing workspace behind a PASS.testscript counts: npm, pnpm, yarn, turbo, lerna or nx, directly or throughnpm runhops.!negations included.measured,inferredordeclared.proposed.forge ledger compactfollows the same rule.authenticatedflag, and nothing is authenticated when no evidence key is available.partial.verify.externaldeclares an intentional boundary, and every verifier event records it.Adversarial round 2 (third commit). Every counterexample the three re-reviews reported is closed. Each fix leans conservative: what cannot be established reads as INCOMPLETE, unknown, proposed or self-reported.
--help,--dry-run,--prefi=…,--tag testand anything after--are no longer credited).cd,exit,trap,export…) andnpm_config_*assignments are refused, and argument-passing hops are not followed..npmrc/env, lerna, nx, turbo.--force, nx/lerna--skip-nx-cache.version;bun run testreplacesbun test.script-shellis INCOMPLETE..gitmodulesare bound, and.gitmodulesis read by git's own parser.core.fsmonitornever runs.requireandimport().end, Python strings and continuations, and JSX apostrophes.Suggestions applied:
maxPossibleCostis renamedestimatedCostIfAllAttemptsRun.assessed_releaseand can link reviewcounterexamples.bump.mjsstampsunreleasedclaims with the version it cuts.claims-status --checkfails on a staleunreleased.Docs are updated in the same change: GUIDE, the Mintlify CLI and concept pages, UNIVERSAL_ROUTING, the substrate plan docs, ARCHITECTURE, README and CHANGELOG (with the rendered Mintlify changelog page). The second and fourth commits re-assess the affected claims against the code commits (2699aa0, then a6be83c).
Checklist
npm testpasses: 1,765 tests, 1,761 pass, 0 fail, 4 platform-gated skips (Node 22)npm run checkpasses (Biome lint + format; 0 errors)recursiveTestInvocation,analyzeRecursiveTestRun,reachesWorkspace,scriptShellProblem,shellCommands,isWorkspaceMember,specDigest,depContractv2,indexedText,literalSpans,layoutFeatures,sameStatement,checkoutId,releaseProblems,stampRelease, …)CHANGELOG.mdupdated under## [Unreleased]forge verify,forge context,forge reuse,forge ledger compact,forge route outcomeand the router all changed, and their docs changed with themRisk & rollback
v2:(older ones read as unknown) and new artifact fieldsmoduleDeps/keyVerbatim;manifest-v3(older stamps no longer verify, so re-runforge verify);checkoutfield (v1 events still verify for their smaller scope, but no longer backverify-eventoutcomes);codeState, so outcomes backed by older events read as self-reported untilforge verifyruns again.INCOMPLETE,unknown,proposedor self-reported results, never as a false PASS. Recursive root runs that were credited before may now be refused, for example turbo without--force; the members then run on their own, and the CLI says why.maxPossibleCost→estimatedCostIfAllAttemptsRuninforge route universal --json. A Bun project's root suite now runsbun run test.Extra checks (tick if applicable)
npm run typecheckpasses🤖 Generated with Claude Code
https://claude.ai/code/session_01GVVG2VDETWsDxMu6MBWPz2