docs(PP-066): session 2026-09-06 — SPEC-2.0 (v2.0, 18-row scope, claims ratchet, executable criteria), DAG rows/edges/amendments, decision receipts, roadmap mints, rescope quorum record - #3025
Conversation
… green (ratchet SKIPPED under pmat 3.31.0 at 04:35Z, ARMED and RED under 3.37.0 from 07:53Z, cited from job logs); only agent/G-10 unpushed (PR-A); C0 uncredited; D-2..D-11 recorded on #2873 Pmat-Ticket: PMAT-966
…, PMAT-1062), G-10b (#3013, PMAT-1063), G-10c (#3014, PMAT-1064), U-1 (#3015, pmat#1200, expiry = C0-3 - 6 d), S-0 (#3016, D-3 speed lane on perf-solo); amendments: P-1.1/P-1.2 -> 0.67 (D-11), C0-3 blocked by U-1 (R4), G-10 expiry 10-03 -> 09-12 (PR-A armed); §5.0 re-rendered (96 rows); roadmap mints PMAT-1061..1064 by hand (pmat#1169) Pmat-Ticket: PMAT-966
…n decision receipts (status: complete, each citing its #2873 comment), DEC rows complete, roadmap tickets PMAT-1018..1023 + PMAT-985 completed with proof, §5.0 re-rendered Pmat-Ticket: PMAT-966
…-3/D-8/D-9/D-10/D-11/D-5) completed with proof (the #2873 comments and their decision receipts) Pmat-Ticket: PMAT-966
…11b (#3018), R-8 (#3019); R-0a/R-0b split folded from #3003 (R-0b #3002 PMAT-1060, design-quorum record carried over); edges L0-1/G-11/G-11b <- G-10, every open 0.66 row <- G-11, C0-3 <- U-1, R-0b <- R-0+L0-1, R-2 <- R-0b+D-9, B-G1/R-5 <- R-2, R-6/R-8 <- R-5, R-7 <- R-6, T-0 <- L0-1, TAG-0.66.0 <- every 0.66 row; amendments for the four slack violations the edges exposed (G-10 -> 09-06, R-3 -> 09-19, R-2 -> 10-02, B-G1 -> 10-09); §5.0 re-rendered (100 rows) Pmat-Ticket: PMAT-966
… = 10-03) -> 2026-10-09 for the R-2 -> I-18 edge; check_dag_invariants.sh exit 0 at 100 rows Pmat-Ticket: PMAT-966
…endered Pmat-Ticket: PMAT-1059
…cope; 39 rows -> 0.67 each with cut_by and the claim it protected; L0-1 split into L0-1a (bounded) + L0-1b (unbounded); SPEC-2.0 row; C0-2 (pin); B-G1 folded into R-7; v5 edges; TAG-0.66.0 <- every kept row; invariants exit 0 at 102 rows; §5.0 re-rendered Pmat-Ticket: PMAT-966
…pe and the 39 cut rows with the claim each protected (generated from the DAG), the claims ratchet, the executable criteria table; scripts/release_criteria.sh (C0 C4 C5 C6 C7 C8 C9 C11 C13 C14 one exit-coded command each, C0 first through the analyser pin, never vacuous, 6-row case table; C1 C2 C3 C10 C12 -> 0.67); scripts/run_clean_room.sh (C8 via ../infra beside the main checkout, ENV exit 2 otherwise) Pmat-Ticket: PMAT-966
…scope-fails premise on C3 refuted by the spec: R-0b ships the resolution) — nine claim removals assigned to R-7, C5 -> 0.67 (Q3 unanimous), the ratchet-universe sentence corrected, L0-1a/L0-1b cite #2971 (#3017 closed as duplicate), SPEC-2.0 = #3023; release_criteria.sh credits nine Pmat-Ticket: PMAT-966
…t repeating the literals (the claims ratchet covers docs/audits too — it went RED on its own record) Pmat-Ticket: PMAT-966
…ed from the DAG and the receipts, README counts exact, status doc, kaizen Pmat-Ticket: PMAT-966
|
§13.11 rung 1 — quorum shadow verdict Shadow mode: this records a verdict and merges nothing. A refusal |
…sts on (R-0b ships the resolution) and names C5's residue row (T-0h, 0.67) — driver v5.2 DONE-IF Pmat-Ticket: PMAT-966
… change vs data defect; pre-commit complexity expansion; arms before instrument; checkout restores the index Pmat-Ticket: PMAT-1071
…a criterion's distribution before designing a fallback; a PATH tool is not a pin Pmat-Ticket: PMAT-1071
…ls — the criteria script is a decision surface and check_guards_are_wired.sh refused it unwired Pmat-Ticket: PMAT-1071
…, env arm first, update-branch pull), PMAT-1072 completed (proof:PR#3030) Pmat-Ticket: PMAT-1071
|
[BSE] Your red required check is one re-serialised roadmap entry, and the fix is one command. Measured, not guessed: That is the whole of it. Remedy, on your branch — BSE has not touched and will not touch this PR:
Cause fixed, not the PR: #3048 gives the guard a Receipt: |
…ytes differ only by a stray blank line Two guards disagreed on the same file and the first fix satisfied the wrong one. check_roadmap_diff_additive.sh (G-6) called PMAT-1077 re-serialised, so 5e2bd40 ran roadmap_trim.py, which restores base ORDER — and base order puts PMAT-1074/1077 before PMAT-1069..1073, which check_roadmap_sorted.sh (BSE-09a) then refused: "PMAT-1069 at line 16527 sorts before PMAT-1077 earlier in the file". The actual cause is neither ordering. `classify_pair` compares the entry BLOCK text, and a block runs to the next `- id:`; PMAT-1079 had been appended with a blank line before it, so PMAT-1077's block carried a trailing empty line that main's copy (where PMAT-1077 is last in the file) does not have. One blank line, no field changed. Dropping that blank line makes PMAT-1077's block byte-identical to main's AND keeps the sorted insertion 1069 → 1079. Both guards pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…mmit.sh — the STATE and SESSION-END tools Both have been in use for several sessions as working-tree files and were never committed. That is the "free pass while untracked" trap running the other way: check_shell_lint_ratchet.sh's universe is `find scripts -maxdepth 1 -name '*.sh'` (the working tree, not `git ls-files`), so an untracked script already counts against the ratchet while no reviewer can see it. session_docs_commit.sh was contributing one such error line — bashrs 7.0.1 reads the `do` of a `for` nested inside a single-line `if ...; then ...; fi` as the `if`'s body opener (SC2136) — so the kaizen block is expanded onto its own lines. 177 scripts, 8 error lines, baseline 8, PASS (ratcheted). Refs PMAT-1066, #3018 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
… (operator ruling 2026-09-08) F-1 (PMAT-1080, #3022, P0): `apr chat` silently loaded its built-in Demo model for a sharded SafeTensors index — exit 0, zero tokens, the real model's path printed in the banner. `Path::extension()` returns the last dot-segment, so on `model.safetensors.index.json` (the exact filename `apr pull` writes and recommends) it is Some("json") and matched no arm. Shipped in PR #3050 as a structural fix, not an added arm: `resolve_chat_format` is one decision — suffix before extension, then magic bytes, then a refusal from error.rs — and Demo is not an outcome for a path that exists. F-2 (PMAT-1081, #3024 ask 3, owner bse): the live matrix runs nightly against real models, and the two axes still uncovered — an `apr serve` column and a sharded-GGUF row (merge_gguf_shards, for which no fixture builder exists) — are named rather than dropped. The rows earn their place in a rescoped 18-row release by claim (1): apr reports truthfully and never silently substitutes what the user named. #3022 is the strongest instance of that failure in the tree — the substitution reported SUCCESS — and #3024 is why it survived: qwen-story-daily was green the same night and structurally could not have caught it (`grep -c "apr chat"` is 0, `grep -c index.json` is 0). DAG 103 -> 105 rows, 0.66 lane 32 -> 34; invariants PASS (violations=0); the spec's §5.0 block re-rendered byte-identical; roadmap 824 -> 826, sorted, unique, additive. Refs #3022, #3024, #2873 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…d, status doc, six kaizen lines Minted by hand (pmat#1169, `pmat work add` collides): PMAT-1082 for L0-1b (#2971 — the root cause is named: a crushed Q8_K activation block routes that matmul to the f32-activation kernel, and the CPU reference was the inaccurate side, not the GPU) and PMAT-1083 for SPEC-2.0 (#3023). PMAT-1062 (G-11) flipped to completed — its receipt says complete and PR #3038 is merged. Kaizen, all six from defects met today: * two guards disagreed about roadmap.yaml and the first fix satisfied the wrong one; the cause was neither ordering but a blank line inside the previous entry's compared BLOCK * `. scripts/apr_bin.sh` does not honour CARGO_TARGET_DIR (it reads cargo metadata's target_directory); APR_BIN is the documented pin * evidence need not carry a machine path — invoke the tool relatively and compare the two outputs field by field before claiming they are the same measurement * N identical evidence files are one file plus a sha256 * a gate's RED leg is cheap: build the merge base in a second worktree * a second host can refute a hypothesis class before a lane is dispatched DAG 105 rows, invariants PASS; §5.0 re-rendered byte-identical; roadmap 828 entries, sorted, unique, additive; README counts exact. Refs #2873 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
… scratch, committed by an over-broad `git add -A docs/` It is a flat concatenation of the spec at v1.5, produced by an earlier session as a reading aid and left untracked. Committing it would put a STALE second copy of PP-066-release-spec.md (v1.5 against the tree's v2.1) under docs/specifications/, where the claims ratchet and every doc guard would then have two sources for the same text. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…oc refreshed All five from defects met in this segment: two of my own PRs shipped a falsifier no workflow executed; a new test-selection tier surfaced a test that had been RED on clean main and unrun; an anchored grep meant to fix a pass-grep was itself a false green; the release had two candidate signing-key paths; and a textual guard was answered with an allowlist entry and a reason rather than by rewording the file to dodge its regex. Refs #2873, #3051 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…ines Kaizen line 5 quoted a throughput literal while explaining how a textual guard flags a must-match fixture — and docs/ is exactly the surface the claims ratchet reads, so the line about textual guards was refused by one. De-literalised; the fact is unchanged. Two new lines: the claims ratchet's aperture is the `///` vs `//` boundary and a refactor can cross it (R-0b promoted rationale comments into rustdoc, turning four numbers main already carried into published speed claims; re-baselining a MOVED line is refused by design because the ratchet is set-based); and: run the shared guards across every branch in one pass — three guards over six worktrees took minutes and found four failures CI had not reported yet, each of which would otherwise have cost a serial ~1h cycle. Refs #2873 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…es to 0.67 Operator, 2026-09-08: "skip it and deprioritize". Signing was the only step on the release path that needed a human, and it was gating a release whose three claims do not include provenance. * KEY (PMAT-1079, #3045) -> lane 0.67, first row of the provenance track * R-5 (PMAT-993) and R-6 (PMAT-994) drop KEY from blockers — both unblocked * C13's command drops "+ minisign signature" * 0.66 lane 34 -> 33 rows; invariants PASS; §5.0 re-rendered byte-identical WHAT THIS COSTS, STATED RATHER THAN ABSORBED. Claim (3) keeps its INTEGRITY reading and loses its AUTHENTICITY one. A sha256 still proves the asset matches the manifest the release job produced, and the four host receipts still prove that sha256 is what was tested. It does not prove WHO produced it: the manifest and the assets live in the same GitHub release, so whoever can replace an asset can replace its checksum. So the release must say so, in the three places a user could otherwise infer otherwise: the notes' install section, R-6's installer output at install time, and the vocabulary — no artefact, script, contract or note may call a 0.66 asset "signed" or "verified". Those four obligations are recorded on the decision (#2873) and are what makes the smaller claim honest rather than merely smaller. Refs #2873, #3045 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…26 -> 5 open rows) Operator, 2026-09-08: "lets reduce scope as goal is mainly to fix CUDA issue". 0.66.0 now makes ONE of the mission's three claims — (2): every model in the manifest computes the same function on GPU as on CPU, or the GPU refuses it. Claims (1) and (3) move to 0.67 IN FULL, and the notes will say so rather than leave it inferable. KEPT (12 rows, 7 already complete, 5 open): L0-1b (the fix — a crushed Q8_K activation block routes that matmul to the f32 kernel; 1.5B 0.9508 -> 0.999761 on lambda and 0.950611 -> 0.999583 on gx10, first_divergence none on both), L0-1a (what makes it checkable and the failure honest: derived manifest, C14, the >=64-position rule, the measured threshold, REG-15's refusal instead of a silent downgrade), F-1 (the same class one layer up, already measured and armed), SPEC-2.0, TAG-0.66.0. CUT to 0.67, each row now carrying `cut_by` and the `claim_protected` it was defending: R-0/R-0b/R-2 (registry, apr devices, dogfood-reads-registry — claim 1), R-3 (training banner), R-5/R-6/R-7/R-8 (assets, installer, README-first, nightly install — claim 3), C0-1/C0-2/C0-4 (credit gates: they gate CREDIT, not correctness — I9), G-10b/G-11b (tooling), F-2 (the nightly format matrix). TAG-0.66.0's blockers reduce from 19 to the 6 kept rows. NO release assets and NO installer in 0.66 — `cargo install aprender` is the only supported install, which with D-13 (checksummed, not signed) means claim (3) is not made at all rather than made weakly. scripts/publish_cascade.sh is cherry-picked from agent/R-5; the rest stays behind. DAG invariants PASS (105 rows, 0 violations); §5.0 re-rendered byte-identical. Refs #2873, #2971, #3022 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…G row — paiml-implement was refused on 80% of the epic
`paiml-implement` refuses at Phase 0 (`kind-gate.sh`, AUTO-IMPL-SKILL-001 T-1) unless the
ticket's roadmap entry carries `kind:<code|triage|docs|measurement>`. Attempting to run it
on R-5/PMAT-993 returned exit 2, and it was not one ticket:
PP-066 tickets in docs/roadmaps/roadmap.yaml : 105
carrying a kind: label : 21
NOT carrying one : 84
The label is now DERIVED, not typed, because the DAG already says it: SPEC-/DEC-/TAG- rows
are docs; a row with a command-shaped acceptance or a contract is code. 103 entries written,
derived=103 missing=0 wrong=0, and `check_kind_labels_derived.sh` refuses drift in either
direction (a missing label AND a label that disagrees with its row).
TWO rows are UNDECIDABLE and are REPORTED rather than guessed — refusing to guess is not
refusing the tree, so they do not fail the guard:
G-2 (PMAT-985) "one line in spec §0 with decided_by and date" — correctly prose, a docs row
R-8 (PMAT-1067) "the workflow is green once on all four hosts" — a PROSE acceptance on a
code row, which is the "a prose test: never runs" defect in another costume
MY OWN FIRST DRAFT SHIPPED THE DEFECT THIS GUARD EXISTS TO PREVENT, and it is worth the
record: it read the roadmap by line regex in both directions. `--update` then CORRUPTED the
file — 284 entries carry the inline `labels: []`, which has no ` - ` block to scan, so the
insert landed after a flow sequence and the YAML stopped loading — and `--check` reported
`missing=0 wrong=0` on the wreckage, because a regex reader cannot see a parse error. A
writer that can break the file its own checker reads is exactly the class this guard is for.
Fixed both ends: the verdict now PARSES the roadmap, the writer handles the inline form, and
it refuses to write a rewrite that does not parse or that loses an entry. The case-table
fixture carries both label forms so the corruption cannot come back.
Case table 7/7, both polarities. No ci.yml edit needed: guard_tree.sh's universe is
`git ls-files 'scripts/check_*.sh'`, and the guard advertises `--self-test` in its usage, so
the dispatcher gives it both rows. check_guards_are_wired PASS, roadmap sorted/unique/additive
PASS (828 entries, 0 non-label fields changed).
Refs #2873, PMAT-1093
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
PMAT-1093 was cited by the kind-label commit and never existed — a Refs pointing at nothing, which is the small version of the defect pmat#1240 describes. Minted retroactively as `completed` with its acceptance command and `proof:` path. PMAT-1094 (#3055): the `*-lint --json` outcome surface emits three shapes for one field, one of them a Rust `Debug` string in a JSON API. Found sweeping #3051; the tests in #3053 accept all three deliberately and document the table, so this ticket's falsifier is the DELETION of `assert_outcome_ok`'s two string arms. Both by hand: `pmat work add` mints colliding ids (pmat#1169), and now pmat#1240 — two agents minting in parallel branches land on the same id and the merge deletes one, with `work validate` passing because uniqueness is preserved by the loss. roadmap 828 -> 830, sorted, unique; kind labels derived=103 missing=0 wrong=0. Refs #2873, #3055 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…ts credited set is derived
`scripts/release_criteria.sh --all` is step 1 of the release sequence. It could not produce
a verdict: measured 2026-09-08, C0, C4, C7 and C8 each hit a 120 s timeout and --all ran
past ten minutes with nothing printed. Three defects, one of them mine to have noticed
sooner.
1. C0 RE-RAN FOR EVERY CRITERION. `run_one` gated each criterion on `bash "$0" C0` (I9's
credited-first rule), and C0 shells out to `pmat comply check` AND a `gh api` call. Nine
criteria therefore paid that cost nine times. It is now evaluated ONCE per process and
only while C0 is itself in the credited set. C7 alone runs in 57 s; through the old gate
it timed out at 120 s.
2. THE CREDITED SET WAS STALE AFTER D-14 and is now DERIVED, not chosen: a criterion is
credited iff at least one row the spec's §4 table names as its owner is still lane 0.66.
C7 SPEC-2.0 KEEP the claims ratchet — load-bearing for D-13/D-14's vocabulary
C8 SPEC-2.0 KEEP clean-room before publish; cargo publish rests on it
C9 C0-7 KEEP every credited row has a complete receipt
C14 L0-1a KEEP GPU = CPU per manifest model, or the GPU refuses it — 0.66's ONE claim
C0 C0-1/2/4 MOVE all 0.67; keeping it made --all both unpassable and unrunnable
C4 R-6,R-2 MOVE four host receipts THROUGH the R-6 installer, which 0.66 does not ship
C6 G-10a… MOVE two owners are not rows at all; its script does not exist (ENV 2)
C11 R-0a/0b MOVE the backend registry is 0.67
C13 KEY,R-5,R-6 MOVE no assets and no installer in 0.66 (D-13, D-14)
3. THE SUCCESS BANNER CARRIED A SECOND, HAND-TYPED COPY of the list, so editing the set
would have left it asserting the old one. It prints $CREDITED now.
Two new self-test rows, both mutation-proven: hand-editing CREDITED to add C13 turns the
derivation row RED (6/7), restoring it returns 7/7; and the banner row greps the source for
a literal `ALL CREDITED (C…` and refuses one.
CORRECTING MY OWN EARLIER READING: not all of the ten minutes was a defect. C8 is
`make -C machines/clean-room clean-room-p1`, a real multi-minute docker build, and it is
SUPPOSED to be slow — it is the hard gate before publish. Worth naming separately: the
release sequence runs `release_criteria.sh --all` and then `run_clean_room.sh`, so the
clean room currently builds twice.
Current state on this branch: C7 CREDITED (57 s), C9 CREDITED (0 s), C14 ENV 2 —
`scripts/check_model_parity.sh` lives on agent/L0-1, which is in the merge queue — and C8
is the long build. C14's ENV is the correct answer, never a pass.
Refs #2873, PMAT-1083
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…on of both sides #3026 merged, so main now carries the three L0-1a manifest steps and this branch carries the release_criteria self-test step. Both are additive steps in the same guard-tree job and neither replaces the other; the resolution is the union, verified by counting each step rather than by eye. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…66's single claim, measured L0-1a landed (#3026), so scripts/check_model_parity.sh is on main and C14 stopped being ENV. Run verbatim as the criterion defines it, on lambda, with a cuda apr built from agent/L0-1b (sha256 776cbbdb4306b5d8 — the tree carrying L0-1b's fix): PASS qwen2.5-coder-1.5b-instruct: 78 positions, min cosine 0.9998 at position 36 PASS qwen2.5-coder-0.5b-instruct: 78 positions, min cosine 0.9996 at position 0 PASS qwen2.5-coder-7b-instruct: 78 positions, min cosine 0.9996 at position 22 C14: measured=3 rc=0 The first row is the model #2971 is about. It read 0.950827 on this host before the fix and reads 0.9998 after, against a threshold whose basis is two measured known-good pairs across two GPU architectures. Fourteen models are UNMEASURED because this host does not hold them, and the script REPORTS rather than fails — no single host holds every model the README names, and the fleet-level rule belongs to the release. From the outside, `apr chat --gpu` on the 1.5B now prints `[GGUF CUDA: NVIDIA GeForce RTX 4090 …]` and answers correctly, with no `falling back to CPU` anywhere in the output. ONE GAP, NAMED RATHER THAN GLOSSED: the driver's Resolved criterion asks for `selected: cuda … parity: PASS` on the success path, and that line does not appear — REG-15's admission line is printed by `on_cuda_load_error`, so it is emitted only when something goes wrong. Reporting the selection only on failure is weaker than the criterion asks. That belongs to claim (1), which D-14 moved to 0.67 with R-0b; 0.66 makes claim (2), and claim (2) is what the table above measures. Refs #2971, #2873 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
… folded into the one with 11 falsifiers Epic #3058 §B3: do not create contracts/apr-gpu-cpu-parity-v1.yaml; extend contracts/apr-cpu-vs-gpu-output-parity-v1.yaml, which already exists with 11 falsification tests. L0-1a shipped the duplicate anyway (#3026, merged), so this removes it. Verified before acting, not taken on faith: the existing contract carries 11 FALSIFY ids and 58 KB of history; mine carried 3 obligations and 3 falsifiers in 10 KB. THE TWO ARE THE SAME ARGUMENT SPLIT IN HALF, which is the real cost B3 names. The existing contract already records the #1864 five-whys whose root cause is "the gate's domain was too narrow (single-step instead of multi-step)" — and L0-1a's >= 64-position horizon rule is the answer to exactly that. Half the reasoning sat in each file. Ported as FALSIFY-CPU-GPU-012/013/014, each with the mutation that turns it RED: 012 the manifest is derived, never typed derive_model_manifest.sh --self-test (6/6) 013 the domain is >= 64 positions, both check_model_parity.sh --self-test (9/9) polarities on BOTH required GPU hosts 014 a forced backend never downgrades cargo test --test reg15_admission (7/7) All three re-run here, green. `pv validate`: 0 errors, 14 falsifiers. THE STALE ANCHOR, AND A SHARPER VERSION OF B3's POINT. B3 is right that the contract cites `mod.rs:268-279` for the SKIP_PARITY_GATE bypass and that :268-279 is something else (a doc comment about qtype resolution). But B3's proposed replacement, `:333`/`:349`, had ALREADY DRIFTED by the time I read it — L0-1a and L0-1b moved the bypass to :390/:406 on this branch. The fix for a drifting line anchor cannot be a different line anchor, so both live citations now anchor on the SYMBOL (`grep for the literal SKIP_PARITY_GATE`) and say why. The 1.1.0 changelog entry keeps its line numbers: it is history, and history is allowed to be stale. References repointed in the DAG (4), the spec table, parity_admission.rs, reg15_admission.rs and .pr/L0-1/accept.sh. DAG invariants PASS; §5.0 re-rendered byte-identical. Refs #3058, #2971, #2873 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
…cells, and one result stronger than the gate asks for Epic #3058 §F1: run it by hand on lambda-labs and gx10, needing none of the missing automation. Four cells for the 7B — {lambda-labs, gx10} x {cpu, gpu} — as §5.5 receipts under docs/audits/release/v0.66.0-pre/models/. M3 parity over >= 64 positions GREEN on both GPU cells: 78 positions, 283 op rows, first_divergence None; lm_head 0.99958 (lambda) and 0.99978 (gx10); worst op ffn_out@L23 0.994065 and attention@L22 0.996844 M4 determinism GREEN, all four cells M6 7B service smoke GREEN, all four cells — loads, non-empty output, zero OOM THE RESULT WORTH MORE THAN ANY SINGLE CELL: the greedy stream sha256 is 6f38f13e5debd285 in ALL FOUR cells. Same model bytes, two architectures (x86_64 sm_89 and aarch64 sm_121) and both backends produce a byte-identical token stream. M4 only asks for determinism WITHIN a cell; this is determinism across the fleet. MY FIRST M4 RUN REPORTED NO, AND MY INSTRUMENT WAS WRONG. It compared whole stdout, whose last line is `Completed in 12.63s` — a clock, not a nondeterminism. The generated text was byte-identical both times. The comparator now strips that line, which is the same rule this repo already enforces on required gates: no wall-clock assertion inside a correctness check. EXCLUSIONS ARE NAMED, NEVER SKIPPED, because §5.5 says a missing cell is NO-GO: M1 artifact identity 0.66 ships no assets and no manifest (D-13, D-14) — nothing to compare M2 registry readback `apr devices` is R-0a, 0.67 M5 refusal semantics PARTIAL — the FX fixtures are R-0a; the admission level is covered by reg15_admission 7/7, now FALSIFY-CPU-GPU-014 in the parity contract M7 transport parity NOT RUN — needs a port and a client; named, not dropped M8 performance NOT RECORDED DELIBERATELY — the claims ratchet refuses a throughput literal on a documented surface and 0.66 makes no speed claim VERDICT WITHHELD, on purpose. Each receipt says "not GO and not NO-GO": this is a dry run on a pre-tag directory, and a receipt calling itself GO with M1 and M2 structurally excluded would be the theater §1.1 forbids. Refs #3058, #2971, #2873 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018RouwmUL7vFfJyCx9qLEoH
check_roadmap_completion_is_cited.sh went RED on this branch: PMAT-1062 said `status: completed` with `github_issue: null` and `notes: null`, so nothing could dereference the claim. Checked before citing rather than after, because the guard's own precedent (PERF-004) is an entry marked completed whose PR was closed unmerged. All three things the title names are on origin/main: guard job runs all, reports all guard_tree.sh + the guard-tree job, #3037 (git log -S' guard-tree:' names ff122fa) aprender timeout-minutes #3037 added timeout-minutes 90/90/22 CARGO_TARGET_DIR per container all six container steps carry -e CARGO_TARGET_DIR=/workspace/target So `completed` is honest and the fix is the citation, not the status. The edit is applied through a YAML load/verify rather than a regex rewrite -- my own kind-label guard corrupted this file once by treating it as text, and reported missing=0 wrong=0 on the wreckage. Entry count asserted unchanged (830) and the citation asserted present after the write. Verified: check_roadmap_completion_is_cited.sh PASS; guard_tree.sh --no-cargo 41 checks, 0 failed. Refs #3025 Pmat-Ticket: PMAT-1083
…licate parity contract guard-cargo's "README claims must match measurement" was RED: the README said 1817 (main's count) while this branch's tree carries 1816. The one-file delta is contracts/apr-gpu-cpu-parity-v1.yaml, deleted here per epic #3058 §B3 — the instruction was to extend the contract that already had 11 falsifiers rather than mint a second one over the same property, and the duplicate had already merged on L0-1a. Re-derived with `make readme-sync`, which reads `find contracts/ -name "*.yaml"` and rewrote both CONTRACT_COUNT blocks. Not hand-edited. Refs #3025 Pmat-Ticket: PMAT-1083
…n the merged tree (1817); #3025 landed under it Pmat-Ticket: PMAT-1080 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjhtNUSensCYpQb3mCYLod
…: roadmap/estimates unions, session_docs_commit.sh DET002+SEC010 fixed, the agent-memory ignore rule withdrawn (main tracks it), receipt phase 3c Pmat-Ticket: PMAT-1096 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BjhtNUSensCYpQb3mCYLod
PP-066 orchestrator docs PR — session of 2026-09-06 (driver v3 → v5.1). Branch
agent/pp-066-spec(orchestrator: the only branch that writes the DAG, the spec block, the roadmap and the README counts under G-11a's rule). Epic #2873.What lands (14 commits, every one verified by the guards it touches):
cut_by/claim_protected), the claims ratchet, the executable criteria table;scripts/release_criteria.sh(nine credited criteria C0 C4 C6 C7 C8 C9 C11 C13 C14, one exit-coded command each, C0 first through the analyser pin, never vacuous — 6-row case table; C1 C2 C3 C5 C10 C12 → 0.67);scripts/run_clean_room.sh(C8 via../infra, ENV exit 2 otherwise); the rescope quorum recorddocs/audits/pp-066-rescope-quorum.md(scope-holds-with-changes; nine claim removals assigned to R-7; C5 → 0.67 by unanimous Q3).check_dag_invariants.shexit 0,render_dag.py --checkbyte-identical): rows L0-1a/L0-1b (GPU inference refuses Qwen2.5-1.5B-Instruct GGUF (hidden=1536/heads=12/kv_heads=2): parity gate fails at cosine 0.94, CPU works fine #2971), G-11 (PP-066 G-11: shared-file write contention — row PRs never write the DAG, roadmap, spec block or README counts; DAG status derived from receipts #3012), G-11b (PP-066 G-11b: one-call state (scripts/pp066_state.sh), the session docs commit (scripts/session_docs_commit.sh), .pr/<row>/accept.sh convention, make fleet-verify ROW=<row> #3018), G-10b (PP-066 G-10b: analyser pin guard — every pmat reference under scripts/ and workflows resolves through scripts/pmat_bin.sh (baseline 281, shrink-only) #3013), G-10c (PP-066 G-10c: analyser reference sweep — the 281 unpinned pmat references resolve through scripts/pmat_bin.sh; baseline to 0 #3014), U-1 (PP-066 U-1: pmat#1200 — work cot derive emits hollow obligations; C0-3 is blocked until the fixed pmat is on the fleet #3015), S-0 (PP-066 S-0: the speed lane runs on perf-solo and its receipt names its own host (D-3; #2720 direction, re-derived from HEAD) #3016), R-8 (PP-066 R-8: nightly-latest-dogfood — install from latest on four hosts, apr devices --json, C14; drift files an issue #3019), SPEC-2.0 (PP-066 SPEC-2.0: spec v2.0 — three claims, 18-row scope, claims ratchet, executable release criteria #3023); R-0a/R-0b split folded from docs(PP-066): R-0 design quorum (3/3 implement-with-changes) — split into R-0a registry + R-0b resolution (#3002, PMAT-1060), expiries moved under §12, spec §12 llamafile citation corrected #3003; the v5 edges; amendments for every slack move; G-10 and the seven decision rows complete.roadmap_trim.py): PMAT-1060..1069 minted by hand (pmat#1169), PMAT-1059 and the decision tickets completed with proof.docs/audits/pp-066-status-2026-09-06.md(fromscripts/pp066_state.sh),docs/audits/driver-kaizen.md(15 lines),docs/audits/impl-estimates.jsonl(+1).Guards run on the branch:
check_dag_invariants.sh·render_dag.py --check·check_receipt_complete.sh --dag·check_roadmap_diff_additive.sh·check_no_claim_literals.sh·check_perf_claims_cite_receipts.sh·release_criteria.sh --self-test·check_shell_lint_ratchet.sh— all PASS.Nothing here is evidence for its own merge (I2):
ci / gateandworkspace-testdecide.