WARS: Bee + TRI arm, a real paired campaign, a readable arena; ROADMAP level II; blog posts - #1184
Merged
Merged
Conversation
…aign Six arms on three real gHashTag/t27 issues (#4614, #4695, #4613), each in its own worktree at afe2186c, judged by the issues' own acceptance commands with t27c 0.4.0. Per run: the patch, the judge transcript, a fixed-mutant review of the new test, and the arm's own report. The campaign folder holds the issue texts the arms read, the judge and mutation scripts, and the judge's run on the unmodified base (which fails all three, as the issues say it should). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
… arena The protocol's variable factor has been the TRI decision layer since #1178, but no arm used it. bee-tri is that arm: the same Bee, prompt, tools and budget, with t27c on PATH while it works. A judged or complete experiment now needs a sealed bee-baseline and bee-tri run; JEV stays a comparison arm. The ledger gains three real gHashTag/t27 issues (4614, 4695, 4613), six sealed runs and 48 measurements: acceptance by the issues' own commands, Queen verdict, elapsed, tool calls, patch lines, and three new metrics (total-tokens, judge-calls, mutants-killed). All six arms pass; each pair kills the same mutants; the TRI arms found three tool defects. A new experiment state, "judged", records that every arm ran without declaring a winner, because the executor model id is not written to this repository. A test shows a judged experiment without a sealed TRI run is refused, and that promoting one to complete without a model id is refused. The view: a campaign-at-a-glance strip, a head-to-head grid of every experiment (a row opens it), a scoreboard that shows only measured metrics with a relative bar per row and lists the rest as not measured, and five lanes. Checked: check:wars, test:wars-spec (26/26), check:aria, check:queen-languages, check:queen-responsive, the typecheck ratchet (no file gained errors) and the viewport contract, whose only failure (.queen27-sectors-empty) fails identically on main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
…were not checked New post (EN + RU, published: true). The AX7203 computed 28,416 of 28,416 rows of tern_tc's 320-input matrices bit-exact, and the harness reported 284,160 authenticated receipts without comparing a tag. The post withdraws that count until the board reruns, shows the fixed harness failing on seven kinds of misbehaving cell, reports the RTL co-simulation (layer 0, all seven matrices, 67,200/67,200 receipts; w_down with int8 activations, 51,840/51,840) and why w_down, wk, wv and int8 activations need no new hardware. Receipts link to gHashTag/trinity-fpga#800 and pinned files. "A small agent needs an exact judge": the open question "No IGLA model has run on a board" and "the 13M model has been neither trained nor run" are updated with dated notes, and the summary's 0.97% is corrected to 1.01%, the mean of the post's own MultiPL-E table. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
check:queen-contrast failed on CI: .queen-wars-glance tiles lay a dark translucent ground over the live hex lattice with nothing taking the lattice's edges out of it. Add backdrop-filter: blur(6px), as the contract asks. Every command in website-checks.yml then passes locally except check:queen-viewport, whose .queen27-sectors-empty failure is identical on main in this environment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
Owner
Author
|
CI status on this PR. One failure was this PR's and is fixed in
The credential failures need someone with access to repository secrets. They cannot be fixed from this PR. Apart from those, the PR is waiting on review. Generated by Claude Code |
The post now reports what the fixed harness did on the AX7203 on 2026-09-27. Layer 5 w_down with int8 activations passed at 8 jobs in flight: 51,840/51,840 receipts verified and 320/320 rows bit-exact. The random matvec passed as well. The two long runs at 64 in flight stopped when the UART link lost 16 and 4 answer bytes. No wrong answer or bad tag came before either hole. The open questions, receipts and next steps (EN and RU) now say that the full 42-matrix run is pending at 24 in flight, and that the cause of the lost bytes is not settled. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
The fixed harness ran all 42 ternary matrices of the trained tern_tc on the AX7203 at 24 jobs in flight. It verified 403,200 of 403,200 receipts, 33,792 of 33,792 rows were bit-exact, and the run took 85.5 s at 4,716 answers/s. The post now leads with that result. The title changes to say so; the slug stays the same. The lost-bytes analysis now cites the measured 22.4 ms host pause, which at 64 jobs in flight is enough to fill the whole queue. The open questions keep what is still missing: the window-64 control on the new harness, the real activations, a forward pass and power. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
The control ran the same full job at 64 jobs in flight on the same harness, setup and session, and lost 59 bytes after 6,744 jobs. So the window explains the pass at 24, not the harness change. Its longest host pause was 4.2 ms, about 1,700 answers before the hole, which is too short to fill a 64-deep queue. That rules out the explanation this post gave an hour ago, that answers pile up while the host looks away. The post now puts the loss in the link itself (the USB adapter, its driver, the hub or the cable), not yet isolated, keeps 24 in flight as the operating point, and adds the control's row and log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
… output A new post (EN and RU) on the pre-registered GFTernary x Z[phi] run on the AX7203 node. The TNF paper's t*phi weights, applied to Z[phi] activations, ran on the unchanged cell with no multiplier. Each output is W.b + (W.a + W.b)*phi: the node computes only ternary dots and the host adds once per output. Layer 0 checked 403,200 of 403,200 receipts, and 5,632 of 5,632 Z[phi] rows matched a plain Z[phi] multiplication oracle bit for bit. The run was registered 38 s before it started. The post ties this to the earlier finding that phi in GFTernary is a scale, not information. It also says what the run does not cover: no TNF accumulator or rounding, one synthetic activation vector, layer 0 only, no speed claim, and receipts that are not publicly verifiable. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
A test was registered before it ran on the owner's board and then run. At 26 jobs in flight (494 bytes of answers) all 403,200 jobs came back clean. At 30 (570 bytes) the stream lost 7 bytes after 150,017 jobs. That puts the loss threshold around the CP2102N bridge's 512-byte receive buffer. The bridge's datasheet asks for handshaking above 1 Mbaud, and the node has none. The post now explains the link failures that way and keeps the side prediction that missed. It adds both runs to the table, updates the counts (185,535 verified answers before four holes), and changes the next step to flow control. EN and RU. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
…t tasks up The ROADMAP tab measured the rewrite and never showed anyone playing it. It now opens with the game board: - the comb: an inverted pyramid in the shape of the t27 mark. The apex is the seed (t27c); every other cell is one port task, laid from the apex upward quickest first (fewest functions), so the comb grows from its point; the widest row is the whole stack in .t27, the BrowserOS browser and every dependency included - ships: one over each cell the Queen's public board reports running, with a beam into it; nothing is drawn that the board did not report, and a board that does not answer is named on the page - priority targets: the quickest free port tasks, with a link into each - sectors: every stage with its state (captured, under attack, open front, no tasks yet, locked and why) Live and public: GitHub search for the port tasks (open, and closed as completed) and /queen/public-board. Without the search every count reads as unknown, not zero. goals.json gains `locked` reasons for stages 3, 4 and 8, and stage 9, the endgame: the whole BrowserOS browser and every third-party dependency, not measured yet. The contrast gate registers the new sheet, its token scope and its four surfaces, sealed by .rm like the tab's own panels. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
… round LEVEL II gains the rules of a game, each computed from public facts and held by qa/roadmap-game-contract.mjs: - the raid of the day: one sector per UTC day over stages 1, 2, 5, 6 and 7; gHashTag/t27's roadmap feeder feeds the same sector first, and both sides pin the list, so the board never names a raid nobody runs - honey: the functions ported in built cells, raid cells closed today counted twice; named as this board's score, not the Queen's XP - cracked cells: a built port file with an open defect that names it - the Queen's round: a band of light climbing the comb, timed by the board's own pulse, with the next round counted down - bosses: stage 8 and the endgame, with their measured source as HP, the board's opening rule, and what really stands in the way - builders: lenders whose lanes built .t27 cells, from the Queen's public leaderboard The rules live in src/lib/roadmapGame.ts; website-checks runs the contract, and the contrast gate seals the raid banner and boss cards. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
The endgame had no goal issue, so its row showed none. It now links gHashTag/t27#4858, and its text carries what was measured on 2026-09-27: BrowserOS dev is 5.48 MB of source, and the running services lock 2,187 npm packages, 650 crates and 35 Go modules. The engine, Chromium 146, is fetched at build time and is still not measured, which is what the lock now says. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
Stage 9 had no HP: the browser is not in any repository the count reads. BrowserOS pins Chromium 146.0.7680.31 (its BASE_COMMIT 4d3225104176d is that tag's commit). Chromium's src at that tag, plus the V8 and Skia commits it pins, were read from the GitHub mirrors and counted by the roadmap-stack.mjs rules, without Config, Docker and Make: - Chromium src @4d32251: 1,144.7 MB (C++ 676.6, C 173.4, Java 101.6) - V8 14.6.202.6 @0a35ee1: 145.6 MB - Skia @9022820: 51.4 MB Together that is 1,341.7 MB, seven times the 192.2 MB the count holds for the whole stack. The rules skip third_party/, so Blink is not in that figure (132.4 MB outside its tests). Neither are the 258 other DEPS repositories. goals.json gains `measured` for stage 9 with the bytes, the date and the commits. A stage with a measurement uses it alone, so the boss bar and the stage row show it without counting stage 2's agent-server twice. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
…field The comb was a flat SVG while the Queen's field is a Babylon.js wall. It is now drawn the same way by default (queenRoadmapScene.ts), from the same cells in the same order: - every cell is a hex prism on a wall facing the player, as tall as its file has come: built honey 1.7, review 1.0, a bee building 0.8, held 0.35, a free target a thin plate in its language's colour, the seed t27c the tallest and lit; rims glow, cracked cells keep their crack; - a ship with a beam hovers over each cell the Queen's board reports running, and the round is a band of light climbing the wall on the board's pulse (the same pulsePhase as the flat comb); - ACES tone mapping and a glow layer as on the field; the pointer lifts the cell under it and shows a card, a click opens the issue, a tap pins the card with a link; a drag tilts the wall, the wheel scrolls the page; - it renders on demand, at 30 FPS only while something moves and the canvas is on screen, and stays still under prefers-reduced-motion. Without WebGL the flat comb is drawn with the reason; a Flat button (or the field's ?engine=canvas) keeps the SVG, remembered per viewer. Babylon loads only in the 3D path, in a chunk of its own (18 KB over the field's core). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
Every claude-review run failed two seconds in: the OAuth token was refused and ANTHROPIC_API_KEY was empty. The owner asked for z.ai. z.ai serves the Anthropic Messages API at https://api.z.ai/api/anthropic, and Claude Code reaches it through ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN. The key comes from a ZAI_API_KEY secret. The model is glm-4.6, the z.ai id in the Queen's own model catalogue, unless a ZAI_MODEL repository variable names another; Haiku-sized calls go to glm-4.5-air. Without the secret (a fork, or before it is added) the job says so in a notice and skips the review instead of failing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
gHashTag
marked this pull request as ready for review
September 27, 2026 11:43
github-actions Bot
added a commit
that referenced
this pull request
Sep 27, 2026
Merge pull request #1184 from gHashTag/claude/peaceful-noether-dperdx WARS: Bee + TRI arm, a real paired campaign, a readable arena; ROADMAP level II in 3D with the measured endgame; two blog posts; the PR review on z.ai Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
This PR makes five changes.
bee-tri: the same Bee, prompt, tools and budget, with t27c 0.4.0 available while it works.gHashTag/t27issues (Test the 1 untested function in specs/boards/arty_a7.t27 t27#4614, Test the 1 untested function in specs/fpga/testbench/simulator_tb.t27 t27#4695, Test the 1 untested function in specs/base/ternary_encoding.t27 t27#4613). Each issue got two arms, and each arm ran in its own worktree atafe2186c. The issues' own acceptance commands judged every arm.7d3a47aunderpublic/queen/runs/.gHashTag/trinity-fpga, fixed in fix(conformance): check tern_tc receipts, and cover every ternary matrix trinity-fpga#800.8a20321).aa35629). The tab measured the goal but never showed anyone playing it. It now opens with:73bca5a,queenRoadmapScene.ts).prefers-reduced-motion.?engine=canvas, keeps the SVG, remembered per viewer.goals.jsongains the reasons stages 3, 4 and 8 are locked. It also gains stage 9, the endgame: the whole BrowserOS browser and every third-party dependency.Its goal is [roadmap] Stage 9: Endgame: the browser and every dependency t27#4858 (
6dc2eeb).The endgame is measured (
35dc52a). BrowserOS pins Chromium 146.0.7680.31: itsBASE_COMMITis that tag's commit. Chromium'ssrcat that tag, plus the V8 and Skia commits it pins, were read from the GitHub mirrors and counted by theroadmap-stack.mjsrules:src4d322510a35ee19022820goals.jsonrecords the measurement asmeasured(bytes, date and commits). A stage with a measurement uses it alone, so the endgame boss has HP 1342 MB instead of "not measured", and stage 2's agent-server is not counted twice.The running services lock 2,187 npm packages, 650 crates and 35 Go modules.
7364a35). Each is computed from public facts, andqa/roadmap-game-contract.mjsholds it:.t27cells, read from the Queen's public leaderboard.fbc3ee5,.github/workflows/claude-code-review.yml). The owner asked for this.claude-reviewrun failed about two seconds in, because the OAuth token was refused andANTHROPIC_API_KEYwas empty.https://api.z.ai/api/anthropic, and Claude Code reaches it throughANTHROPIC_BASE_URLandANTHROPIC_AUTH_TOKEN. The key comes from a newZAI_API_KEYsecret.glm-4.6by default, the z.ai id in the Queen's own model catalogue. AZAI_MODELrepository variable overrides it, and Haiku-sized calls go toglm-4.5-air.Related Issue
vargets a shadowing local, so the spec does not compile t27#4847 andvalidate-vacuitynever looks inside a braceless test, so anassert truetest written without braces goes uncounted t27#4848.Specification Link
Spec:
apps/website/specs/queen/wars.t27. It is the source of truth.wars.json,wars.t27(public) andqueenWars.generated.tsare regenerated byscripts/queen-wars-from-spec.mjs.Changes Made
bee-triarm. New experiment statejudged(every arm ran, no winner declared). Three new metrics:total-tokens,judge-calls,mutants-killed. Three experiments, six sealed runs, 48 measurements.ZAI_API_KEYsecret is missing.Files Changed
apps/website/specs/queen/wars.t27: the ledger (source of truth)apps/website/public/queen/wars.{json,t27},src/lib/queenWars.generated.ts: generatedapps/website/scripts/queen-wars-from-spec.mjs(+ test): arm set,judged, paired-arm ruleapps/website/qa/queen-wars-contract.mjs: five arms; observed runs must pin artifactsapps/website/src/components/QueenWars.{tsx,css}: the viewapps/website/public/queen/runs/**: sealed campaign artifactsapps/website/src/data/blog/**: two new posts, and corrections to the previous oneapps/website/src/components/QueenRoadmapGame.tsx,queenRoadmapGame.css: the comb (3D by default, flat as fallback), ships, targets and sectorsapps/website/src/components/queenRoadmapScene.ts: the Babylon.js wallapps/website/src/components/QueenRoadmap.tsx:measureduses that measurement, and its row names the source.apps/website/public/roadmap/goals.json:lockedreasons, and stage 9, the endgame, linked to [roadmap] Stage 9: Endgame: the browser and every dependency t27#4858 with its measurementapps/website/qa/queen-contrast-contract.mjs: the new sheet, its token scope, and its surfaces sealed by.rmapps/website/src/lib/roadmapGame.ts: the rules (raid, honey, boss opening, cracks, round pulse, titles, rows)apps/website/qa/roadmap-game-contract.mjs,package.json(check:roadmap-game),.github/workflows/website-checks.yml: the contract, run in CI.github/workflows/claude-code-review.yml: the review on z.aiGolden Chain Checklist
wars.t27, and the projections are generated from itTesting Checklist
npm run check:wars: 5 configurations, 4 experiments, 7 runs, 53 measurements; spec tests 5, asserts 38, all holdnpm run test:wars-spec: 26/2673bca5a:node scripts/typecheck-ratchet.mjsreports 179 errors across 26 files, the same as baseline; no file gained errors73bca5a:vite build, then the EN and RU language audits and the mobile audit on the built site: PASS, 35 routes each73bca5a: these pass:check:queen-contrast: 24 pairs, worst 5.39:1, the new card and buttons low-passed by a backdrop blur;check:roadmap-game;check:queen-languages: 377 EN and 377 RU keys;check:render: no uncaught errors.7364a35: everynpm runcheck inwebsite-checks.ymlpasses locally, 38 of them including the newcheck:roadmap-game.73bca5a):#4384 · tools/jtag/mpsse_jtag.py · built · 3 fn.claude-code-review.ymlparses, with three steps; the review step runs only when the key check says ready, and no OAuth input is left.fbc3ee5theclaude-reviewjob is green for the first time: the key check found no secret and skipped the review with its notice.ZAI_API_KEYsecret, and this container cannot reach api.z.ai to try the endpoint.check:queen-viewport(Chromium,warsview): fails, but identically onmain. The one failure is.queen27-sectors-emptywithsectors=0, which is not a WARS or ROADMAP element.Reviewer Notes
ZAI_API_KEYsecret under Settings > Secrets and variables > Actions for the z.ai review to run. Optionally, set aZAI_MODELvariable, for exampleglm-4.7, if the account has it.judgedtocomplete, and the generator enforces that rule.ternary_encoding.t27. This was already open as A recovered assertion that actually fails: trits_to_bits round-trip in ternary_encoding t27#2778; the byte pair and the four blockers were added there.varassigned in a test body. Filed as A test or bench body that assigns a module-levelvargets a shadowing local, so the spec does not compile t27#4847.validate-vacuityskips braceless tests. Filed asvalidate-vacuitynever looks inside a braceless test, so anassert truetest written without braces goes uncounted t27#4848.goals.json, not part ofroadmap-stack.mjs. That script reads the GitHub tree API, which truncates Chromium's 481,709-file tree, so it cannot take this measurement. The figure was taken from shallow clones and is dated in the file.craft_*), but only 14 other models of it were ever committed, and those were removed in feat(queen): one kit, one meaning #949.queenRoadmapScene.ts, so a kit model can replace it when the file is available.{ "version": 1, "head_sha": "fbc3ee54ca2f65a394a7ffa78e1550284ce432d8", "summary": "The WARS arena had a TRI decision layer as its variable but no arm that used it, and one blocked run. It now has a bee-tri arm and a real paired campaign on three gHashTag/t27 issues, judged by the issues' own acceptance commands, plus a view that reads at a glance; the blog gains a post correcting an FPGA receipt claim; the ROADMAP tab opens with the game board of the rewrite, drawn in 3D in Babylon.js like the Queen's field, with an endgame boss measured at 1,341.7 MB of Chromium, V8 and Skia source; and the PR review runs on z.ai instead of a missing Anthropic key.", "changes": [ "specs/queen/wars.t27: bee-tri arm, experiment state judged, metrics total-tokens/judge-calls/mutants-killed, experiments for t27 issues 4614/4695/4613, six sealed runs and 48 measurements; projections regenerated", "scripts/queen-wars-from-spec.mjs: a judged or complete experiment needs sealed bee-baseline and bee-tri runs; complete still needs an observed model id", "src/components/QueenWars.tsx and .css: campaign glance strip, head-to-head grid, measured-only scoreboard with per-row bars, five lanes", "public/queen/runs/: per-run patch, judge transcript, mutant review and arm report; campaign README, issue texts and judge scripts", "src/data/blog: new post trained-weights-ran-receipts-were-not-checked (EN+RU), with the fixed harness's 2026-09-27 board results; new post golden-ratio-weights-ran-on-the-board (EN+RU) on the pre-registered GFTernary x Z[phi] board run; dated corrections and a 0.97% to 1.01% fix in a-small-agent-needs-an-exact-judge", "src/components/QueenRoadmapGame.tsx and queenRoadmapGame.css: ROADMAP level II - the apex-down comb of port tasks (quickest first from the seed), ships over cells the public board reports running, priority targets, sectors; unknown counts read as unknown", "src/components/queenRoadmapScene.ts: the comb drawn in Babylon.js like the Queen's field - hex prisms by state on a wall, ships with beams, the round's band, ACES tone mapping and glow, lift and card under the pointer, click to the issue; on-demand rendering; the flat SVG when WebGL is missing or Flat is chosen", "public/roadmap/goals.json: locked reasons for stages 3, 4 and 8, and stage 9 (the endgame: BrowserOS and every dependency) linked to its goal gHashTag/t27#4858, with a measured field: Chromium 146.0.7680.31 src 1,144.7 MB, V8 145.6 MB, Skia 51.4 MB by the roadmap-stack.mjs rules; QueenRoadmap.tsx uses a stage's own measurement when it has one; qa/queen-contrast-contract.mjs registers the new sheet", "src/lib/roadmapGame.ts, QueenRoadmapGame.tsx: the raid of the day (shared with gHashTag/t27's feeder), honey, cracked cells, the Queen's round pulse, boss cards and builders from the public leaderboard; qa/roadmap-game-contract.mjs holds the rules and runs in website-checks", ".github/workflows/claude-code-review.yml: the review runs on z.ai (ANTHROPIC_BASE_URL https://api.z.ai/api/anthropic, ANTHROPIC_AUTH_TOKEN from the ZAI_API_KEY secret, model glm-4.6 or the ZAI_MODEL variable); without the secret it is skipped with a notice" ], "tests": [ { "command": "npm run check:wars", "status": "passed", "result": "5 configurations, 4 experiments, 7 runs, 53 measurements; spec tests 5, asserts 38 hold; contract passes", "evidence": "local run in apps/website at 7364a35; later commits touch only the ROADMAP files and one workflow" }, { "command": "npm run test:wars-spec", "status": "passed", "result": "26 of 26, including a negative control for the judged state", "evidence": "local run in apps/website at 7364a35; later commits touch only the ROADMAP files and one workflow" }, { "command": "node scripts/typecheck-ratchet.mjs", "status": "passed", "result": "179 errors across 26 files, equal to baseline; no file gained errors", "evidence": "local run in apps/website at 73bca5a; fbc3ee5 changes only a workflow" }, { "command": "npx vite build, then npm run audit:en, audit:ru and audit:mobile against the built site", "status": "passed", "result": "build succeeds; EN and RU language audits PASS on 35 routes each; no route scrolls sideways", "evidence": "local run in apps/website at 73bca5a; fbc3ee5 changes only a workflow" }, { "command": "npm run check:queen-contrast, check:roadmap-game, check:queen-languages and check:render", "status": "passed", "result": "contrast 24 pairs, worst 5.39:1, the new card and view buttons low-passed; the game rules hold; 377 EN and 377 RU keys; no uncaught errors", "evidence": "local run in apps/website at 73bca5a; fbc3ee5 changes only a workflow" }, { "command": "every npm run check in website-checks.yml", "status": "passed", "result": "38 of 38 pass", "evidence": "local run in apps/website at 7364a35" }, { "command": "headless Chromium (SwiftShader) render of #/queen?tab=roadmap at 1366 and 390 px, EN and RU, in 3D", "status": "passed", "result": "the scene is ready with 67 task cells, 3 ships and 2 raid cells, the whole frame in view; pointing at a cell shows its card in both languages; no sideways scroll", "evidence": "local run at 73bca5a" }, { "command": "headless Chromium with WebGL disabled; the Flat button and a reload; reduced motion", "status": "passed", "result": "no WebGL draws the flat comb and names the reason; Flat survives a reload and 3D returns; reduced motion draws the wall still", "evidence": "local run at 73bca5a" }, { "command": "yaml.safe_load of .github/workflows/claude-code-review.yml, then the claude-review job on this head", "status": "passed", "result": "three steps; the review step runs only when the key check says ready; the base URL is https://api.z.ai/api/anthropic; no OAuth input is left; the job is green, the review skipped with its notice because the secret is not set", "evidence": "local parse at this head and the claude-review check run on fbc3ee5" }, { "command": "the claude-review job against api.z.ai", "status": "not_run", "result": "needs the ZAI_API_KEY secret; api.z.ai is not reachable from this container", "evidence": "not run: the ZAI_API_KEY secret is not set yet" }, { "command": "git ls-tree -r -l over shallow clones of github.com/chromium/chromium at tag 146.0.7680.31, v8/v8 at 0a35ee1 and google/skia at 9022820, counted by the roadmap-stack.mjs extension map and skip rules", "status": "passed", "result": "Chromium src 1,144,688,069 bytes, V8 145,602,411, Skia 51,375,089; 1,341,665,569 in all, recorded in goals.json", "evidence": "local measurement on 2026-09-27; the commits are named in goals.json" }, { "command": "CHROME_PATH=chromium npm run check:queen-viewport", "status": "failed", "result": "fails on .queen27-sectors-empty with sectors=0 at five sizes, identically on main", "evidence": "local runs on 7364a35 and on main 45e3c9f" }, { "command": "judge/accept.py and judge/mutate.py over six arm worktrees of gHashTag/t27 at afe2186c", "status": "passed", "result": "6 of 6 arms pass acceptance; new tests kill 3/3, 4/4 and 2/4 mutants (2 blocked by an existing invariant)", "evidence": "apps/website/public/queen/runs/ at commit 7d3a47a" } ], "limitations": [ "No winner is declared: the executor model id is not recorded, so experiments are judged, not complete.", "Three issues are too few to rank the arms; the decision layer changed no acceptance outcome here.", "Issue 4613's tests pass only in a review copy with four pre-existing blockers removed.", "The board runs used random test activation vectors, not the model's real activations, and no forward pass; the UART link has no flow control, so answers in flight are kept under the CP2102N's 512-byte buffer, and what stalls the bridge is not measured.", "Level II was not rendered against the live GitHub search and Queen board, which this container cannot reach; it was rendered against files holding real t27 issue titles, and with both sources refused.", "The endgame's figure covers Chromium src, V8 and Skia only; Blink sits under the skipped third_party/, and the other 258 DEPS repositories and the npm, crate and Go dependencies are not measured in bytes.", "The 3D ships are primitives, not the Kenney Space Kit craft models: those were never committed and kenney.nl is not reachable from this container.", "The z.ai review is configured but has not run: it needs the ZAI_API_KEY secret, and the endpoint and model id were taken from z.ai's Claude Code guide and the Queen's catalogue, not tried from here." ], "tags": [ "WARS", "Verification", "FPGA", "Agents", "Roadmap" ], "blog": { "title": "The board's receipts were never checked. Now all 403,200 are.", "summary": "This PR publishes that article: a harness reported authenticated receipts it never compared, and the fix is shown able to fail.", "outline": [ "The board computed 28,416 of 28,416 trained-weight rows bit-exact, while the harness counted a status byte as authentication.", "The fixed harness fails on seven misbehaving cells, and RTL co-simulation covers all seven matrix kinds and int8 activations with no new hardware.", "On the board, the fixed harness verified all 403,200 receipts across the 42 trained matrices, 33,792 of 33,792 rows bit-exact, and 51,840 more for w_down with int8 activations; runs with too many jobs in flight stopped on lost UART bytes, not on a wrong answer, and a registered test put the threshold at the CP2102N bridge's 512-byte receive buffer.", "Still open: what stalls the USB bridge, real activations, a forward pass on the board, a power figure, and receipts that a third party can verify." ] } }🤖 Generated with Claude Code
https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY