Skip to content

WARS: Bee + TRI arm, a real paired campaign, a readable arena; ROADMAP level II; blog posts - #1184

Merged
gHashTag merged 15 commits into
mainfrom
claude/peaceful-noether-dperdx
Sep 27, 2026
Merged

gHashTag merged 15 commits into
mainfrom
claude/peaceful-noether-dperdx

Conversation

@gHashTag

@gHashTag gHashTag commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Description

This PR makes five changes.

  1. WARS gets a real paired campaign and an arm for its own variable.
  2. The WARS view is easier to read. Changes:
    • a campaign-at-a-glance strip
    • a head-to-head grid of every experiment; clicking a row opens that experiment
    • a scoreboard that shows only the metrics that were measured, with a relative bar per row, and lists the rest as not measured
    • five lanes
  3. Blog.
    • New post (EN + RU): The board's receipts were never checked. Now all 403,200 are.
      • It corrects a claim on gHashTag/trinity-fpga, fixed in fix(conformance): check tern_tc receipts, and cover every ternary matrix trinity-fpga#800.
      • It reports the fixed harness on the board (2026-09-27). All 42 ternary matrices of the trained model passed: 403,200/403,200 receipts verified, 33,792/33,792 rows bit-exact, 85.5 s. Layer 5 w_down with int8 activations passed with all 51,840 receipts verified.
      • With too many jobs in flight, long runs stopped when the UART link lost answer bytes. No wrong answer came before any hole. A test registered before it ran put the loss threshold between 494 and 570 bytes in flight (window 26 clean, window 30 slipped). That is around the CP2102N bridge's 512-byte receive buffer, and the node has no flow control.
    • Second new post (EN + RU): Golden-ratio weights ran on the board. The phi cost one add per output.
      • It covers the pre-registered GFTernary × Z[φ] run on the same node (fix(conformance): check tern_tc receipts, and cover every ternary matrix trinity-fpga#800, 8a20321).
      • Layer 0 verified 403,200/403,200 receipts, and 5,632/5,632 Z[φ] rows matched a plain Z[φ] multiplication oracle.
      • No multiplier was used: the node computes ternary dots, and the host adds once per output.
      • It links the earlier finding that φ in GFTernary is a scale, not information.
    • The previous post's stale open questions get dated updates.
    • That post's summary said 0.97%. The mean of its own table is 1.01%, so the summary now says 1.01%.
  4. ROADMAP, level II: the game board of the rewrite (aa35629). The tab measured the goal but never showed anyone playing it. It now opens with:
    • The comb. An inverted pyramid in the shape of the t27 mark. The apex is the seed, t27c. Every other cell is one port task, laid from the apex upward with the quickest first (fewest functions), so the comb grows from its point. The widest row is the whole stack in .t27, with the BrowserOS browser and every dependency included.
    • Drawn in 3D, in Babylon.js like the Queen's field (73bca5a, queenRoadmapScene.ts).
      • The same cells, as hex prisms on a wall that faces the player. A prism is as tall as its file has come: built honey 1.7, review 1.0, a bee building 0.8, held 0.35. A free target is a thin plate in its language's colour, and the seed t27c is the tallest cell and lit.
      • The look. Rims glow, and cracked cells keep their crack. Tone mapping (ACES) and a glow layer match the field.
      • Interaction.
        • The pointer lifts the cell under it and shows a card; a click opens the issue.
        • A tap pins the card, with a link to the issue.
        • A drag tilts the wall, and the wheel still scrolls the page.
      • Cost. It renders on demand, at 30 FPS only while something moves and the canvas is on screen, and it holds still under prefers-reduced-motion.
      • Fallback.
        • Without WebGL, the flat SVG comb is drawn with the reason.
        • A Flat button, or the field's own ?engine=canvas, keeps the SVG, remembered per viewer.
        • Babylon loads only on the 3D path, in its own 18 KB chunk over the field's core.
    • Ships. One flies over each cell that the Queen's public board reports running, with a beam into it. Nothing is drawn that the board did not report. If the board does not answer, the page says so.
    • Priority targets. The quickest free port tasks, each linked to its issue.
    • Sectors. Every stage with its state: captured, under attack, open front, no tasks yet, or locked with the reason.
    • goals.json gains the reasons stages 3, 4 and 8 are locked. It also gains stage 9, the endgame: the whole BrowserOS browser and every third-party dependency.
      • Its goal is [roadmap] Stage 9: Endgame: the browser and every dependency t27#4858 (6dc2eeb).

      • The endgame is measured (35dc52a). BrowserOS pins Chromium 146.0.7680.31: its BASE_COMMIT is that tag's commit. Chromium's src at that tag, plus the V8 and Skia commits it pins, were read from the GitHub mirrors and counted by the roadmap-stack.mjs rules:

        Part Commit Size
        Chromium src 4d32251 1,144.7 MB
        V8 0a35ee1 145.6 MB
        Skia 9022820 51.4 MB
        Total 1,341.7 MB
        • That total is seven times the 192.2 MB the count holds for the whole stack.
        • Blink (132.4 MB outside its tests) and 258 other DEPS repositories are not in it.
      • goals.json records the measurement as measured (bytes, date and commits). A stage with a measurement uses it alone, so the endgame boss has HP 1342 MB instead of "not measured", and stage 2's agent-server is not counted twice.

      • The running services lock 2,187 npm packages, 650 crates and 35 Go modules.

    • The tasks the comb shows are filed by the roadmap feeder in feat(queen): feed the swarm from the roadmap, with a brief that says what the game is t27#4855.
    • The rules of the game (7364a35). Each is computed from public facts, and qa/roadmap-game-contract.mjs holds it:
      • Raid of the day. One sector per UTC day, taken in turn from stages 1, 2, 5, 6 and 7. The t27 feeder files that sector first, and both sides pin the list, so the board never names a raid nobody runs. Today the raid is stage 6.
      • Honey. The functions ported in built cells. Raid cells closed today count twice. It is this board's own score, not the Queen's XP.
      • Cracked cells. A built port file that has an open defect naming it.
      • The Queen's round. A band of light climbs the comb on the board's own pulse, and a countdown shows when the next round starts.
      • Bosses. Stage 8 and the endgame. Their HP is the measured source still to rewrite; the card also shows the board's rule for opening the fight and what really stands in the way.
      • Builders. The lenders whose lanes built .t27 cells, read from the Queen's public leaderboard.
      • Not done. Colouring each cell by the model or key that built it: no public source names the builder per issue.
  5. The PR review runs on z.ai (fbc3ee5, .github/workflows/claude-code-review.yml). The owner asked for this.
    • Why: every claude-review run failed about two seconds in, because the OAuth token was refused and ANTHROPIC_API_KEY was empty.
    • How it connects: z.ai serves the Anthropic Messages API at https://api.z.ai/api/anthropic, and Claude Code reaches it through ANTHROPIC_BASE_URL and ANTHROPIC_AUTH_TOKEN. The key comes from a new ZAI_API_KEY secret.
    • Model: glm-4.6 by default, the z.ai id in the Queen's own model catalogue. A ZAI_MODEL repository variable overrides it, and Haiku-sized calls go to glm-4.5-air.
    • Without the secret (a fork, or before the secret is added), the job says so in a notice and skips the review instead of failing.

Related Issue

Specification Link

Spec: apps/website/specs/queen/wars.t27. It is the source of truth. wars.json, wars.t27 (public) and queenWars.generated.ts are regenerated by scripts/queen-wars-from-spec.mjs.

Changes Made

  • Feature: bee-tri arm. New experiment state judged (every arm ran, no winner declared). Three new metrics: total-tokens, judge-calls, mutants-killed. Three experiments, six sealed runs, 48 measurements.
  • Feature: new WARS view sections: glance strip, head-to-head grid, measured-only scoreboard with bars.
  • Feature: ROADMAP level II. The apex-down comb, drawn in 3D in Babylon.js like the field, with a flat fallback. Ships over running cells, priority targets and sectors, all from live public sources.
  • CI: the PR review runs on z.ai through its Anthropic-compatible endpoint, and it is skipped with a notice when the ZAI_API_KEY secret is missing.
  • Tests: the generator test covers the new ledger. A new negative control checks two refusals: a judged experiment without a sealed TRI run, and a judged experiment promoted to complete without a model id. The contract now checks that observed runs pin their artifacts to a commit. The contrast gate now registers the roadmap comb's stylesheet.
  • Documentation: two new blog posts with the board results, plus corrections to the previous post.

Files Changed

  • apps/website/specs/queen/wars.t27: the ledger (source of truth)
  • apps/website/public/queen/wars.{json,t27}, src/lib/queenWars.generated.ts: generated
  • apps/website/scripts/queen-wars-from-spec.mjs (+ test): arm set, judged, paired-arm rule
  • apps/website/qa/queen-wars-contract.mjs: five arms; observed runs must pin artifacts
  • apps/website/src/components/QueenWars.{tsx,css}: the view
  • apps/website/public/queen/runs/**: sealed campaign artifacts
  • apps/website/src/data/blog/**: two new posts, and corrections to the previous one
  • apps/website/src/components/QueenRoadmapGame.tsx, queenRoadmapGame.css: the comb (3D by default, flat as fallback), ships, targets and sectors
  • apps/website/src/components/queenRoadmapScene.ts: the Babylon.js wall
  • apps/website/src/components/QueenRoadmap.tsx:
    • mounts the game first, and adds the endgame to the goal text;
    • a stage with no measured bytes says "not measured yet";
    • a stage with measured uses that measurement, and its row names the source.
  • apps/website/public/roadmap/goals.json: locked reasons, and stage 9, the endgame, linked to [roadmap] Stage 9: Endgame: the browser and every dependency t27#4858 with its measurement
  • apps/website/qa/queen-contrast-contract.mjs: the new sheet, its token scope, and its surfaces sealed by .rm
  • apps/website/src/lib/roadmapGame.ts: the rules (raid, honey, boss opening, cracks, round pulse, titles, rows)
  • apps/website/qa/roadmap-game-contract.mjs, package.json (check:roadmap-game), .github/workflows/website-checks.yml: the contract, run in CI
  • .github/workflows/claude-code-review.yml: the review on z.ai

Golden Chain Checklist

  • Spec First: the arena is authored in wars.t27, and the projections are generated from it
  • No Manual Edits to generated files

Testing Checklist

  • npm run check:wars: 5 configurations, 4 experiments, 7 runs, 53 measurements; spec tests 5, asserts 38, all hold
  • npm run test:wars-spec: 26/26
  • At 73bca5a: node scripts/typecheck-ratchet.mjs reports 179 errors across 26 files, the same as baseline; no file gained errors
  • At 73bca5a: vite build, then the EN and RU language audits and the mobile audit on the built site: PASS, 35 routes each
  • At 73bca5a: these pass:
    • check:queen-contrast: 24 pairs, worst 5.39:1, the new card and buttons low-passed by a backdrop blur;
    • check:roadmap-game;
    • check:queen-languages: 377 EN and 377 RU keys;
    • check:render: no uncaught errors.
  • At 7364a35: every npm run check in website-checks.yml passes locally, 38 of them including the new check:roadmap-game.
  • The 3D comb, rendered in headless Chromium (SwiftShader) at 1366 px and 390 px, in EN and RU (at 73bca5a):
    • the scene reports ready with 67 task cells, 3 ships and 2 raid cells;
    • the whole frame is in view at the default tilt, with no sideways scroll;
    • pointing at a cell shows its card in both languages, for example #4384 · tools/jtag/mpsse_jtag.py · built · 3 fn.
  • The other paths:
    • with WebGL disabled, the flat comb is drawn with "The 3D wall stopped (WebGL not supported)";
    • Flat switches to the SVG and survives a reload, and 3D switches back;
    • under reduced motion the wall is drawn and holds still.
  • The flat comb (Level II) was rendered in headless Chromium at 1366 px and 390 px, in EN and RU, in two modes:
    • Live. GitHub search, the board and the leaderboard were answered from files holding real t27 port-issue titles, because this container cannot reach those hosts. Result: 78 cells, 3 ships, 8 targets, 4 cracked cells, the round band, 3 builders and 2 bosses, with no sideways scroll. The bosses read HP 20 MB (stage 8) and 1342 MB (the endgame).
    • Down. Both sources refused. Result: every count reads "—", the page names each source that did not answer, and no ship is drawn.
  • claude-code-review.yml parses, with three steps; the review step runs only when the key check says ready, and no OAuth input is left.
  • On fbc3ee5 the claude-review job is green for the first time: the key check found no secret and skipped the review with its notice.
  • The z.ai review has not run yet. It needs the ZAI_API_KEY secret, and this container cannot reach api.z.ai to try the endpoint.
  • check:queen-viewport (Chromium, wars view): fails, but identically on main. The one failure is .queen27-sectors-empty with sectors=0, which is not a WARS or ROADMAP element.

Reviewer Notes

{
  "version": 1,
  "head_sha": "fbc3ee54ca2f65a394a7ffa78e1550284ce432d8",
  "summary": "The WARS arena had a TRI decision layer as its variable but no arm that used it, and one blocked run. It now has a bee-tri arm and a real paired campaign on three gHashTag/t27 issues, judged by the issues' own acceptance commands, plus a view that reads at a glance; the blog gains a post correcting an FPGA receipt claim; the ROADMAP tab opens with the game board of the rewrite, drawn in 3D in Babylon.js like the Queen's field, with an endgame boss measured at 1,341.7 MB of Chromium, V8 and Skia source; and the PR review runs on z.ai instead of a missing Anthropic key.",
  "changes": [
    "specs/queen/wars.t27: bee-tri arm, experiment state judged, metrics total-tokens/judge-calls/mutants-killed, experiments for t27 issues 4614/4695/4613, six sealed runs and 48 measurements; projections regenerated",
    "scripts/queen-wars-from-spec.mjs: a judged or complete experiment needs sealed bee-baseline and bee-tri runs; complete still needs an observed model id",
    "src/components/QueenWars.tsx and .css: campaign glance strip, head-to-head grid, measured-only scoreboard with per-row bars, five lanes",
    "public/queen/runs/: per-run patch, judge transcript, mutant review and arm report; campaign README, issue texts and judge scripts",
    "src/data/blog: new post trained-weights-ran-receipts-were-not-checked (EN+RU), with the fixed harness's 2026-09-27 board results; new post golden-ratio-weights-ran-on-the-board (EN+RU) on the pre-registered GFTernary x Z[phi] board run; dated corrections and a 0.97% to 1.01% fix in a-small-agent-needs-an-exact-judge",
    "src/components/QueenRoadmapGame.tsx and queenRoadmapGame.css: ROADMAP level II - the apex-down comb of port tasks (quickest first from the seed), ships over cells the public board reports running, priority targets, sectors; unknown counts read as unknown",
    "src/components/queenRoadmapScene.ts: the comb drawn in Babylon.js like the Queen's field - hex prisms by state on a wall, ships with beams, the round's band, ACES tone mapping and glow, lift and card under the pointer, click to the issue; on-demand rendering; the flat SVG when WebGL is missing or Flat is chosen",
    "public/roadmap/goals.json: locked reasons for stages 3, 4 and 8, and stage 9 (the endgame: BrowserOS and every dependency) linked to its goal gHashTag/t27#4858, with a measured field: Chromium 146.0.7680.31 src 1,144.7 MB, V8 145.6 MB, Skia 51.4 MB by the roadmap-stack.mjs rules; QueenRoadmap.tsx uses a stage's own measurement when it has one; qa/queen-contrast-contract.mjs registers the new sheet",
    "src/lib/roadmapGame.ts, QueenRoadmapGame.tsx: the raid of the day (shared with gHashTag/t27's feeder), honey, cracked cells, the Queen's round pulse, boss cards and builders from the public leaderboard; qa/roadmap-game-contract.mjs holds the rules and runs in website-checks",
    ".github/workflows/claude-code-review.yml: the review runs on z.ai (ANTHROPIC_BASE_URL https://api.z.ai/api/anthropic, ANTHROPIC_AUTH_TOKEN from the ZAI_API_KEY secret, model glm-4.6 or the ZAI_MODEL variable); without the secret it is skipped with a notice"
  ],
  "tests": [
    {
      "command": "npm run check:wars",
      "status": "passed",
      "result": "5 configurations, 4 experiments, 7 runs, 53 measurements; spec tests 5, asserts 38 hold; contract passes",
      "evidence": "local run in apps/website at 7364a35; later commits touch only the ROADMAP files and one workflow"
    },
    {
      "command": "npm run test:wars-spec",
      "status": "passed",
      "result": "26 of 26, including a negative control for the judged state",
      "evidence": "local run in apps/website at 7364a35; later commits touch only the ROADMAP files and one workflow"
    },
    {
      "command": "node scripts/typecheck-ratchet.mjs",
      "status": "passed",
      "result": "179 errors across 26 files, equal to baseline; no file gained errors",
      "evidence": "local run in apps/website at 73bca5a; fbc3ee5 changes only a workflow"
    },
    {
      "command": "npx vite build, then npm run audit:en, audit:ru and audit:mobile against the built site",
      "status": "passed",
      "result": "build succeeds; EN and RU language audits PASS on 35 routes each; no route scrolls sideways",
      "evidence": "local run in apps/website at 73bca5a; fbc3ee5 changes only a workflow"
    },
    {
      "command": "npm run check:queen-contrast, check:roadmap-game, check:queen-languages and check:render",
      "status": "passed",
      "result": "contrast 24 pairs, worst 5.39:1, the new card and view buttons low-passed; the game rules hold; 377 EN and 377 RU keys; no uncaught errors",
      "evidence": "local run in apps/website at 73bca5a; fbc3ee5 changes only a workflow"
    },
    {
      "command": "every npm run check in website-checks.yml",
      "status": "passed",
      "result": "38 of 38 pass",
      "evidence": "local run in apps/website at 7364a35"
    },
    {
      "command": "headless Chromium (SwiftShader) render of #/queen?tab=roadmap at 1366 and 390 px, EN and RU, in 3D",
      "status": "passed",
      "result": "the scene is ready with 67 task cells, 3 ships and 2 raid cells, the whole frame in view; pointing at a cell shows its card in both languages; no sideways scroll",
      "evidence": "local run at 73bca5a"
    },
    {
      "command": "headless Chromium with WebGL disabled; the Flat button and a reload; reduced motion",
      "status": "passed",
      "result": "no WebGL draws the flat comb and names the reason; Flat survives a reload and 3D returns; reduced motion draws the wall still",
      "evidence": "local run at 73bca5a"
    },
    {
      "command": "yaml.safe_load of .github/workflows/claude-code-review.yml, then the claude-review job on this head",
      "status": "passed",
      "result": "three steps; the review step runs only when the key check says ready; the base URL is https://api.z.ai/api/anthropic; no OAuth input is left; the job is green, the review skipped with its notice because the secret is not set",
      "evidence": "local parse at this head and the claude-review check run on fbc3ee5"
    },
    {
      "command": "the claude-review job against api.z.ai",
      "status": "not_run",
      "result": "needs the ZAI_API_KEY secret; api.z.ai is not reachable from this container",
      "evidence": "not run: the ZAI_API_KEY secret is not set yet"
    },
    {
      "command": "git ls-tree -r -l over shallow clones of github.com/chromium/chromium at tag 146.0.7680.31, v8/v8 at 0a35ee1 and google/skia at 9022820, counted by the roadmap-stack.mjs extension map and skip rules",
      "status": "passed",
      "result": "Chromium src 1,144,688,069 bytes, V8 145,602,411, Skia 51,375,089; 1,341,665,569 in all, recorded in goals.json",
      "evidence": "local measurement on 2026-09-27; the commits are named in goals.json"
    },
    {
      "command": "CHROME_PATH=chromium npm run check:queen-viewport",
      "status": "failed",
      "result": "fails on .queen27-sectors-empty with sectors=0 at five sizes, identically on main",
      "evidence": "local runs on 7364a35 and on main 45e3c9f"
    },
    {
      "command": "judge/accept.py and judge/mutate.py over six arm worktrees of gHashTag/t27 at afe2186c",
      "status": "passed",
      "result": "6 of 6 arms pass acceptance; new tests kill 3/3, 4/4 and 2/4 mutants (2 blocked by an existing invariant)",
      "evidence": "apps/website/public/queen/runs/ at commit 7d3a47a"
    }
  ],
  "limitations": [
    "No winner is declared: the executor model id is not recorded, so experiments are judged, not complete.",
    "Three issues are too few to rank the arms; the decision layer changed no acceptance outcome here.",
    "Issue 4613's tests pass only in a review copy with four pre-existing blockers removed.",
    "The board runs used random test activation vectors, not the model's real activations, and no forward pass; the UART link has no flow control, so answers in flight are kept under the CP2102N's 512-byte buffer, and what stalls the bridge is not measured.",
    "Level II was not rendered against the live GitHub search and Queen board, which this container cannot reach; it was rendered against files holding real t27 issue titles, and with both sources refused.",
    "The endgame's figure covers Chromium src, V8 and Skia only; Blink sits under the skipped third_party/, and the other 258 DEPS repositories and the npm, crate and Go dependencies are not measured in bytes.",
    "The 3D ships are primitives, not the Kenney Space Kit craft models: those were never committed and kenney.nl is not reachable from this container.",
    "The z.ai review is configured but has not run: it needs the ZAI_API_KEY secret, and the endpoint and model id were taken from z.ai's Claude Code guide and the Queen's catalogue, not tried from here."
  ],
  "tags": [
    "WARS",
    "Verification",
    "FPGA",
    "Agents",
    "Roadmap"
  ],
  "blog": {
    "title": "The board's receipts were never checked. Now all 403,200 are.",
    "summary": "This PR publishes that article: a harness reported authenticated receipts it never compared, and the fix is shown able to fail.",
    "outline": [
      "The board computed 28,416 of 28,416 trained-weight rows bit-exact, while the harness counted a status byte as authentication.",
      "The fixed harness fails on seven misbehaving cells, and RTL co-simulation covers all seven matrix kinds and int8 activations with no new hardware.",
      "On the board, the fixed harness verified all 403,200 receipts across the 42 trained matrices, 33,792 of 33,792 rows bit-exact, and 51,840 more for w_down with int8 activations; runs with too many jobs in flight stopped on lost UART bytes, not on a wrong answer, and a registered test put the threshold at the CP2102N bridge's 512-byte receive buffer.",
      "Still open: what stalls the USB bridge, real activations, a forward pass on the board, a power figure, and receipts that a third party can verify."
    ]
  }
}

🤖 Generated with Claude Code

https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY

…aign

Six arms on three real gHashTag/t27 issues (#4614, #4695, #4613), each in its
own worktree at afe2186c, judged by the issues' own acceptance commands with
t27c 0.4.0. Per run: the patch, the judge transcript, a fixed-mutant review of
the new test, and the arm's own report. The campaign folder holds the issue
texts the arms read, the judge and mutation scripts, and the judge's run on
the unmodified base (which fails all three, as the issues say it should).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
… arena

The protocol's variable factor has been the TRI decision layer since #1178,
but no arm used it. bee-tri is that arm: the same Bee, prompt, tools and
budget, with t27c on PATH while it works. A judged or complete experiment
now needs a sealed bee-baseline and bee-tri run; JEV stays a comparison arm.

The ledger gains three real gHashTag/t27 issues (4614, 4695, 4613), six
sealed runs and 48 measurements: acceptance by the issues' own commands,
Queen verdict, elapsed, tool calls, patch lines, and three new metrics
(total-tokens, judge-calls, mutants-killed). All six arms pass; each pair
kills the same mutants; the TRI arms found three tool defects. A new
experiment state, "judged", records that every arm ran without declaring a
winner, because the executor model id is not written to this repository.
A test shows a judged experiment without a sealed TRI run is refused, and
that promoting one to complete without a model id is refused.

The view: a campaign-at-a-glance strip, a head-to-head grid of every
experiment (a row opens it), a scoreboard that shows only measured metrics
with a relative bar per row and lists the rest as not measured, and five
lanes. Checked: check:wars, test:wars-spec (26/26), check:aria,
check:queen-languages, check:queen-responsive, the typecheck ratchet (no
file gained errors) and the viewport contract, whose only failure
(.queen27-sectors-empty) fails identically on main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
…were not checked

New post (EN + RU, published: true). The AX7203 computed 28,416 of 28,416
rows of tern_tc's 320-input matrices bit-exact, and the harness reported
284,160 authenticated receipts without comparing a tag. The post withdraws
that count until the board reruns, shows the fixed harness failing on seven
kinds of misbehaving cell, reports the RTL co-simulation (layer 0, all seven
matrices, 67,200/67,200 receipts; w_down with int8 activations, 51,840/51,840)
and why w_down, wk, wv and int8 activations need no new hardware. Receipts
link to gHashTag/trinity-fpga#800 and pinned files.

"A small agent needs an exact judge": the open question "No IGLA model has
run on a board" and "the 13M model has been neither trained nor run" are
updated with dated notes, and the summary's 0.97% is corrected to 1.01%, the
mean of the post's own MultiPL-E table.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
@github-actions github-actions Bot added the status:in-progress 🔵 Agent working label Sep 26, 2026
check:queen-contrast failed on CI: .queen-wars-glance tiles lay a dark
translucent ground over the live hex lattice with nothing taking the
lattice's edges out of it. Add backdrop-filter: blur(6px), as the contract
asks. Every command in website-checks.yml then passes locally except
check:queen-viewport, whose .queen27-sectors-empty failure is identical on
main in this environment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY

gHashTag commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner Author

CI status on this PR. One failure was this PR's and is fixed in 4ae20f1. The others were not caused by this PR.

  • checks (website-checks): this was mine. check:queen-contrast flagged the .queen-wars-glance tiles, a translucent ground over the live lattice with no blur. 4ae20f1 adds backdrop-filter: blur(6px). The whole job, including check:queen-viewport, now passes in CI on 4ae20f1. The PR description says the viewport contract fails; that was true only in my local sandbox, which has no sectors data, and it failed there identically on main.
  • ⚡ Brain Health Check and 📋 Brain Health Report: not caused by this diff, which touches no Zig. tri stress --health prints stress-test: TODO - not implemented yet, so no Score: line is produced. The workflow's own comment says this gate "has failed on every run in its history, including on main". The report job is red on the already-merged fix(website): repair blog merge fallout that blocked the site build #1180 as well. There is no fix to port, because the command is unimplemented.
  • pr-opened (project auto-status): fails with gh: Bad credentials (HTTP 401) from the project-board token. The cause is a workflow secret or token, not code.
  • claude-review: fails on both heads. Each run ends after a few seconds with is_error:true and $0 cost, and ANTHROPIC_API_KEY is empty in the job env. The cause is workflow credentials or configuration, not the diff.
  • The red report run on 4ae20f1: this ran before the PR description's head_sha was updated. The run after the update passed, and T27 work report is green on 4ae20f1.

The credential failures need someone with access to repository secrets. They cannot be fixed from this PR. Apart from those, the PR is waiting on review.


Generated by Claude Code

The post now reports what the fixed harness did on the AX7203 on
2026-09-27. Layer 5 w_down with int8 activations passed at 8 jobs in
flight: 51,840/51,840 receipts verified and 320/320 rows bit-exact. The
random matvec passed as well. The two long runs at 64 in flight stopped
when the UART link lost 16 and 4 answer bytes. No wrong answer or bad
tag came before either hole.

The open questions, receipts and next steps (EN and RU) now say that the
full 42-matrix run is pending at 24 in flight, and that the cause of the
lost bytes is not settled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
The fixed harness ran all 42 ternary matrices of the trained tern_tc on
the AX7203 at 24 jobs in flight. It verified 403,200 of 403,200 receipts,
33,792 of 33,792 rows were bit-exact, and the run took 85.5 s at 4,716
answers/s. The post now leads with that result. The title changes to say
so; the slug stays the same.

The lost-bytes analysis now cites the measured 22.4 ms host pause, which
at 64 jobs in flight is enough to fill the whole queue. The open
questions keep what is still missing: the window-64 control on the new
harness, the real activations, a forward pass and power.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
The control ran the same full job at 64 jobs in flight on the same
harness, setup and session, and lost 59 bytes after 6,744 jobs. So the
window explains the pass at 24, not the harness change. Its longest host
pause was 4.2 ms, about 1,700 answers before the hole, which is too short
to fill a 64-deep queue. That rules out the explanation this post gave an
hour ago, that answers pile up while the host looks away. The post now
puts the loss in the link itself (the USB adapter, its driver, the hub or
the cable), not yet isolated, keeps 24 in flight as the operating point,
and adds the control's row and log.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
… output

A new post (EN and RU) on the pre-registered GFTernary x Z[phi] run on the
AX7203 node. The TNF paper's t*phi weights, applied to Z[phi] activations,
ran on the unchanged cell with no multiplier. Each output is W.b + (W.a +
W.b)*phi: the node computes only ternary dots and the host adds once per
output. Layer 0 checked 403,200 of 403,200 receipts, and 5,632 of 5,632
Z[phi] rows matched a plain Z[phi] multiplication oracle bit for bit. The
run was registered 38 s before it started.

The post ties this to the earlier finding that phi in GFTernary is a
scale, not information. It also says what the run does not cover: no TNF
accumulator or rounding, one synthetic activation vector, layer 0 only,
no speed claim, and receipts that are not publicly verifiable.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
A test was registered before it ran on the owner's board and then run. At
26 jobs in flight (494 bytes of answers) all 403,200 jobs came back clean.
At 30 (570 bytes) the stream lost 7 bytes after 150,017 jobs. That puts
the loss threshold around the CP2102N bridge's 512-byte receive buffer.
The bridge's datasheet asks for handshaking above 1 Mbaud, and the node has
none.

The post now explains the link failures that way and keeps the side
prediction that missed. It adds both runs to the table, updates the counts
(185,535 verified answers before four holes), and changes the next step to
flow control. EN and RU.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
…t tasks up

The ROADMAP tab measured the rewrite and never showed anyone playing it.
It now opens with the game board:

- the comb: an inverted pyramid in the shape of the t27 mark. The apex is
  the seed (t27c); every other cell is one port task, laid from the apex
  upward quickest first (fewest functions), so the comb grows from its
  point; the widest row is the whole stack in .t27, the BrowserOS browser
  and every dependency included
- ships: one over each cell the Queen's public board reports running, with
  a beam into it; nothing is drawn that the board did not report, and a
  board that does not answer is named on the page
- priority targets: the quickest free port tasks, with a link into each
- sectors: every stage with its state (captured, under attack, open front,
  no tasks yet, locked and why)

Live and public: GitHub search for the port tasks (open, and closed as
completed) and /queen/public-board. Without the search every count reads
as unknown, not zero.

goals.json gains `locked` reasons for stages 3, 4 and 8, and stage 9, the
endgame: the whole BrowserOS browser and every third-party dependency,
not measured yet. The contrast gate registers the new sheet, its token
scope and its four surfaces, sealed by .rm like the tab's own panels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
@gHashTag gHashTag changed the title WARS: Bee + TRI arm, a real paired campaign, a readable arena; blog correction post WARS: Bee + TRI arm, a real paired campaign, a readable arena; ROADMAP level II; blog posts Sep 27, 2026
… round

LEVEL II gains the rules of a game, each computed from public facts and
held by qa/roadmap-game-contract.mjs:

- the raid of the day: one sector per UTC day over stages 1, 2, 5, 6 and
  7; gHashTag/t27's roadmap feeder feeds the same sector first, and both
  sides pin the list, so the board never names a raid nobody runs
- honey: the functions ported in built cells, raid cells closed today
  counted twice; named as this board's score, not the Queen's XP
- cracked cells: a built port file with an open defect that names it
- the Queen's round: a band of light climbing the comb, timed by the
  board's own pulse, with the next round counted down
- bosses: stage 8 and the endgame, with their measured source as HP, the
  board's opening rule, and what really stands in the way
- builders: lenders whose lanes built .t27 cells, from the Queen's public
  leaderboard

The rules live in src/lib/roadmapGame.ts; website-checks runs the
contract, and the contrast gate seals the raid banner and boss cards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
The endgame had no goal issue, so its row showed none. It now links
gHashTag/t27#4858, and its text carries what was measured on 2026-09-27:
BrowserOS dev is 5.48 MB of source, and the running services lock 2,187
npm packages, 650 crates and 35 Go modules. The engine, Chromium 146, is
fetched at build time and is still not measured, which is what the lock
now says.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
Stage 9 had no HP: the browser is not in any repository the count reads.
BrowserOS pins Chromium 146.0.7680.31 (its BASE_COMMIT 4d3225104176d is
that tag's commit). Chromium's src at that tag, plus the V8 and Skia
commits it pins, were read from the GitHub mirrors and counted by the
roadmap-stack.mjs rules, without Config, Docker and Make:

- Chromium src @4d32251: 1,144.7 MB (C++ 676.6, C 173.4, Java 101.6)
- V8 14.6.202.6 @0a35ee1: 145.6 MB
- Skia @9022820: 51.4 MB

Together that is 1,341.7 MB, seven times the 192.2 MB the count holds for
the whole stack. The rules skip third_party/, so Blink is not in that
figure (132.4 MB outside its tests). Neither are the 258 other DEPS
repositories.

goals.json gains `measured` for stage 9 with the bytes, the date and the
commits. A stage with a measurement uses it alone, so the boss bar and the
stage row show it without counting stage 2's agent-server twice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
…field

The comb was a flat SVG while the Queen's field is a Babylon.js wall. It is
now drawn the same way by default (queenRoadmapScene.ts), from the same
cells in the same order:

- every cell is a hex prism on a wall facing the player, as tall as its
  file has come: built honey 1.7, review 1.0, a bee building 0.8, held 0.35,
  a free target a thin plate in its language's colour, the seed t27c the
  tallest and lit; rims glow, cracked cells keep their crack;
- a ship with a beam hovers over each cell the Queen's board reports
  running, and the round is a band of light climbing the wall on the
  board's pulse (the same pulsePhase as the flat comb);
- ACES tone mapping and a glow layer as on the field; the pointer lifts
  the cell under it and shows a card, a click opens the issue, a tap pins
  the card with a link; a drag tilts the wall, the wheel scrolls the page;
- it renders on demand, at 30 FPS only while something moves and the canvas
  is on screen, and stays still under prefers-reduced-motion.

Without WebGL the flat comb is drawn with the reason; a Flat button (or the
field's ?engine=canvas) keeps the SVG, remembered per viewer. Babylon loads
only in the 3D path, in a chunk of its own (18 KB over the field's core).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
Every claude-review run failed two seconds in: the OAuth token was refused
and ANTHROPIC_API_KEY was empty. The owner asked for z.ai.

z.ai serves the Anthropic Messages API at https://api.z.ai/api/anthropic,
and Claude Code reaches it through ANTHROPIC_BASE_URL and
ANTHROPIC_AUTH_TOKEN. The key comes from a ZAI_API_KEY secret. The model is
glm-4.6, the z.ai id in the Queen's own model catalogue, unless a ZAI_MODEL
repository variable names another; Haiku-sized calls go to glm-4.5-air.

Without the secret (a fork, or before it is added) the job says so in a
notice and skips the review instead of failing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
@gHashTag
gHashTag marked this pull request as ready for review September 27, 2026 11:43
@gHashTag
gHashTag merged commit f172d3d into main Sep 27, 2026
20 of 24 checks passed
@github-actions github-actions Bot added status:completed Done and removed status:in-progress 🔵 Agent working labels Sep 27, 2026
github-actions Bot added a commit that referenced this pull request Sep 27, 2026
Merge pull request #1184 from gHashTag/claude/peaceful-noether-dperdx

WARS: Bee + TRI arm, a real paired campaign, a readable arena; ROADMAP level II in 3D with the measured endgame; two blog posts; the PR review on z.ai

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011vMJx1ZWk2hq58jyEecRBY
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants