Skip to content

feat(repo): model scorecard verdicts, the ratchet and the published schema - #44

Closed
systemfsoftware-maker wants to merge 7 commits into
mainfrom
verify/scorecard-model
Closed

systemfsoftware-maker wants to merge 7 commits into
mainfrom
verify/scorecard-model

Conversation

@systemfsoftware-maker

@systemfsoftware-maker systemfsoftware-maker commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Layer 1 of the rat-stack scorecard stack (plan: docs/plans/2026-10-05-2151-feat-ratstack-scorecard-plan.md, approved by Kiro on 2026-10-05). This layer is inert: nothing runs it until the CI layer (U3) wires it in.

What lands

  • The plan document. It is committed in its own docs(global) commit.
  • judgeRow (src/model/judge-row.workflow.ts) decides each row's verdict under KTD1:
    • Beaten only when the starter's worst run is strictly better than rat-stack's best run.
    • Tie when the two run ranges are identical.
    • Instrument error when a count does not reproduce, when a side produced the wrong number of runs, or when a rat-stack citation no longer holds at the pin.
  • compareWithMain (src/model/compare-with-main.workflow.ts) is the KTD2 ratchet.
    • It fails a PR on an instrument error, on a row beaten on main that the PR no longer beats, and on a starter median worse than main's worst run. A starter value that disappears also fails.
    • A missing secret, a fork PR without a preview, and a changed metric definition are neutral.
    • With no main artifact, the result is a first baseline.
  • Assembly and rendering. assembleScorecard builds the R10 document. renderSummary writes the job-summary table. scorecard.schema.json is the JSON contract.
  • The registry. src/metrics/registry.ts holds the M1-M29 rows and their sub-rows as data.

Deviations from the plan, declared (CONST-W3)

  • Workflow file names. I named the two decision modules after their decisions (judge-row, compare-with-main) instead of the plan's verdict and ratchet. Rendering and assembly make no decision, so they are plain pure modules (scorecard-document.ts, summary-table.ts) instead of a render.workflow.ts.
  • Local imports in the workflows. The workflows import dispatch.ts, runs.ts and type-only cell.ts. These are local, pure, cast-free helpers. KTD3's "import nothing" was there to keep third-party code out, and that still holds.
  • The Stryker config moves to U3. U3 adds the release-gate mutation job and its packages, so its config lands there too.
  • Acceptance examples live outside the property files. AE1-AE5 and AE7 are spec literals in acceptance-examples.test.ts. The *.property.test.ts files hold only generated laws.
  • No Effect Schema. The laws draw from hand-written fast-check arbitraries in scorecard.arbitrary.ts, because KTD3 keeps the instrument free of third-party runtime code.

Verification (run in this session)

All commands were run from evals/ratstack-scorecard, with SANDBOX_PROJECT=$PWD.

The launcher is github:systemfsoftware/pnpm-release-management/180122866dd537fa728b5563fb1820fbd2af88cc#sandbox (PR #5). pnpm 12.3.4 and Node 24.20.0 come from nixpkgs 4975466.

$ sandbox --allow-host registry.npmjs.org -- pnpm install
devDependencies: + @fast-check/vitest 0.5.0  + ajv 8.20.0  + fast-check 4.10.2  + vitest 5.0.3
Done in 1.2s using pnpm v12.3.4

$ sandbox -- pnpm vitest run --project model
 ✓ |model| src/model/acceptance-examples.test.ts (6 tests)
 ✓ |model| src/model/judge-row.workflow.property.test.ts (7 tests)
 ✓ |model| src/model/compare-with-main.workflow.property.test.ts (7 tests)
 ✓ |model| src/model/summary-table.property.test.ts (2 tests)
 ✓ |model| src/model/scorecard-document.property.test.ts (3 tests)
 Test Files  5 passed (5)
      Tests  25 passed (25)

$ deno check src/
Check src/metrics/registry.ts … Check src/model/summary-table.ts   (exit 0)

$ pnpm format:check        (exit 0)
$ pnpm gate:tasks          Tasks: 5 successful, 5 total
$ pnpm gate:dist           Tasks: 1 successful, 1 total

I did not run pnpm mutation (Kiro's rule: mutation never runs locally). The scorecard's own mutation job lands in U3.

The first property run found a real defect: an unsupported rat-stack citation let a non-reproducing starter count be judged beaten, because the starter's runs were only validated against a measured rat-stack. The fix validates each side's runs independently (reproducible in judge-row.workflow.ts). A law now pins it.

Sabotage

  1. Beaten rule. Changing Math.max(...s) < Math.min(...r) to <= in judge-row.workflow.ts turns a starter whose every run only equals rat-stack's best run is not beaten red (1 failed, 24 passed). Reverted, and the suite is green again.
  2. Regression band. Comparing the PR median with main's best run instead of its worst in compare-with-main.workflow.ts turns AE5, the self-comparison law and the band law red (3 failed, 22 passed). Reverted, and the suite is green again (25 passed).

Review fix (Kiro, 2026-10-06): offline install, pnpm 12.9.0

Commit aaf7caa on this layer; #46 and #48 rebased on it.

  • The instrument gets its own flake.nix here (it previously arrived in feat(repo): pin rat-stack and measure the static scorecard rows #46): nixpkgs 494ce7fd23ff6a5dff39e1fb11e9b6f2ac74bf25 (nixos-unstable, includes pnpm_12: 12.3.4 -> 12.9.0 NixOS/nixpkgs#566850, merged 2026-10-03), the launcher at prm#5 1801228, and tools-store, a fixed-output pnpm store built from pnpm-lock.yaml through mkPnpmWorkspacePackages.
  • Installs run in the launcher with no --allow-host, which means no network at all (loopback only). The README no longer shows any registry access, and storeDir is dropped from pnpm-workspace.yaml because it empties the fetcher output.
  • Versions checked at that rev: pnpm_12 12.9.0, nodejs_24 24.21.0, deno 2.9.7.
$ nix build .#tools-store                 → sha256-TylxLEQflTlKx6QDk9mFlBOEInUvvBBeOVFUYCURmYo=, 43M
$ nix develop -c sh -c 'pnpm --version'   → 12.9.0
$ sandbox --pnpm-store "$SANDBOX_PNPM_STORE" -- pnpm install --frozen-lockfile
Done in 149ms using pnpm v12.9.0
$ sandbox -- pnpm vitest run              → Tests 25 passed (25)

pnpm 12.9.0 re-resolved the lockfile (--lockfile-only) to the same pnpm-lock.yaml, byte for byte. On #46 the store hash did not change either, and nix build --rebuild .#tools-store confirmed it reproduces under pnpm 12.9.0. That rebuild was needed because a fixed-output path is reused without being rebuilt while its hash is unchanged. #46: 30/30 tests, and scorecard measure --family static gives the same M23 and M28 numbers. #48: 35/35 tests, actionlint clean, dprint clean, gate:tasks and gate:dist green.

Review fixes (Kiro rulings, 2026-10-06)

Commits on this stack, per ruling. Local gate at bd50211: pnpm format:check ok, gate:tasks and gate:dist green, pnpm scorecard:check 40/40 (import rules, then journeys), tsc -p evals/ratstack-scorecard clean, deno check src/ clean, actionlint clean. Mutation runs in CI only.

Ruling Commit Proof
#1+#3 one vitest project, journeys by import b937803 (#44), 849d77b (#46) renaming static-family.test.ts to measure.test.ts still runs it from its record; a src/model/misplaced.test.ts importing the fixture fails scorecard check (exit 1)
#4 absolute bar for Unsupported rows 9bfdcfd (#44) sabotage meetsBar => true: 2 tests red; reverted green
#18 duplicate cell refused 9bfdcfd, bd50211 aggregate on families with a second ratstack M23 cell: DuplicateCell {"id":"M23","side":"ratstack"}, exit 1
#11 static family end to end 849d77b (#46) deleted record: MissingLauncherRecord, red; edited fixture repo without re-producing: StaleLauncherRecord, red; failing tool gives InstrumentError with exited 3
#2 cache never holds InstrumentError bd50211 cache-check on a poisoned entry: CacheUnusable … M23, M28, exit 1; clean entry exit 0; save is gated on the same check
#5 SHA decode + leased bot push bd50211 pin check live: PinMoved 54d3560 -> 65e9465; GitRefSchema requires 40 hex; push uses --force-with-lease=<branch>:<tip> and refuses a tip by another author
#6 token via credential helper bd50211 git config --get credential.helper stores the literal ${GH_TOKEN}; git credential fill yields the secret from env only
#8 HttpClient timeout + retry, typed bd50211 missing repo: GithubApiError naming the request, exit 1 (no silent baseline); 15 s timeout, 3 transient retries
#9 definition hash scope, base-branch instrument, drift, preview ac93e77, bd50211 editing src/harness/sandbox.ts leaves M23/M28 hashes unchanged; editing src/families/static.ts moves both; grading jobs check out base.sha; instrument preview job is continue-on-error
Orchestrator split 849d77b, bd50211 adding import 'effect' to src/families/static.ts fails the import check (exit 1); driver flake has no --allow-net

@systemfsoftware-maker systemfsoftware-maker changed the title verify/scorecard model feat(repo): model scorecard verdicts, the ratchet and the published schema Oct 5, 2026
@systemfsoftware-maker
systemfsoftware-maker added this pull request to stack #47 October 5, 2026 22:40
Kiro approved this plan on 2026-10-05. It defines 29 metric families
measured on both the starter and rat-stack, the verdict and ratchet
rules, and the U1-U9 stack that builds them
…chema

The scorecard's pure core. judgeRow decides beaten, not beaten, tie or
instrument error for one row: the starter's worst run must be strictly
better than rat-stack's best (KTD1), a count must reproduce exactly, and
an unsupported rat-stack invariant counts only while its citation holds.
compareWithMain is the ratchet (KTD2): it fails on instrument errors, on
a row beaten on main and lost on the PR, and on a starter value worse
than main's worst run, and treats missing secrets, fork PRs and changed
metric definitions as neutral.

scorecard.schema.json is the JSON contract CI publishes and /scorecard
will render. registry.ts declares the M1-M29 rows as data.

The decision modules import only each other, so Deno checks them with no
third-party code. Their laws run under vitest inside the sandbox
launcher from pnpm-release-management
The instrument gets its own flake: nixpkgs 494ce7f (pnpm 12.9.0, per
NixOS/nixpkgs#566850), the pnpm-release-management launcher, and a
fixed-output pnpm store built from pnpm-lock.yaml. Installs run in the
launcher with no allowed hosts, so nothing reaches the registry. Drops
the storeDir override, which empties the fetcher output

Review fix for #44 from Kiro
Review fix #1/#3 (CONST-T12): no project split by folder or suffix. A test becomes a journey by importing the journey fixture, added with the first journey
…w counts as beaten

Review fix #4: every row declares ratstackSupport. Required rows turn an unsupported rat-stack cell into an instrument error; MayBeUnsupported rows (M12, bar 0) are beaten only when the starter's worst run meets the bar. Review fix #18: assembled cells are keyed by row and side, so a duplicate cell cannot reach assembly and cellFor no longer folds
starter-brainstorm superseded the Lake 1 plan with 2026-10-06-1703; the scorecard plan's source list now points at it, with the unit and ruling anchors it uses there
…bsolute bar

Since the absolute bar for unsupported rat-stack rows, judging is no longer symmetric under negating runs and flipping direction: the bar does not negate with them, so the law failed on generated MaybeUnsupported definitions (about one run in four). The direction cases stay covered by the beaten, tie and instrument-error laws beside it; the same law is already gone from the CI layer
systemfsoftware-maker added a commit that referenced this pull request Oct 6, 2026
…r probes

ce-code-review over #44/#46/#48 (7 Claude reviewers, validator), every finding listed with severity and file:line, plus the verifier's sabotage runs and probes; the earlier round at 74e664d stays in scorecard-stack-findings.md. Nothing applied; Kiro rules each finding
@ryanleecode

Copy link
Copy Markdown
Contributor

Closing: outside the starter's scope (Ryan, 2026-10-06: a proper starter, not a copy of every rat-stack feature).

@ryanleecode ryanleecode closed this Oct 6, 2026
@ryanleecode
ryanleecode deleted the verify/scorecard-model branch October 6, 2026 21:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants