Repository navigation
feat(repo): model scorecard verdicts, the ratchet and the published schema - #44
Closed
systemfsoftware-maker wants to merge 7 commits into
Closed
systemfsoftware-maker wants to merge 7 commits into
systemfsoftware-maker wants to merge 7 commits into
Conversation
systemfsoftware-maker
added this pull request to stack #47
October 5, 2026 22:40
This was referenced Oct 5, 2026
Kiro approved this plan on 2026-10-05. It defines 29 metric families measured on both the starter and rat-stack, the verdict and ratchet rules, and the U1-U9 stack that builds them
…chema The scorecard's pure core. judgeRow decides beaten, not beaten, tie or instrument error for one row: the starter's worst run must be strictly better than rat-stack's best (KTD1), a count must reproduce exactly, and an unsupported rat-stack invariant counts only while its citation holds. compareWithMain is the ratchet (KTD2): it fails on instrument errors, on a row beaten on main and lost on the PR, and on a starter value worse than main's worst run, and treats missing secrets, fork PRs and changed metric definitions as neutral. scorecard.schema.json is the JSON contract CI publishes and /scorecard will render. registry.ts declares the M1-M29 rows as data. The decision modules import only each other, so Deno checks them with no third-party code. Their laws run under vitest inside the sandbox launcher from pnpm-release-management
The instrument gets its own flake: nixpkgs 494ce7f (pnpm 12.9.0, per NixOS/nixpkgs#566850), the pnpm-release-management launcher, and a fixed-output pnpm store built from pnpm-lock.yaml. Installs run in the launcher with no allowed hosts, so nothing reaches the registry. Drops the storeDir override, which empties the fetcher output Review fix for #44 from Kiro
…w counts as beaten Review fix #4: every row declares ratstackSupport. Required rows turn an unsupported rat-stack cell into an instrument error; MayBeUnsupported rows (M12, bar 0) are beaten only when the starter's worst run meets the bar. Review fix #18: assembled cells are keyed by row and side, so a duplicate cell cannot reach assembly and cellFor no longer folds
starter-brainstorm superseded the Lake 1 plan with 2026-10-06-1703; the scorecard plan's source list now points at it, with the unit and ruling anchors it uses there
systemfsoftware-maker
force-pushed
the
verify/scorecard-model
branch
from
October 6, 2026 17:27
9bfdcfd to
b48af9b
Compare
…bsolute bar Since the absolute bar for unsupported rat-stack rows, judging is no longer symmetric under negating runs and flipping direction: the bar does not negate with them, so the law failed on generated MaybeUnsupported definitions (about one run in four). The direction cases stay covered by the beaten, tie and instrument-error laws beside it; the same law is already gone from the CI layer
Contributor
|
Closing: outside the starter's scope (Ryan, 2026-10-06: a proper starter, not a copy of every rat-stack feature). |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Layer 1 of the rat-stack scorecard stack (plan:
docs/plans/2026-10-05-2151-feat-ratstack-scorecard-plan.md, approved by Kiro on 2026-10-05). This layer is inert: nothing runs it until the CI layer (U3) wires it in.What lands
docs(global)commit.judgeRow(src/model/judge-row.workflow.ts) decides each row's verdict under KTD1:compareWithMain(src/model/compare-with-main.workflow.ts) is the KTD2 ratchet.mainthat the PR no longer beats, and on a starter median worse thanmain's worst run. A starter value that disappears also fails.mainartifact, the result is a first baseline.assembleScorecardbuilds the R10 document.renderSummarywrites the job-summary table.scorecard.schema.jsonis the JSON contract.src/metrics/registry.tsholds the M1-M29 rows and their sub-rows as data.Deviations from the plan, declared (CONST-W3)
judge-row,compare-with-main) instead of the plan'sverdictandratchet. Rendering and assembly make no decision, so they are plain pure modules (scorecard-document.ts,summary-table.ts) instead of arender.workflow.ts.dispatch.ts,runs.tsand type-onlycell.ts. These are local, pure, cast-free helpers. KTD3's "import nothing" was there to keep third-party code out, and that still holds.acceptance-examples.test.ts. The*.property.test.tsfiles hold only generated laws.Schema. The laws draw from hand-written fast-check arbitraries inscorecard.arbitrary.ts, because KTD3 keeps the instrument free of third-party runtime code.Verification (run in this session)
All commands were run from
evals/ratstack-scorecard, withSANDBOX_PROJECT=$PWD.The launcher is
github:systemfsoftware/pnpm-release-management/180122866dd537fa728b5563fb1820fbd2af88cc#sandbox(PR #5). pnpm 12.3.4 and Node 24.20.0 come from nixpkgs4975466.I did not run
pnpm mutation(Kiro's rule: mutation never runs locally). The scorecard's own mutation job lands in U3.The first property run found a real defect: an unsupported rat-stack citation let a non-reproducing starter count be judged
beaten, because the starter's runs were only validated against a measured rat-stack. The fix validates each side's runs independently (reproducibleinjudge-row.workflow.ts). A law now pins it.Sabotage
Math.max(...s) < Math.min(...r)to<=injudge-row.workflow.tsturnsa starter whose every run only equals rat-stack's best run is not beatenred (1 failed, 24 passed). Reverted, and the suite is green again.main's best run instead of its worst incompare-with-main.workflow.tsturns AE5, the self-comparison law and the band law red (3 failed, 22 passed). Reverted, and the suite is green again (25 passed).Review fix (Kiro, 2026-10-06): offline install, pnpm 12.9.0
Commit
aaf7caaon this layer; #46 and #48 rebased on it.flake.nixhere (it previously arrived in feat(repo): pin rat-stack and measure the static scorecard rows #46): nixpkgs494ce7fd23ff6a5dff39e1fb11e9b6f2ac74bf25(nixos-unstable, includes pnpm_12: 12.3.4 -> 12.9.0 NixOS/nixpkgs#566850, merged 2026-10-03), the launcher at prm#51801228, andtools-store, a fixed-output pnpm store built frompnpm-lock.yamlthroughmkPnpmWorkspacePackages.--allow-host, which means no network at all (loopback only). The README no longer shows any registry access, andstoreDiris dropped frompnpm-workspace.yamlbecause it empties the fetcher output.pnpm_1212.9.0,nodejs_2424.21.0,deno2.9.7.pnpm 12.9.0 re-resolved the lockfile (
--lockfile-only) to the samepnpm-lock.yaml, byte for byte. On #46 the store hash did not change either, andnix build --rebuild .#tools-storeconfirmed it reproduces under pnpm 12.9.0. That rebuild was needed because a fixed-output path is reused without being rebuilt while its hash is unchanged. #46: 30/30 tests, andscorecard measure --family staticgives the same M23 and M28 numbers. #48: 35/35 tests, actionlint clean, dprint clean,gate:tasksandgate:distgreen.Review fixes (Kiro rulings, 2026-10-06)
Commits on this stack, per ruling. Local gate at
bd50211:pnpm format:checkok,gate:tasksandgate:distgreen,pnpm scorecard:check40/40 (import rules, then journeys),tsc -p evals/ratstack-scorecardclean,deno check src/clean, actionlint clean. Mutation runs in CI only.b937803(#44),849d77b(#46)static-family.test.tstomeasure.test.tsstill runs it from its record; asrc/model/misplaced.test.tsimporting the fixture failsscorecard check(exit 1)9bfdcfd(#44)meetsBar => true: 2 tests red; reverted green9bfdcfd,bd50211aggregateon families with a second ratstack M23 cell:DuplicateCell {"id":"M23","side":"ratstack"}, exit 1849d77b(#46)MissingLauncherRecord, red; edited fixture repo without re-producing:StaleLauncherRecord, red; failing tool givesInstrumentErrorwithexited 3bd50211cache-checkon a poisoned entry:CacheUnusable … M23, M28, exit 1; clean entry exit 0; save is gated on the same checkbd50211pin checklive:PinMoved 54d3560 -> 65e9465;GitRefSchemarequires 40 hex; push uses--force-with-lease=<branch>:<tip>and refuses a tip by another authorbd50211git config --get credential.helperstores the literal${GH_TOKEN};git credential fillyields the secret from env onlybd50211GithubApiErrornaming the request, exit 1 (no silent baseline); 15 s timeout, 3 transient retriesac93e77,bd50211src/harness/sandbox.tsleaves M23/M28 hashes unchanged; editingsrc/families/static.tsmoves both; grading jobs check outbase.sha;instrument previewjob iscontinue-on-error849d77b,bd50211import 'effect'tosrc/families/static.tsfails the import check (exit 1); driver flake has no--allow-net