Repository navigation
ci(ci): publish the scorecard on every PR to main - #48
systemfsoftware-maker wants to merge 6 commits into
Conversation
a4f756f to
e4623a0
Compare
e4623a0 to
74e664d
Compare
74e664d to
bd50211
Compare
|
2 similar comments
|
|
a6674e7 to
0248251
Compare
|
0248251 to
a308d58
Compare
|
1 similar comment
|
c80ff15 to
402e35e
Compare
|
402e35e to
d17dc8d
Compare
|
d17dc8d to
e96d357
Compare
|
scorecard.yml plans one matrix entry per built family, measures each family in its own GitHub-hosted job with the rat-stack side cached by rat-stack commit, instrument tree hash and nixpkgs rev, then judges and ratchets against the latest successful main run's scorecard artifact. The table goes to the job summary and scorecard.json is uploaded. Off pull requests a pin job opens the rat-stack pin-bump PR through a GitHub App token scoped to contents and pull requests. The CLI gains plan, aggregate, latest-main-run and pin; the family results carry definition hashes so a changed metric re-baselines
… driver Review fix #9 (part): a row's definition hash covers its registry entry and its family's measurement code only; harness, sides, tools, lockfiles and nixpkgs move the cache key instead. J3 becomes a two-phase journey: the host driver runs the real scorecard aggregate CLI and records it; the sandboxed test decodes the record. Restack repair: assemble keys cells by row and side (DuplicateCell refused at decode, #18), the codec decodes ratstackSupport, and the journey manifest decodes through the shared combinators
…ndency driver Review fixes #2, #5, #6, #8 and the rest of #9 under Kiro's rulings. The host driver only starts launcher invocations; decide/ (Effect Schema, effect/http) decodes, judges, ratchets and calls GitHub with a 15 s timeout, three transient retries and a typed GithubApiError. The rat-stack cache never saves or restores an instrument error. pin check decodes main's SHA as 40 hex; the bot-owned pin branch is pushed with a lease and a credential helper. Grading jobs build the instrument from the base tree, re-baselined rows list both hashes, and an informational job previews a PR's own instrument
…, bootstrapping only without one Kiro's bootstrap ruling: the workflow builds the grading instrument and runs scorecard ci, which measures, caches, fetches main's scorecard, judges and publishes inside the instrument. ci/instrument-ref.sh, run from the base tree when the base has it, grades a PR with its base's instrument and uses the PR head only when the base has no scorecard workflow, announcing BOOTSTRAP in the summary and a PR comment. The instrument-ref journey proves on real git history that editing or deleting the instrument or the workflow on a base that has one is still graded by the base. instrumentHash now hashes the grading instrument's own files
With supportedArchitectures and the per-system hash table from the layer below, this layer's lockfile (typescript 7 for the decide step) carries its own linux and darwin hashes; the darwin value is the one the macOS CI leg computed
…change is done Kiro approved this AGENTS.md line on 2026-10-06 after the scorecard's macOS leg exposed a launcher path that only the darwin sandbox exercises (pnpm 12's store lock in /tmp)
e96d357 to
d70e047
Compare
|
|
Closing: outside the starter's scope (Ryan, 2026-10-06: a proper starter, not a copy of every rat-stack feature). |
Layer 3 of the rat-stack scorecard stack (plan:
docs/plans/2026-10-05-2151-feat-ratstack-scorecard-plan.md, unit U3). It sits on #46. This layer wires in what U1 and U2 built: from here on every PR tomainpublishes the scorecard as a job summary and ascorecard.jsonartifact.Evaluator change (GATE1, CONST-E9). This PR adds a judgment surface:
.github/workflows/scorecard.ymlfails a PR on scorecard regressions. Its approval is Kiro's gate ruling of 2026-10-05 (the gate fails on instrument errors, on a starter regression againstmainbeyond the noise band, and on a bin beaten onmainthat the PR no longer beats), given when Kiro approved the plan. The verifier session (starter-verify) owns the instrument; no builder-authored code is graded by a gate written in the same PR.What lands
.github/workflows/scorecard.yml, triggered bypull_requesttomain,pushtomain, a dailyscheduleandworkflow_dispatch. Per Kiro's ruling of 2026-10-05, every job runs on GitHub-hostedubuntu-latest, because the self-hosted fleet's runner group excludes public repositories. Heavy measurement is split across parallel family jobs instead of moving to a bigger runner.planrunsscorecard plan→{ include: [{ family, timeoutMinutes, cacheKey }] }.measure (<family>): a matrix withfail-fast: falseand a per-family timeout from the registry. It allows unprivileged user namespaces (the same sysctl step the launcher's own CI uses), restores the rat-stack side fromactions/cachekeyedscorecard-ratstack/<family>/rat-stack=<sha>/instrument=<tree>/nixpkgs=<rev>, measures that side on a miss, and measures the starter side on every run. Only non-PR runs save the cache: a cache written on a PR is invisible to other PRs.aggregaterunsscorecard latest-main-run, downloads that run'sscorecardartifact when one exists, runsscorecard aggregate, appends the Markdown table to$GITHUB_STEP_SUMMARY, uploadsscorecard.json, and exits 1 when the ratchet fails.pin(not on PRs):scorecard pin check. When rat-stackmainhas moved, it prefetches the narHash, rewritesratstack.pin.jsononly, and opens or updatesscorecard/pin-rat-stackthrough a GitHub App token restricted topermission-contents: writeandpermission-pull-requests: write(security finding D3).src/main.ts) gainsplan,aggregate,latest-main-runandpin check|write. Family results now carry per-row definition hashes, so a changed metric is re-baselined, not compared.plan-pin-bump.workflow.ts(PinCurrent | PinMoved) andcache-key.ts, each with properties.journeys/aggregate.journey.test.ts) runs the realdeno … src/main.ts aggregateagainst hand-written families and a hand-written main scorecard. M23 was beaten on main and is tied on the PR; M28 is still beaten. The journey expects exit 1, stderr naming M23 and not M28, and output that validates againstscorecard.schema.json. Without a main artifact, the same inputs are a passing first baseline.actionlintis in the instrument's dev shell, pinned by itsflake.lock.Evidence (local, this worktree)
check:cialso runspnpm mutation. Per Kiro's rule, mutation never runs locally; this PR touches onlyevals/**and the workflow, and neither is in the root mutation set.The CI path, run locally end to end
The first
latest-main-runattempt failed:listing scorecard runs failed: 404. A workflow file not yet onmainmakes the runs endpoint return 404, which would have failed this PR's own aggregate job. Fixed: a 404 now means "no baseline", and the PR is ratcheted as a first baseline. The decode path was checked against a real runs payload (ci.ymlon main):37383780942.Sabotage
runs: 2(the plan's sabotage)scorecard aggregateexit 1,scorecard: 1 failing rows: M23, row shows**instrument error** | **failed**: rat-stack produced 3 of 2 runsaggregate=0aggregateexits 1 only whenfailures.length > 99× a row beaten on main and tied on the PR fails the job…(1 failed, 3 passed)Notes for Kiro
planPinBumpreturns the moved commit, not commit plus narHash. Prefetching the narHash is I/O, and doing it inside the decision would mean prefetching on every run. Thepinjob prefetches only onPinMovedand passes the hash toscorecard pin write. File names follow the naming lint (plan-pin-bump.workflow.ts,cache-key.ts, which is not a workflow because it decides nothing), not the plan's working names.pinjob needsvars.SCORECARD_APP_IDandsecrets.SCORECARD_APP_PRIVATE_KEYfor the pin-bump App (D3). Until they exist,pinfails only on a run where rat-stack has moved; on every other run it reportsPinCurrent.app-idis used, not v3's newerclient-id.actionlint1.7.12 (the latest release) still requiresapp-idforcreate-github-app-token@v3, and rejectsclient-idas undefined.app-idremains supported in v3.2.0 and prints a deprecation notice.@systemfsoftware/stryker-jsandstryker-js-vitest-runneras flake outputs (your ruling, routed to that repo's session). It gets pinned by flake rev the moment that PR exists.Review fixes (Kiro rulings, 2026-10-06)
Commits on this stack, per ruling. Local gate at
bd50211:pnpm format:checkok,gate:tasksandgate:distgreen,pnpm scorecard:check40/40 (import rules, then journeys),tsc -p evals/ratstack-scorecardclean,deno check src/clean, actionlint clean. Mutation runs in CI only.b937803(#44),849d77b(#46)static-family.test.tstomeasure.test.tsstill runs it from its record; asrc/model/misplaced.test.tsimporting the fixture failsscorecard check(exit 1)9bfdcfd(#44)meetsBar => true: 2 tests red; reverted green9bfdcfd,bd50211aggregateon families with a second ratstack M23 cell:DuplicateCell {"id":"M23","side":"ratstack"}, exit 1849d77b(#46)MissingLauncherRecord, red; edited fixture repo without re-producing:StaleLauncherRecord, red; failing tool givesInstrumentErrorwithexited 3bd50211cache-checkon a poisoned entry:CacheUnusable … M23, M28, exit 1; clean entry exit 0; save is gated on the same checkbd50211pin checklive:PinMoved 54d3560 -> 65e9465;GitRefSchemarequires 40 hex; push uses--force-with-lease=<branch>:<tip>and refuses a tip by another authorbd50211git config --get credential.helperstores the literal${GH_TOKEN};git credential fillyields the secret from env onlybd50211GithubApiErrornaming the request, exit 1 (no silent baseline); 15 s timeout, 3 transient retriesac93e77,bd50211src/harness/sandbox.tsleaves M23/M28 hashes unchanged; editingsrc/families/static.tsmoves both; grading jobs check outbase.sha;instrument previewjob iscontinue-on-error849d77b,bd50211import 'effect'tosrc/families/static.tsfails the import check (exit 1); driver flake has no--allow-netBootstrap ruling (Kiro, 2026-10-06): one entrypoint, base instrument grades
Commit
dd38740.scorecard check45/45,tscanddeno checkclean, actionlint clean.scorecard ci. Measuring every family, the rat-stack cache, findingmain's scorecard (the decide step downloads and unzips the artifact itself), judging, the ratchet and the summary all live in the instrument.plan,latest-main-run,cache-checkand the per-family matrix are gone from the YAML and the CLI.ci/instrument-ref.shruns from the base tree whenever the base has it. It grades with the base instrument; it uses the PR head only when the base has no.github/workflows/scorecard.yml, and then the summary and a PR comment say BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT. A base commit missing from the clone exits 1 rather than bootstrapping.instrument-ref, real git history): with the workflow on the base and the base instrument grading a rowred, a PR that edits the instrument togreen, deletes the instrument, or deletes the workflow is still graded by the base (ref=<base>,mode=base, graded verdictred). Sabotage: making the selector always bootstrap turns those 3 journeys red; reverted, 45/45.scorecard cion this checkout: push event measured both sides, saved the rat-stack side and found nomainbaseline (first baseline, M23 and M28 beaten). PR event with--mainreused the cache without rewriting it, and the ratchet held both rows. Push event with a poisoned cache entry:CacheUnusable … M23, M28, re-measured, clean entry saved, stale key pruned.verify/scorecard-statichas no scorecard workflow. The trusted run is the push run on the branch it merges into, with that branch's own instrument; its link goes here once it exists.