Skip to content

ci(ci): publish the scorecard on every PR to main - #48

Closed
systemfsoftware-maker wants to merge 6 commits into
verify/scorecard-staticfrom
verify/scorecard-ci
Closed

systemfsoftware-maker wants to merge 6 commits into
verify/scorecard-staticfrom
verify/scorecard-ci

Conversation

@systemfsoftware-maker

@systemfsoftware-maker systemfsoftware-maker commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Layer 3 of the rat-stack scorecard stack (plan: docs/plans/2026-10-05-2151-feat-ratstack-scorecard-plan.md, unit U3). It sits on #46. This layer wires in what U1 and U2 built: from here on every PR to main publishes the scorecard as a job summary and a scorecard.json artifact.

Evaluator change (GATE1, CONST-E9). This PR adds a judgment surface: .github/workflows/scorecard.yml fails a PR on scorecard regressions. Its approval is Kiro's gate ruling of 2026-10-05 (the gate fails on instrument errors, on a starter regression against main beyond the noise band, and on a bin beaten on main that the PR no longer beats), given when Kiro approved the plan. The verifier session (starter-verify) owns the instrument; no builder-authored code is graded by a gate written in the same PR.

What lands

  • .github/workflows/scorecard.yml, triggered by pull_request to main, push to main, a daily schedule and workflow_dispatch. Per Kiro's ruling of 2026-10-05, every job runs on GitHub-hosted ubuntu-latest, because the self-hosted fleet's runner group excludes public repositories. Heavy measurement is split across parallel family jobs instead of moving to a bigger runner.
    • plan runs scorecard plan → { include: [{ family, timeoutMinutes, cacheKey }] }.
    • measure (<family>): a matrix with fail-fast: false and a per-family timeout from the registry. It allows unprivileged user namespaces (the same sysctl step the launcher's own CI uses), restores the rat-stack side from actions/cache keyed scorecard-ratstack/<family>/rat-stack=<sha>/instrument=<tree>/nixpkgs=<rev>, measures that side on a miss, and measures the starter side on every run. Only non-PR runs save the cache: a cache written on a PR is invisible to other PRs.
    • aggregate runs scorecard latest-main-run, downloads that run's scorecard artifact when one exists, runs scorecard aggregate, appends the Markdown table to $GITHUB_STEP_SUMMARY, uploads scorecard.json, and exits 1 when the ratchet fails.
    • pin (not on PRs): scorecard pin check. When rat-stack main has moved, it prefetches the narHash, rewrites ratstack.pin.json only, and opens or updates scorecard/pin-rat-stack through a GitHub App token restricted to permission-contents: write and permission-pull-requests: write (security finding D3).
  • CLI (src/main.ts) gains plan, aggregate, latest-main-run and pin check|write. Family results now carry per-row definition hashes, so a changed metric is re-baselined, not compared.
  • Decisions (pure, CC 1): plan-pin-bump.workflow.ts (PinCurrent | PinMoved) and cache-key.ts, each with properties.
  • J3 journey (journeys/aggregate.journey.test.ts) runs the real deno … src/main.ts aggregate against hand-written families and a hand-written main scorecard. M23 was beaten on main and is tied on the PR; M28 is still beaten. The journey expects exit 1, stderr naming M23 and not M28, and output that validates against scorecard.schema.json. Without a main artifact, the same inputs are a passing first baseline.
  • actionlint is in the instrument's dev shell, pinned by its flake.lock.

Evidence (local, this worktree)

$ ./bin/dprint check                                   → fmt-ok
$ DENO_NO_PACKAGE_JSON=1 deno check src/               → no output (clean)
$ nix develop --command sh -c 'SANDBOX_PROJECT=$PWD sandbox -- pnpm vitest run'
 Test Files  11 passed (11)
      Tests  35 passed (35)
$ nix develop --command actionlint ../../.github/workflows/scorecard.yml   → actionlint-ok (exit 0)
$ TURBO_CONCURRENCY=100% pnpm gate:tasks   → 5 cached, 5 total
$ pnpm gate:dist                           → 1 cached, 1 total

check:ci also runs pnpm mutation. Per Kiro's rule, mutation never runs locally; this PR touches only evals/** and the workflow, and neither is in the root mutation set.

The CI path, run locally end to end

$ scorecard plan
{"include":[{"family":"static","timeoutMinutes":20,"cacheKey":"scorecard-ratstack/static/rat-stack=54d356037c99…/instrument=8ce45987…/nixpkgs=4975466d…"}]}
$ scorecard pin check
{"_tag":"PinCurrent","commit":"54d356037c994f89698760a4727be71d0005a087"}
$ scorecard measure --family static --side ratstack --out families/ratstack-static.json
$ scorecard measure --family static --side starter  --out families/starter-static.json
$ scorecard aggregate --families families --out scorecard.json --summary summary.md   → exit 0
| M23 | fence: debt ledger | Suppression directives in tracked source | 219 directives | 0 directives | **beaten** | new |  |
| M28 | install | Distinct name@version packages in the side's lockfile | 957 packages | 392 packages | **beaten** | new |  |
$ GITHUB_REPOSITORY=systemfsoftware/effect-endgame-starter-kit scorecard latest-main-run   → "" exit 0

The first latest-main-run attempt failed: listing scorecard runs failed: 404. A workflow file not yet on main makes the runs endpoint return 404, which would have failed this PR's own aggregate job. Fixed: a 404 now means "no baseline", and the PR is ratcheted as a first baseline. The decode path was checked against a real runs payload (ci.yml on main): 37383780942.

Sabotage

Break Result Reverted
Registry M23 declares runs: 2 (the plan's sabotage) scorecard aggregate exit 1, scorecard: 1 failing rows: M23, row shows **instrument error** | **failed**: rat-stack produced 3 of 2 runs yes; aggregate=0
aggregate exits 1 only when failures.length > 99 J3 red: × a row beaten on main and tied on the PR fails the job… (1 failed, 3 passed) yes; 35/35 green

Notes for Kiro

  • Deviation from the plan's U3 property: planPinBump returns the moved commit, not commit plus narHash. Prefetching the narHash is I/O, and doing it inside the decision would mean prefetching on every run. The pin job prefetches only on PinMoved and passes the hash to scorecard pin write. File names follow the naming lint (plan-pin-bump.workflow.ts, cache-key.ts, which is not a workflow because it decides nothing), not the plan's working names.
  • The pin job needs vars.SCORECARD_APP_ID and secrets.SCORECARD_APP_PRIVATE_KEY for the pin-bump App (D3). Until they exist, pin fails only on a run where rat-stack has moved; on every other run it reports PinCurrent.
  • app-id is used, not v3's newer client-id. actionlint 1.7.12 (the latest release) still requires app-id for create-github-app-token@v3, and rejects client-id as undefined. app-id remains supported in v3.2.0 and prints a deprecation notice.
  • The mutation leg (KTD15) waits on stryker-js-effect shipping @systemfsoftware/stryker-js and stryker-js-vitest-runner as flake outputs (your ruling, routed to that repo's session). It gets pinned by flake rev the moment that PR exists.

Review fixes (Kiro rulings, 2026-10-06)

Commits on this stack, per ruling. Local gate at bd50211: pnpm format:check ok, gate:tasks and gate:dist green, pnpm scorecard:check 40/40 (import rules, then journeys), tsc -p evals/ratstack-scorecard clean, deno check src/ clean, actionlint clean. Mutation runs in CI only.

Ruling Commit Proof
#1+#3 one vitest project, journeys by import b937803 (#44), 849d77b (#46) renaming static-family.test.ts to measure.test.ts still runs it from its record; a src/model/misplaced.test.ts importing the fixture fails scorecard check (exit 1)
#4 absolute bar for Unsupported rows 9bfdcfd (#44) sabotage meetsBar => true: 2 tests red; reverted green
#18 duplicate cell refused 9bfdcfd, bd50211 aggregate on families with a second ratstack M23 cell: DuplicateCell {"id":"M23","side":"ratstack"}, exit 1
#11 static family end to end 849d77b (#46) deleted record: MissingLauncherRecord, red; edited fixture repo without re-producing: StaleLauncherRecord, red; failing tool gives InstrumentError with exited 3
#2 cache never holds InstrumentError bd50211 cache-check on a poisoned entry: CacheUnusable … M23, M28, exit 1; clean entry exit 0; save is gated on the same check
#5 SHA decode + leased bot push bd50211 pin check live: PinMoved 54d3560 -> 65e9465; GitRefSchema requires 40 hex; push uses --force-with-lease=<branch>:<tip> and refuses a tip by another author
#6 token via credential helper bd50211 git config --get credential.helper stores the literal ${GH_TOKEN}; git credential fill yields the secret from env only
#8 HttpClient timeout + retry, typed bd50211 missing repo: GithubApiError naming the request, exit 1 (no silent baseline); 15 s timeout, 3 transient retries
#9 definition hash scope, base-branch instrument, drift, preview ac93e77, bd50211 editing src/harness/sandbox.ts leaves M23/M28 hashes unchanged; editing src/families/static.ts moves both; grading jobs check out base.sha; instrument preview job is continue-on-error
Orchestrator split 849d77b, bd50211 adding import 'effect' to src/families/static.ts fails the import check (exit 1); driver flake has no --allow-net

Bootstrap ruling (Kiro, 2026-10-06): one entrypoint, base instrument grades

Commit dd38740. scorecard check 45/45, tsc and deno check clean, actionlint clean.

  • Thin caller. The workflow runs one command, scorecard ci. Measuring every family, the rat-stack cache, finding main's scorecard (the decide step downloads and unzips the artifact itself), judging, the ratchet and the summary all live in the instrument. plan, latest-main-run, cache-check and the per-family matrix are gone from the YAML and the CLI.
  • Which instrument grades. ci/instrument-ref.sh runs from the base tree whenever the base has it. It grades with the base instrument; it uses the PR head only when the base has no .github/workflows/scorecard.yml, and then the summary and a PR comment say BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT. A base commit missing from the clone exits 1 rather than bootstrapping.
  • Proof it can't be abused (journey instrument-ref, real git history): with the workflow on the base and the base instrument grading a row red, a PR that edits the instrument to green, deletes the instrument, or deletes the workflow is still graded by the base (ref=<base>, mode=base, graded verdict red). Sabotage: making the selector always bootstrap turns those 3 journeys red; reverted, 45/45.
  • instrumentHash now hashes the grading instrument's own files, not the measured checkout's tree, so base-graded results and cache keys name the instrument that produced them.
  • Smoke of scorecard ci on this checkout: push event measured both sides, saved the rat-stack side and found no main baseline (first baseline, M23 and M28 beaten). PR event with --main reused the cache without rewriting it, and the ratchet held both rows. Push event with a poisoned cache entry: CacheUnusable … M23, M28, re-measured, clean entry saved, stale key pruned.
  • This PR is graded in bootstrap mode: its base verify/scorecard-static has no scorecard workflow. The trusted run is the push run on the branch it merges into, with that branch's own instrument; its link goes here once it exists.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

2 similar comments
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

1 similar comment
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

scorecard.yml plans one matrix entry per built family, measures each
family in its own GitHub-hosted job with the rat-stack side cached by
rat-stack commit, instrument tree hash and nixpkgs rev, then judges and
ratchets against the latest successful main run's scorecard artifact.
The table goes to the job summary and scorecard.json is uploaded. Off
pull requests a pin job opens the rat-stack pin-bump PR through a GitHub
App token scoped to contents and pull requests.

The CLI gains plan, aggregate, latest-main-run and pin; the family
results carry definition hashes so a changed metric re-baselines
… driver

Review fix #9 (part): a row's definition hash covers its registry entry and its family's measurement code only; harness, sides, tools, lockfiles and nixpkgs move the cache key instead. J3 becomes a two-phase journey: the host driver runs the real scorecard aggregate CLI and records it; the sandboxed test decodes the record. Restack repair: assemble keys cells by row and side (DuplicateCell refused at decode, #18), the codec decodes ratstackSupport, and the journey manifest decodes through the shared combinators
…ndency driver

Review fixes #2, #5, #6, #8 and the rest of #9 under Kiro's rulings. The host driver only starts launcher invocations; decide/ (Effect Schema, effect/http) decodes, judges, ratchets and calls GitHub with a 15 s timeout, three transient retries and a typed GithubApiError. The rat-stack cache never saves or restores an instrument error. pin check decodes main's SHA as 40 hex; the bot-owned pin branch is pushed with a lease and a credential helper. Grading jobs build the instrument from the base tree, re-baselined rows list both hashes, and an informational job previews a PR's own instrument
…, bootstrapping only without one

Kiro's bootstrap ruling: the workflow builds the grading instrument and runs scorecard ci, which measures, caches, fetches main's scorecard, judges and publishes inside the instrument. ci/instrument-ref.sh, run from the base tree when the base has it, grades a PR with its base's instrument and uses the PR head only when the base has no scorecard workflow, announcing BOOTSTRAP in the summary and a PR comment. The instrument-ref journey proves on real git history that editing or deleting the instrument or the workflow on a base that has one is still graded by the base. instrumentHash now hashes the grading instrument's own files
With supportedArchitectures and the per-system hash table from the layer below, this layer's lockfile (typescript 7 for the decide step) carries its own linux and darwin hashes; the darwin value is the one the macOS CI leg computed
…change is done

Kiro approved this AGENTS.md line on 2026-10-06 after the scorecard's macOS leg exposed a launcher path that only the darwin sandbox exercises (pnpm 12's store lock in /tmp)
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no .github/workflows/scorecard.yml, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument.

systemfsoftware-maker added a commit that referenced this pull request Oct 6, 2026
…r probes

ce-code-review over #44/#46/#48 (7 Claude reviewers, validator), every finding listed with severity and file:line, plus the verifier's sabotage runs and probes; the earlier round at 74e664d stays in scorecard-stack-findings.md. Nothing applied; Kiro rules each finding
@ryanleecode

Copy link
Copy Markdown
Contributor

Closing: outside the starter's scope (Ryan, 2026-10-06: a proper starter, not a copy of every rat-stack feature).

@ryanleecode ryanleecode closed this Oct 6, 2026
@ryanleecode
ryanleecode deleted the verify/scorecard-ci branch October 6, 2026 21:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants