Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
198 changes: 198 additions & 0 deletions .github/workflows/scorecard.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,198 @@
name: Scorecard

on:
pull_request:
branches: [main]
push:
branches: [main]
schedule:
- cron: "17 4 * * *"
workflow_dispatch:

permissions:
contents: read
actions: read

concurrency:
group: scorecard-${{ github.ref }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}

jobs:
scorecard:
runs-on: ubuntu-latest
timeout-minutes: 360
permissions:
contents: read
actions: read
pull-requests: write
outputs:
mode: ${{ steps.instrument.outputs.mode }}
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
- id: instrument
name: Choose the instrument that grades this run
env:
BASE: ${{ github.event.pull_request.base.sha }}
run: |
if [ "$GITHUB_EVENT_NAME" != pull_request ]; then
printf 'ref=%s\nmode=base\n' "$GITHUB_SHA" >> "$GITHUB_OUTPUT"
exit 0
fi
selector=evals/ratstack-scorecard/ci/instrument-ref.sh
if git cat-file -e "$BASE:$selector" 2>/dev/null; then
git show "$BASE:$selector" | bash -s -- "$BASE" "$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
else
bash "$selector" "$BASE" "$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
fi
cat "$GITHUB_OUTPUT"
- uses: actions/checkout@v7
with:
ref: ${{ steps.instrument.outputs.ref }}
path: .scorecard-instrument
persist-credentials: false
- name: "BOOTSTRAP: self-graded, base had no instrument"
if: steps.instrument.outputs.mode == 'bootstrap'
env:
GH_TOKEN: ${{ github.token }}
PR: ${{ github.event.pull_request.number }}
run: |
banner="# ⚠️ BOOTSTRAP: SELF-GRADED, BASE HAD NO INSTRUMENT

The base commit has no \`.github/workflows/scorecard.yml\`, so this run grades the pull request with its own instrument. Every later pull request is graded by its base's instrument."
printf '%s\n\n' "$banner" >> "$GITHUB_STEP_SUMMARY"
gh pr comment "$PR" --body "$banner"
- uses: cachix/install-nix-action@v31
- name: Allow the launcher's unprivileged user namespaces
run: sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
- name: Build the grading instrument
run: echo "SCORECARD=$(nix build --no-link --print-out-paths ./.scorecard-instrument/evals/ratstack-scorecard#scorecard)/bin/scorecard" >> "$GITHUB_ENV"
- name: Restore rat-stack results
uses: actions/cache/restore@v6
with:
path: .scorecard-cache
key: scorecard-ratstack-${{ github.run_id }}
restore-keys: scorecard-ratstack-
- name: Measure, judge, ratchet and publish
env:
GITHUB_TOKEN: ${{ github.token }}
run: '"$SCORECARD" ci --checkout "$GITHUB_WORKSPACE" --cache .scorecard-cache --out scorecard-out'
- uses: actions/upload-artifact@v7
if: ${{ !cancelled() }}
with:
name: scorecard
path: scorecard-out/scorecard.json
if-no-files-found: error
- name: Save rat-stack results
if: ${{ !cancelled() && github.event_name != 'pull_request' }}
uses: actions/cache/save@v6
with:
path: .scorecard-cache
key: scorecard-ratstack-${{ github.run_id }}

instrument-preview:
name: instrument preview (informational)
needs: scorecard
if: ${{ !cancelled() && github.event_name == 'pull_request' && needs.scorecard.outputs.mode == 'base' }}
continue-on-error: true
runs-on: ubuntu-latest
timeout-minutes: 360
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
- id: changed
name: Does this PR edit the instrument
run: |
if git diff --quiet "${{ github.event.pull_request.base.sha }}" HEAD -- evals/ratstack-scorecard; then
echo "edits=false" >> "$GITHUB_OUTPUT"
echo "This PR leaves the instrument alone; nothing to preview." >> "$GITHUB_STEP_SUMMARY"
else
echo "edits=true" >> "$GITHUB_OUTPUT"
fi
- uses: cachix/install-nix-action@v31
if: steps.changed.outputs.edits == 'true'
- name: Allow the launcher's unprivileged user namespaces
if: steps.changed.outputs.edits == 'true'
run: sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
- uses: actions/download-artifact@v8
if: steps.changed.outputs.edits == 'true'
with:
name: scorecard
path: graded
- name: Grade with this PR's own instrument against the graded scorecard
if: steps.changed.outputs.edits == 'true'
env:
GITHUB_TOKEN: ${{ github.token }}
run: |
scorecard="$(nix build --no-link --print-out-paths ./evals/ratstack-scorecard#scorecard)/bin/scorecard"
"$scorecard" ci --checkout "$GITHUB_WORKSPACE" --cache .scorecard-cache --out preview \
--main graded/scorecard.json --title "Instrument preview (does not gate)"

pin:
if: github.event_name != 'pull_request'
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: cachix/install-nix-action@v31
- name: Allow the launcher's unprivileged user namespaces
run: sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
- name: Build the instrument
run: echo "SCORECARD=$(nix build --no-link --print-out-paths ./evals/ratstack-scorecard#scorecard)/bin/scorecard" >> "$GITHUB_ENV"
- id: check
name: Compare the pin with rat-stack main
env:
GITHUB_TOKEN: ${{ github.token }}
run: |
plan="$("$SCORECARD" pin check)"
echo "$plan" | jq .
echo "tag=$(jq -r ._tag <<<"$plan")" >> "$GITHUB_OUTPUT"
echo "to=$(jq -r '.to // empty' <<<"$plan")" >> "$GITHUB_OUTPUT"
echo "Pin: $(jq -r 'if ._tag == "PinCurrent" then "pin current at \(.commit)" else "rat-stack moved \(.from) -> \(.to)" end' <<<"$plan")" >> "$GITHUB_STEP_SUMMARY"
- id: token
if: steps.check.outputs.tag == 'PinMoved'
name: Pin-bump app token
uses: actions/create-github-app-token@v3
with:
app-id: ${{ vars.SCORECARD_APP_ID }}
private-key: ${{ secrets.SCORECARD_APP_PRIVATE_KEY }}
permission-contents: write
permission-pull-requests: write
# scorecard/pin-rat-stack is bot-owned: the job only ever replaces a tip the bot
# itself wrote, through a lease pinned to that tip, so a human commit there is
# never overwritten. The App token reaches git through a credential helper that
# reads it from the environment, never through argv or a URL.
- name: Open or update the pin-bump PR
if: steps.check.outputs.tag == 'PinMoved'
env:
GH_TOKEN: ${{ steps.token.outputs.token }}
TO: ${{ steps.check.outputs.to }}
BRANCH: scorecard/pin-rat-stack
BOT_EMAIL: scorecard-pin[bot]@users.noreply.github.com
run: |
git config credential.helper "!f() { echo username=x-access-token; echo \"password=\${GH_TOKEN}\"; }; f"
nar_hash="$(nix flake prefetch --json "github:joelhooks/rat-stack/$TO" | jq -r .hash)"
"$SCORECARD" pin write --commit "$TO" --nar-hash "$nar_hash"
lease="$(git ls-remote origin "refs/heads/$BRANCH" | cut -f1)"
if [ -n "$lease" ]; then
git fetch --quiet origin "$lease"
author="$(git log -1 --format=%ae "$lease")"
if [ "$author" != "$BOT_EMAIL" ]; then
echo "::error::$BRANCH tip $lease was written by $author, not the pin bot; refusing to replace it"
exit 1
fi
fi
git config user.name "scorecard-pin[bot]"
git config user.email "$BOT_EMAIL"
git switch -c "$BRANCH"
git commit -m "chore(deps): pin rat-stack to ${TO:0:7} for the scorecard" -- evals/ratstack-scorecard/ratstack.pin.json
git push --force-with-lease="refs/heads/$BRANCH:$lease" origin "HEAD:refs/heads/$BRANCH"
if [ -z "$(gh pr list --head "$BRANCH" --json number --jq '.[0].number')" ]; then
gh pr create --base main --head "$BRANCH" \
--title "chore(deps): pin rat-stack to ${TO:0:7} for the scorecard" \
--body "rat-stack \`main\` moved to \`$TO\`. This PR moves \`evals/ratstack-scorecard/ratstack.pin.json\` only; its scorecard run measures the new rat-stack. Kiro merges."
fi
13 changes: 7 additions & 6 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,13 @@ Starter template for TypeScript / Effect libraries and tools.

## Definition of Done

| ID | Rule | Gate |
| --------- | --------------------------------------------------- | ------------------- |
| `START-1` | Formatting passes dprint with no diffs | `pnpm format:check` |
| `START-2` | Typechecking succeeds workspace-wide with no errors | `pnpm typecheck` |
| `START-3` | All test suites pass | `pnpm test` |
| `START-4` | Full CI validation passes before completion | `pnpm check:ci` |
| ID | Rule | Gate |
| --------- | ---------------------------------------------------------------------------------------------------- | ------------------------- |
| `START-1` | Formatting passes dprint with no diffs | `pnpm format:check` |
| `START-2` | Typechecking succeeds workspace-wide with no errors | `pnpm typecheck` |
| `START-3` | All test suites pass | `pnpm test` |
| `START-4` | Full CI validation passes before completion | `pnpm check:ci` |
| `START-5` | A launcher change counts as done only when its darwin path has run a real offline pnpm install in CI | `check (macos)` job in CI |

Workspace roots: `packages/` holds libraries, `apps/` holds publishable
applications — both are workspace globs in `pnpm-workspace.yaml`. Turbo declares
Expand Down
29 changes: 24 additions & 5 deletions evals/ratstack-scorecard/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,12 +28,31 @@ scorecard=$(nix build --no-link --print-out-paths ./evals/ratstack-scorecard#sco
$scorecard/bin/scorecard measure --family static --out static.json
```

The instrument's own tests run inside the launcher against the same offline store:
The instrument's own checks run with one command: `scorecard check` (also `pnpm scorecard:check` at the repo root, part of `check:ci`). It first refuses any third-party import in the host driver's graph and any import of the journey fixture from `src/`, then runs `scorecard journeys`.

The program is split in two:

- **The host driver** (`src/main.ts` and what it imports) uses only Deno APIs, `node:` builtins and its own files, so `DENO_NO_PACKAGE_JSON=1 deno check src/` type-checks it without third-party code. It starts launcher invocations, mounts their inputs read-only, collects output files and exit codes, and writes the pin file. It decodes and judges nothing. Its only network access is through the launcher.
- **The decide step** (`decide/`, Effect 4 with Schema and `effect/http`) runs as its own launcher invocation from a copy of the grading instrument: `aggregate` (decode the family results and `main`'s scorecard, judge, ratchet, render), `cache-check`, `main-scorecard` and `pin-check`. Only the last two reach the network, and only `api.github.com` and the artifact store (`*.blob.core.windows.net`); GitHub calls time out after 15 s, retry transient failures three times, and fail as a typed `GithubApiError`. Pass or fail is the step's exit code. `pnpm exec tsc` type-checks `decide/` and `journeys/` inside the launcher.

Journeys run in two phases, because the launcher cannot nest. `scorecard journeys` produces every journey declared in `journeys/manifest.json` through the real launcher work and records it to `journeys/__records__/<id>.json`, then runs the one vitest project inside the launcher. A test is a journey because it imports `journeys/launcher-run.ts`; a missing or stale record fails it red.

```sh
cd evals/ratstack-scorecard
nix develop --command sh -c 'SANDBOX_PROJECT=$PWD sandbox --pnpm-store "$SANDBOX_PNPM_STORE" -- pnpm install --frozen-lockfile'
nix develop --command sh -c 'SANDBOX_PROJECT=$PWD sandbox -- pnpm vitest run'
scorecard=$(nix build --no-link --print-out-paths ./evals/ratstack-scorecard#scorecard)
$scorecard/bin/scorecard check
```

`src/main.ts` and everything it imports use only Deno APIs, `node:` builtins and each other, so `DENO_NO_PACKAGE_JSON=1 deno check src/` type-checks the orchestrator and the decision core without third-party code. The Node scripts in `src/tools/` are the only code that loads npm packages, and they only ever run inside the launcher.
The Node scripts in `src/tools/` and everything in `decide/` load npm packages, and they only ever run inside the launcher.

## In CI

`.github/workflows/scorecard.yml` runs on every PR to `main`, every push to `main`, daily, and on demand. All jobs run on GitHub-hosted runners (`ubuntu-latest`): the self-hosted fleet excludes public repositories by design.

The workflow is a thin caller: it chooses the grading instrument, builds it, and runs `scorecard ci`, the one stable entrypoint. Measuring every family, the rat-stack cache, finding `main`'s scorecard, judging, the ratchet and the job summary all happen inside the instrument, so a pull request that adds a step never needs its base to know that step; the step starts grading after it merges.

- **Which instrument grades.** Pushes, the schedule and manual runs use their own commit. A pull request is graded by its base commit's instrument, chosen by `ci/instrument-ref.sh` (run from the base tree whenever the base has it). Only when the base tree has no `.github/workflows/scorecard.yml` at all is the PR head's instrument used; the job summary and a PR comment then say **BOOTSTRAP: self-graded, base had no instrument**. A base commit missing from the clone fails the job instead of falling back to bootstrap. The `instrument-ref` journey proves it on real git history: with the workflow on the base, a PR that edits the instrument to flip a verdict, deletes the instrument, or deletes the workflow is still graded by the base.
- **scorecard** builds the chosen instrument and runs `scorecard ci`. The rat-stack side comes from the cache when its key (rat-stack commit, instrument content hash, nixpkgs rev) matches and `cache-check` accepts it; an entry holding a rat-stack instrument error is refused and measured again. Only pushes, the schedule and manual runs save the cache, only results without instrument errors, and stale keys are pruned. The starter side is measured every run. `scorecard ci` compares with the latest successful `main` run's `scorecard` artifact (none yet means a first baseline), writes the table to the job summary, uploads `scorecard.json`, and fails when the ratchet fails. Re-baselined rows are listed under definition drift with both hashes. A definition hash covers only a row's registry entry and its family's measurement code.
- **instrument preview** (PRs graded by their base that edit the instrument; never gates) runs the PR's own `scorecard ci` against the graded scorecard and lists what it would re-baseline.
- **pin** (not on PRs) compares `ratstack.pin.json` with rat-stack `main`, decoded as a 40-hex SHA. When rat-stack has moved it opens or updates the `scorecard/pin-rat-stack` PR through the pin-bump GitHub App (`vars.SCORECARD_APP_ID`, `secrets.SCORECARD_APP_PRIVATE_KEY`; contents and pull requests write only). The branch is bot-owned: the job refuses to replace a tip another author wrote and pushes with a lease pinned to the tip it read. The token reaches git through a credential helper that reads it from the environment. That PR changes only the pin; Kiro merges it.

`actionlint` is in the dev shell: `nix develop ./evals/ratstack-scorecard --command actionlint .github/workflows/scorecard.yml`.
26 changes: 26 additions & 0 deletions evals/ratstack-scorecard/ci/instrument-ref.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,26 @@
#!/usr/bin/env bash
set -euo pipefail

if [ "$#" -ne 2 ]; then
echo "usage: instrument-ref.sh <base-sha> <head-sha>" >&2
exit 2
fi
base=$1
head=$2

git cat-file -e "${base}^{commit}" 2>/dev/null || {
echo "base commit ${base} is not in this clone; fetch it before choosing an instrument" >&2
exit 1
}
git cat-file -e "${head}^{commit}" 2>/dev/null || {
echo "head commit ${head} is not in this clone" >&2
exit 1
}

if git cat-file -e "${base}:.github/workflows/scorecard.yml" 2>/dev/null; then
echo "ref=${base}"
echo "mode=base"
else
echo "ref=${head}"
echo "mode=bootstrap"
fi
Loading
Loading