From 2494f1e3c2895cc0c4993aa03043505053f9693b Mon Sep 17 00:00:00 2001 From: Shuxin Lin Date: Wed, 30 Sep 2026 19:21:51 -0400 Subject: [PATCH 1/3] fix(harbor): give each run.sh job its own name, tasks and exit status Four run.sh problems that hold even with a stable runtime image. - A job was named after the model alone, and a resume runs the job's saved config. So `-m "X high" -m "X max"` resumed the high job for max and ran no max trials, and another -p in the same leaderboard directory hit Harbor's lock and could not resume. Jobs are now stirrup_agent____[__]. - Tasks were regenerated on every run, and Harbor refuses to resume a job whose tasks differ from its lock. Any change to the template (comments included), the suite's scenario files or the generator made every job unresumable. Each new job now gets its own copy of the tasks, generated once when it starts and reused on resume. The copies hold the answers, so they go to the repo's gitignored benchmarks/harbor/datasets/jobs/, keyed by the job's path, not into the leaderboard directory. - One shared task folder meant a second run.sh deleted tasks that a running job was still reading. Per-job copies remove the shared folder. - `harbor run ... || true` hid the only failures harbor run reports: it exits 0 when trials fail and non-zero when the job itself cannot run (a rejected config, or StirrupAgent's credential check aborting the job as its first trial starts). That, and a model skipped for an unreachable router, now set a non-zero exit. A resume whose task copy is gone says so. Jobs started under the old names are not resumed; a rerun starts new jobs. Co-Authored-By: Claude Opus 5.5 Signed-off-by: Shuxin Lin --- benchmarks/harbor/README.md | 37 ++++++++----- benchmarks/harbor/run.sh | 100 ++++++++++++++++++++++++++---------- 2 files changed, 98 insertions(+), 39 deletions(-) diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index ad2dec26a..1d87a4bb9 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -18,6 +18,7 @@ benchmarks/harbor/ overlays/private-data.yaml Opt-in read-only mount of a private suite's shared/ (AOB_PRIVATE_DIR) metric.py Category rollups datasets/assetopsbench-/ Generated by generate_tasks.py (gitignored), one per profile + datasets/jobs/-/ run.sh's per-job copy of the tasks (gitignored), same layout dataset.toml Harbor dataset manifest metric.py Copied from benchmarks/harbor/metric.py wosr-1/ One task = one scenario @@ -110,20 +111,32 @@ the number of concurrent trials, and `-r` the runtime image (otherwise `scenario_suite_runner`; 3. builds the code sandbox image and saves it to `~/assetops-code.tar` for the per-trial Docker-in-Docker daemon (`overlays/code-sandbox.yaml`); -4. generates one task per scenario with `--scenario-root` and - `--skip-missing`, skipping profile entries the suite lacks; -5. runs one Harbor job per model at - `/harbor-jobs/stirrup_agent__`, with both overlays and - with credentials loaded from `.env` by `uv run --env-file` into the Harbor - process only. +4. runs one Harbor job per profile, model and reasoning effort at + `/harbor-jobs/stirrup_agent____[__]`, + with both overlays and with credentials loaded from `.env` by + `uv run --env-file` into the Harbor process only; +5. gives each new job its own copy of the tasks, generated with + `--scenario-root` and `--skip-missing` (skipping profile entries the suite + lacks) into `benchmarks/harbor/datasets/jobs/`. The copies hold the + answers, so they stay in the repo's gitignored `datasets/`, not in the + leaderboard directory. Re-running resumes an existing job and finishes only its incomplete trials, -the equivalent of `--skip-existing`. Harbor refuses to resume a job whose tasks -or overlays have changed since it started, such as one from before `run.sh` -mounted only `shared/`; the script says so, and moving that job aside reruns -the model from scratch. The script exits non-zero when any model's job could -not start or resume, so a wrapper sees it. Each trial with the code sandbox runs a -privileged `dind` sidecar, so keep `-n` around 4 on a laptop-sized Docker VM. +the equivalent of `--skip-existing`. A resume reuses the job's own tasks and +settings: later changes to the template, the suite's scenario files or the +generator apply to new jobs only, and so does a changed `-n`. The profile and +effort are part of the job name so that a second effort, or another profile in +the same leaderboard directory, starts its own job instead of resuming the +first. Harbor still refuses to resume a job whose overlays have changed since it +started; the script says so, and moving that job aside reruns the model from +scratch. Several `run.sh` processes can share a leaderboard directory as long +as they run different jobs. + +The script exits non-zero when any model's job could not start or resume, +including a model skipped because its router is unreachable, so a wrapper sees +it. A trial that fails inside a job does not count. Each trial with the code +sandbox runs a privileged `dind` sidecar, so keep `-n` around 4 on a +laptop-sized Docker VM. ## Generating other profiles by hand diff --git a/benchmarks/harbor/run.sh b/benchmarks/harbor/run.sh index 17f2c63a0..51f61b67e 100755 --- a/benchmarks/harbor/run.sh +++ b/benchmarks/harbor/run.sh @@ -22,9 +22,14 @@ # the Harbor process only; StirrupAgent forwards them to the agent phase. They # never enter an image. # -# One Harbor job per model, at LEADERBOARD_DIR/harbor-jobs/stirrup_agent__. +# One Harbor job per profile, model and reasoning effort, at +# LEADERBOARD_DIR/harbor-jobs/stirrup_agent____[__]. # Re-running resumes that job, finishing only the trials it has not completed, -# which is the equivalent of run.sh's --skip-existing. +# which is the equivalent of run.sh's --skip-existing. A resume keeps the +# settings the job started with, so a changed -n applies to new jobs only. +# +# Exits non-zero when a model's job could not start or resume, including a +# model skipped because its router is unreachable. set -euo pipefail @@ -86,7 +91,15 @@ fi code_image=assetops-code:dev code_tar="${AOB_CODE_TAR:-$HOME/assetops-code.tar}" -dataset_dir=benchmarks/harbor/datasets/assetopsbench-suite + +# Each job builds from its own copy of the tasks, generated when the job starts +# and never regenerated. Harbor refuses to resume a job whose tasks differ from +# its lock, so regenerating on every run left a job unresumable after any change +# to the template, the suite's scenario files or the generator; and one shared +# folder let a second run.sh delete tasks a running job was still reading. The +# copies hold every scenario's answers (tests/, solution/), so they stay in the +# repo's gitignored datasets/ rather than beside the results. +tasks_root="$repo_root/benchmarks/harbor/datasets/jobs" # The suite's shared/ data reaches each trial through # overlays/private-data.yaml, a read-only bind mount of this directory's shared/. @@ -167,16 +180,16 @@ if [[ ! -s "$code_tar" ]]; then fi export AOB_CODE_TAR="$code_tar" AOB_CODE_IMAGE="$code_image" -# Regenerate from scratch: the generator overwrites tasks but never removes -# them, so a folder left over from a larger profile would join this run. -rm -rf "$dataset_dir" -uv run python benchmarks/harbor/adapter/generate_tasks.py \ - --scenario-root "$scenario_dir" \ - --profile "$profile" \ - --output-dir "$dataset_dir" \ - --dataset-name assetopsbench/suite \ - --skip-missing \ - --overwrite >/dev/null +# A name safe for a directory: anything but letters, digits and ._- becomes a +# single dash, and a trailing dash is dropped. +slug() { + local name + name="$(printf '%s' "$1" | tr -c 'A-Za-z0-9._-' '-' | tr -s '-')" + printf '%s' "${name%-}" +} + +profile_name="$(basename "$profile")" +profile_slug="$(slug "${profile_name%.*}")" # Fail fast when a model's router is unreachable. Otherwise every trial builds # its containers, loads its data, retries the model for minutes and exits 1, @@ -203,17 +216,26 @@ except Exception as exc: PY } -# Non-zero when a model's job could not run or resume. A trial that fails -# inside a job does not count: Harbor records it and the loop moves on. +# Non-zero when a model's job could not start or resume, or was skipped. A +# trial that fails inside a job does not count: Harbor records it and the loop +# moves on. status=0 for model_config in "${model_configs[@]}"; do read -r model_id reasoning_effort <<< "$model_config" [[ -z "${model_id:-}" ]] && continue - model_slug="$(printf '%s' "$model_id" | tr -c 'A-Za-z0-9._-' '-' | tr -s '-')" - job_name="stirrup_agent__${model_slug%-}" + # The profile and effort are in the name because a resume runs the job's own + # saved settings. Named after the model alone, a second effort or another + # profile resumed the first job instead of starting its own. + job_name="stirrup_agent__${profile_slug}__$(slug "$model_id")" + if [[ -n "${reasoning_effort:-}" ]]; then + job_name+="__$(slug "$reasoning_effort")" + fi job_path="$jobs_dir/$job_name" + # Keyed by the job's full path, so the same job name under another + # LEADERBOARD_DIR gets its own copy. + tasks_dir="$tasks_root/$job_name-$(printf '%s' "$job_path" | cksum | cut -d' ' -f1)" # The runtime image a job started on, beside the job rather than in it. # Harbor's resume lock covers the task files but not the base they build # FROM, so without this a resume with another -r would mix two images in @@ -222,6 +244,7 @@ for model_config in "${model_configs[@]}"; do if ! router_reachable "$model_id"; then echo "Skipping $model_id: its router is unreachable" >&2 + status=1 continue fi @@ -236,21 +259,39 @@ for model_config in "${model_configs[@]}"; do status=1 continue fi + if [[ ! -d "$tasks_dir" ]]; then + printf 'The tasks %s started with are gone (%s).\n' "$job_path" "$tasks_dir" >&2 + printf 'Move the job aside to rerun %s from scratch.\n' "$model_id" >&2 + status=1 + continue + fi # Drop trials whose agent crashed (e.g. the model was unreachable) so # they run again; scored trials are kept, as run.sh's --skip-existing # kept scenarios that already had a trajectory. - # Harbor refuses to resume once the tasks or overlays differ from the - # job's lock, e.g. a job started before run.sh switched to the shared/ - # mount. Say so rather than skip the model silently. + # Harbor refuses to resume once the overlays differ from the job's lock. + # The tasks cannot differ: the job builds from its own copy. Say so rather + # than skip the model silently. if ! uv run --env-file "$env_file" harbor jobs resume -p "$job_path" \ --filter-error-type NonZeroAgentExitCodeError; then - printf 'Could not resume %s. If its tasks or overlays changed since it\n' "$job_path" >&2 - printf 'started, move it aside to rerun %s from scratch.\n' "$model_id" >&2 + printf 'Could not resume %s. If its overlays changed since it started,\n' "$job_path" >&2 + printf 'move it aside to rerun %s from scratch.\n' "$model_id" >&2 status=1 fi continue fi + # A new job, so its own tasks from scratch: the generator overwrites tasks + # but never removes them, so a folder left from an attempt that never + # started would join this one. + rm -rf "$tasks_dir" + uv run python benchmarks/harbor/adapter/generate_tasks.py \ + --scenario-root "$scenario_dir" \ + --profile "$profile" \ + --output-dir "$tasks_dir" \ + --dataset-name assetopsbench/suite \ + --skip-missing \ + --overwrite >/dev/null + mkdir -p "$jobs_dir" printf '%s\t%s\n' "$runtime_id" "$runtime_image" >"$image_record" @@ -259,10 +300,12 @@ for model_config in "${model_configs[@]}"; do effort_args=(--ak "reasoning_effort=$reasoning_effort") fi - # --continue-on-error equivalent: a failed trial is recorded in the job and - # the loop moves on to the next model. - uv run --env-file "$env_file" harbor run -y \ - -p "$dataset_dir" \ + # harbor run exits 0 when trials fail, recording them in the job, so the loop + # moves on to the next model as run.sh's --continue-on-error did. Non-zero + # means the job itself could not run, e.g. a rejected config or StirrupAgent + # refusing its credentials, which aborts the job as the first trial starts. + if ! uv run --env-file "$env_file" harbor run -y \ + -p "$tasks_dir" \ --agent assetops_harbor.stirrup:StirrupAgent \ --model "$model_id" \ --ak code_enabled=true \ @@ -274,7 +317,10 @@ for model_config in "${model_configs[@]}"; do --extra-docker-compose benchmarks/harbor/overlays/code-sandbox.yaml \ --n-concurrent "$n_concurrent" \ --job-name "$job_name" \ - -o "$jobs_dir" || true + -o "$jobs_dir"; then + printf 'Harbor could not run %s; see the error above.\n' "$job_path" >&2 + status=1 + fi done exit "$status" From 66eb514444b852837f3e13db5d4397200a82f266 Mon Sep 17 00:00:00 2001 From: Shuxin Lin Date: Wed, 30 Sep 2026 19:34:34 -0400 Subject: [PATCH 2/3] fix(harbor): run.sh retries, code tar, key checks, env file, lock, paths The rest of the run.sh problems found reviewing #558. - Resume reran only NonZeroAgentExitCodeError. Harbor matches exact class names, so its subclasses (rate limits, network, auth), Ctrl-C, and environment and verifier failures stayed failed. run.sh now lists every type that is not the model's own work; timeouts, context and output limits and safety refusals stay as results. test_run_sh.py checks the list against the installed Harbor and fails when a new agent error subclass is unplaced. - The code sandbox tar was built once and reused forever. It is now built each run (cached), saved once per image id under AOB_CODE_TAR_DIR (default ~/.cache/assetopsbench), and recorded per job; a resume loads the job's own tar, so a job never switches code image and a rebuild never rewrites a tar in use. - The router check only probed the base URL. It now rejects a key the router refuses (401/403 from /models), also checks FMSR_MODEL_ID's router, and skips a model with no router prefix unless FMSR_MODEL_ID names one, since FMSR would reject it and fmsr scenarios would run without generate_failure_modes. - ENV_FILE was not the only credential source: StirrupAgent also loaded the repo's .env and filled its gaps. StirrupAgent now reads AOB_ENV_FILE instead of searching when it is set, and run.sh sets it. - Two run.sh processes could work on one job. A per-job lock (mkdir, with the owner's PID so a dead owner's lock is taken over) makes the second skip it. - Relative -s, -l, -p, ENV_FILE and AOB_CODE_TAR_DIR resolved against the repo root. They now resolve against the caller's directory; the defaults are still the repo's. Co-Authored-By: Claude Opus 5.5 Signed-off-by: Shuxin Lin --- benchmarks/harbor/CODE-SANDBOX.md | 10 +- benchmarks/harbor/README.md | 48 +++- benchmarks/harbor/run.sh | 287 +++++++++++++++++----- src/assetops_harbor/stirrup.py | 12 +- src/assetops_harbor/tests/test_run_sh.py | 89 +++++++ src/assetops_harbor/tests/test_stirrup.py | 23 ++ 6 files changed, 393 insertions(+), 76 deletions(-) create mode 100644 src/assetops_harbor/tests/test_run_sh.py diff --git a/benchmarks/harbor/CODE-SANDBOX.md b/benchmarks/harbor/CODE-SANDBOX.md index 4e103a3d2..a0b692444 100644 --- a/benchmarks/harbor/CODE-SANDBOX.md +++ b/benchmarks/harbor/CODE-SANDBOX.md @@ -70,8 +70,9 @@ well below what you use for the tools-only arm. ## The easy path: run.sh `benchmarks/harbor/run.sh` already does all of it. It builds the code image, -saves the tar, exports both variables, generates the dataset and passes the -overlay with the four required `--ak` flags. +saves the tar (one per image id, in `~/.cache/assetopsbench`), points each job +at its own tar, generates the dataset and passes the overlay with the four +required `--ak` flags. ```bash ./benchmarks/harbor/run.sh \ @@ -197,8 +198,9 @@ one this overlay is known to work on. **"no AOB_CODE_TAR; the daemon will pull ..."** in the loader log. Expected when you went the registry route. If you meant to use a tar, the path in -`AOB_CODE_TAR` is wrong or the file is empty. It defaults to -`$HOME/assetops-code.tar` under `run.sh`. +`AOB_CODE_TAR` is wrong or the file is empty. Under `run.sh` it is the job's +own tar in `~/.cache/assetopsbench`, whose path is in `.code-tar` beside +the job. **Runs are much slower than the tools-only arm.** Each trial pays for a container and an image load. Lower `--n-concurrent`. diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index 1d87a4bb9..9c5fd63f5 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -94,7 +94,19 @@ bash benchmarks/harbor/run.sh \ `-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), `-n` the number of concurrent trials, and `-r` the runtime image (otherwise `AOB_RUNTIME_IMAGE` from the shell, then from `.env`, then -`assetopsbench/runtime:dev`). The script: +`assetopsbench/runtime:dev`). Credentials come from `ENV_FILE`, the repo's +`.env` by default, and only from that file: with another file, the repo's +`.env` fills none of its gaps. Relative paths in `-s`, `-l`, `-p`, `ENV_FILE` +and `AOB_CODE_TAR_DIR` are relative to the directory you run the script from. + +Before each model, the script checks that it can be served, and skips it +otherwise. The model's router and `FMSR_MODEL_ID`'s must answer and must not +reject their key (a 401 or 403 from `/models`). A model with no `litellm_proxy/` +or `tokenrouter/` prefix also needs `FMSR_MODEL_ID` set to one, because the FMSR +server would otherwise reject it and every fmsr scenario would run without +`generate_failure_modes`. + +The script: 1. resolves the runtime image, which every task image builds FROM. A published image, one whose local copy came from that registry or that is not local at @@ -109,8 +121,14 @@ the number of concurrent trials, and `-r` the runtime image (otherwise `overlays/private-data.yaml` mounts its `shared/` at `/opt/suite/scenarios_data/shared`. Only `init_data.py` reads it, as in `scenario_suite_runner`; -3. builds the code sandbox image and saves it to `~/assetops-code.tar` for the - per-trial Docker-in-Docker daemon (`overlays/code-sandbox.yaml`); +3. builds the code sandbox image for the per-trial Docker-in-Docker daemon + (`overlays/code-sandbox.yaml`) on every run, which the build cache makes + cheap, so a change to `Dockerfile.code` is picked up. It saves the image as + a tar named after its id, in `AOB_CODE_TAR_DIR` (default + `~/.cache/assetopsbench`), once per image, and never rewrites it. Each job + records its tar, and a resume loads that one, so a job keeps one code image + from start to finish. `AOB_CODE_TAR` from your shell is not used. Old tars + stay until you delete them, and `~/assetops-code.tar` is no longer used; 4. runs one Harbor job per profile, model and reasoning effort at `/harbor-jobs/stirrup_agent____[__]`, with both overlays and with credentials loaded from `.env` by @@ -129,14 +147,23 @@ effort are part of the job name so that a second effort, or another profile in the same leaderboard directory, starts its own job instead of resuming the first. Harbor still refuses to resume a job whose overlays have changed since it started; the script says so, and moving that job aside reruns the model from -scratch. Several `run.sh` processes can share a leaderboard directory as long -as they run different jobs. +scratch. + +A resume runs again every trial that failed for a reason other than the model's +own work: its API or the network, the environment, the verifier, or Ctrl-C. +Harbor matches exact class names, so `run.sh` lists each one +(`retry_error_types`), and `src/assetops_harbor/tests/test_run_sh.py` checks the +list against the installed Harbor. A timeout, an exceeded context window or +output limit, and a safety refusal are kept as results. + +Several `run.sh` processes can share a leaderboard directory. Only one works on +a given job at a time: the other skips it and names the lock +(`.lock`). A lock whose process has died is taken over. The script exits non-zero when any model's job could not start or resume, -including a model skipped because its router is unreachable, so a wrapper sees -it. A trial that fails inside a job does not count. Each trial with the code -sandbox runs a privileged `dind` sidecar, so keep `-n` around 4 on a -laptop-sized Docker VM. +including a model skipped by the checks above, so a wrapper sees it. A trial +that fails inside a job does not count. Each trial with the code sandbox runs a +privileged `dind` sidecar, so keep `-n` around 4 on a laptop-sized Docker VM. ## Generating other profiles by hand @@ -320,7 +347,8 @@ the runner's behaviour and the SDK default it compensates for. The agent forwards credentials from the Harbor process into the container for the agent phase only. Putting them in the repo's `.env` is enough when you run from the repo root: `StirrupAgent` loads the nearest `.env` above the working -directory before it checks for credentials. +directory before it checks for credentials. `AOB_ENV_FILE`, when set, names the +one file it loads instead; `run.sh` sets it to its `ENV_FILE`. ```bash uv run harbor run -p benchmarks/harbor/datasets/assetopsbench-open \ diff --git a/benchmarks/harbor/run.sh b/benchmarks/harbor/run.sh index 51f61b67e..d6d6f6d4f 100755 --- a/benchmarks/harbor/run.sh +++ b/benchmarks/harbor/run.sh @@ -18,18 +18,23 @@ # or published, passed as -r (or AOB_RUNTIME_IMAGE in the shell or ENV_FILE), # e.g. -r quay.io/assetopsbench/runtime:dev. # -# Credentials are read from ENV_FILE (default .env) by `uv run --env-file`, into -# the Harbor process only; StirrupAgent forwards them to the agent phase. They -# never enter an image. +# Credentials are read from ENV_FILE (default: the repo's .env) by +# `uv run --env-file`, into the Harbor process only; StirrupAgent forwards them +# to the agent phase. They never enter an image. ENV_FILE is the only file read: +# with another file, the repo's .env fills none of its gaps. +# +# Relative paths in -s, -l, -p, ENV_FILE and AOB_CODE_TAR_DIR are relative to +# the directory the script is run from. # # One Harbor job per profile, model and reasoning effort, at # LEADERBOARD_DIR/harbor-jobs/stirrup_agent____[__]. # Re-running resumes that job, finishing only the trials it has not completed, # which is the equivalent of run.sh's --skip-existing. A resume keeps the -# settings the job started with, so a changed -n applies to new jobs only. +# settings the job started with, so a changed -n applies to new jobs only. Only +# one run.sh works on a job at a time; another skips it. # # Exits non-zero when a model's job could not start or resume, including a -# model skipped because its router is unreachable. +# model skipped by the checks before it. set -euo pipefail @@ -40,8 +45,8 @@ usage() { scenario_dir="${SCENARIO_DIR:-}" leaderboard_dir="${LEADERBOARD_DIR:-}" n_concurrent="${N_CONCURRENT:-4}" -profile="${PROFILE:-benchmarks/scenario_suite/all.yaml}" -env_file="${ENV_FILE:-.env}" +profile="${PROFILE:-}" +env_file="${ENV_FILE:-}" runtime_image="${AOB_RUNTIME_IMAGE:-}" model_configs=() @@ -77,20 +82,53 @@ if [[ -z "${model_configs[*]+set}" ]]; then fi repo_root="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)" + +# Paths the caller gives are relative to where the script was run, as for any +# command; the defaults are the repo's own. Resolved before the cd below. +caller_dir="$PWD" +absolute() { + case "$1" in + /*) printf '%s' "$1" ;; + *) printf '%s/%s' "$caller_dir" "$1" ;; + esac +} +scenario_dir="$(absolute "$scenario_dir")" +leaderboard_dir="$(absolute "$leaderboard_dir")" +if [[ -n "$profile" ]]; then + profile="$(absolute "$profile")" +else + profile="$repo_root/benchmarks/scenario_suite/all.yaml" +fi +if [[ -n "$env_file" ]]; then + env_file="$(absolute "$env_file")" +else + env_file="$repo_root/.env" +fi +code_tar_dir="$(absolute "${AOB_CODE_TAR_DIR:-$HOME/.cache/assetopsbench}")" + cd "$repo_root" +if [[ ! -d "$scenario_dir" ]]; then + printf 'Scenario directory not found: %s\n' "$scenario_dir" >&2 + exit 2 +fi scenario_dir="$(cd "$scenario_dir" && pwd)" +if [[ ! -f "$profile" ]]; then + printf 'Profile not found: %s\n' "$profile" >&2 + exit 2 +fi mkdir -p "$leaderboard_dir" leaderboard_dir="$(cd "$leaderboard_dir" && pwd)" jobs_dir="$leaderboard_dir/harbor-jobs" +mkdir -p "$jobs_dir" if [[ ! -f "$env_file" ]]; then printf 'Credentials file not found: %s (set ENV_FILE)\n' "$env_file" >&2 exit 2 fi - -code_image=assetops-code:dev -code_tar="${AOB_CODE_TAR:-$HOME/assetops-code.tar}" +# StirrupAgent reads this file in place of the nearest .env, so the repo's .env +# cannot fill the gaps of another ENV_FILE. +export AOB_ENV_FILE="$env_file" # Each job builds from its own copy of the tasks, generated when the job starts # and never regenerated. Harbor refuses to resume a job whose tasks differ from @@ -101,6 +139,18 @@ code_tar="${AOB_CODE_TAR:-$HOME/assetops-code.tar}" # repo's gitignored datasets/ rather than beside the results. tasks_root="$repo_root/benchmarks/harbor/datasets/jobs" +# Everything the script creates for itself goes on exit: the runtime image pin, +# a code tar it did not finish writing, and the lock of the job it was on. +runtime_pin="" +code_tar_partial="" +job_lock="" +cleanup() { + if [[ -n "$runtime_pin" ]]; then docker rmi "$runtime_pin" >/dev/null 2>&1 || true; fi + if [[ -n "$code_tar_partial" ]]; then rm -f "$code_tar_partial"; fi + if [[ -n "$job_lock" ]]; then rm -rf "$job_lock"; fi +} +trap cleanup EXIT + # The suite's shared/ data reaches each trial through # overlays/private-data.yaml, a read-only bind mount of this directory's shared/. # Compose reads the variable on `harbor run` and again on `harbor jobs resume`. @@ -165,20 +215,34 @@ runtime_id="$(docker image inspect --format '{{.Id}}' "$runtime_image")" runtime_id="${runtime_id#sha256:}" runtime_pin="aob-runtime-pin:${runtime_id:0:12}-$$" docker tag "$runtime_image" "$runtime_pin" -trap 'docker rmi "$runtime_pin" >/dev/null 2>&1 || true' EXIT export AOB_RUNTIME_IMAGE="$runtime_pin" printf 'Runtime image: %s (%s)\n' "$runtime_image" "${runtime_id:0:12}" # The code sandbox image, as a tar each trial's Docker-in-Docker daemon loads -# (benchmarks/harbor/overlays/code-sandbox.yaml). +# (benchmarks/harbor/overlays/code-sandbox.yaml). It is built on every run, +# which the build cache makes cheap, so a change to Dockerfile.code is picked +# up. The tar is named after the image id and written once, so a new build never +# rewrites the file a running job's trials load. Each job records its tar and a +# resume loads that one; a new job loads this run's. AOB_CODE_TAR from the +# environment is ignored here, since every job sets its own. Old tars stay in +# AOB_CODE_TAR_DIR until removed. +code_image=assetops-code:dev +docker build -q -t "$code_image" \ + -f src/agent/stirrup_agent/Dockerfile.code src/agent/stirrup_agent >/dev/null +code_id="$(docker image inspect --format '{{.Id}}' "$code_image")" +code_id="${code_id#sha256:}" +code_tar="$code_tar_dir/assetops-code-${code_id:0:12}.tar" if [[ ! -s "$code_tar" ]]; then - echo "Building $code_image and saving it to $code_tar" - docker build -q -t "$code_image" \ - -f src/agent/stirrup_agent/Dockerfile.code src/agent/stirrup_agent - docker save "$code_image" -o "$code_tar" - chmod 644 "$code_tar" + mkdir -p "$code_tar_dir" + echo "Saving $code_image to $code_tar" + code_tar_partial="$code_tar.partial.$$" + docker save "$code_image" -o "$code_tar_partial" + chmod 644 "$code_tar_partial" + mv "$code_tar_partial" "$code_tar" + code_tar_partial="" fi -export AOB_CODE_TAR="$code_tar" AOB_CODE_IMAGE="$code_image" +export AOB_CODE_IMAGE="$code_image" +printf 'Code image: %s (%s)\n' "$code_image" "${code_id:0:12}" # A name safe for a directory: anything but letters, digits and ._- becomes a # single dash, and a trailing dash is dropped. @@ -191,31 +255,117 @@ slug() { profile_name="$(basename "$profile")" profile_slug="$(slug "${profile_name%.*}")" -# Fail fast when a model's router is unreachable. Otherwise every trial builds -# its containers, loads its data, retries the model for minutes and exits 1, -# which is how a dropped VPN turns into a job of failed trials. -router_reachable() { - local base_var - case "$1" in - litellm_proxy/*) base_var=LITELLM_BASE_URL ;; - tokenrouter/*) base_var=TOKENROUTER_BASE_URL ;; - *) return 0 ;; - esac - uv run --env-file "$env_file" python - "$base_var" <<'PY' -import os, sys, urllib.error, urllib.request -name = sys.argv[1] -url = os.environ.get(name, "") -if not url: - sys.exit(f"{name} is not set in the env file") -try: - urllib.request.urlopen(url, timeout=15) -except urllib.error.HTTPError: - pass # the router answered; any HTTP status proves it is reachable -except Exception as exc: - sys.exit(f"cannot reach {name} ({exc}); check the VPN or network") +# Fail fast when a model cannot be served. Otherwise every trial builds its +# containers, loads its data, retries the model for minutes and exits 1, which +# is how a dropped VPN or an expired key turns into a job of failed trials. +# +# Each router in use must answer, and must not reject its key: GET /models +# answers 401 or 403 to a bad key, and any other answer, 404 included, passes. +# The routers in use are the model's and FMSR_MODEL_ID's. A model with no router +# prefix is not probed, but it needs FMSR_MODEL_ID: the FMSR server otherwise +# runs generate_failure_modes on the agent's model, accepts only router models, +# and every fmsr scenario would run without that tool. +check_model() { + uv run --env-file "$env_file" python - "$1" <<'PY' +import os +import sys +import urllib.error +import urllib.request + +# src/llm/routers.py PROXY_ROUTERS; src/assetops_harbor/tests checks they match. +ROUTERS = { + "litellm_proxy/": ("LITELLM_BASE_URL", "LITELLM_API_KEY"), + "tokenrouter/": ("TOKENROUTER_BASE_URL", "TOKENROUTER_API_KEY"), +} + + +def router(model): + return next((prefix for prefix in ROUTERS if model.startswith(prefix)), None) + + +model = sys.argv[1] +fmsr_model = os.environ.get("FMSR_MODEL_ID", "").strip() +if fmsr_model and not router(fmsr_model): + sys.exit(f"FMSR_MODEL_ID={fmsr_model} needs a {' or '.join(ROUTERS)} prefix") +if not fmsr_model and not router(model): + sys.exit( + f"{model} has no {' or '.join(ROUTERS)} prefix, so the FMSR server would " + "reject it; set FMSR_MODEL_ID to a router model" + ) + +for prefix in dict.fromkeys(p for p in (router(model), router(fmsr_model)) if p): + base_var, key_var = ROUTERS[prefix] + base, key = os.environ.get(base_var, ""), os.environ.get(key_var, "") + missing = [name for name, value in ((base_var, base), (key_var, key)) if not value] + if missing: + sys.exit(f"{' and '.join(missing)} not set for {prefix} models") + request = urllib.request.Request( + base.rstrip("/") + "/models", headers={"Authorization": f"Bearer {key}"} + ) + try: + urllib.request.urlopen(request, timeout=15) + except urllib.error.HTTPError as exc: + if exc.code in (401, 403): + sys.exit(f"{base_var} rejected {key_var} (HTTP {exc.code})") + except Exception as exc: + sys.exit(f"cannot reach {base_var} ({exc}); check the VPN or network") PY } +# One run.sh per job at a time: two resuming the same job would run its trials +# twice into the same directories. mkdir is atomic, and the lock holds its +# owner's PID, so a lock left by a run.sh that died is taken over. +lock_job() { + local lock="$1" owner + if mkdir "$lock" 2>/dev/null; then + printf '%s\n' "$$" >"$lock/pid" + return 0 + fi + owner="$(cat "$lock/pid" 2>/dev/null || true)" + if [[ -z "$owner" ]] || kill -0 "$owner" 2>/dev/null; then + return 1 + fi + rm -rf "$lock" + mkdir "$lock" 2>/dev/null || return 1 + printf '%s\n' "$$" >"$lock/pid" +} + +# Trials a resume runs again: those that failed for a reason other than the +# model's own work, i.e. its API or the network, the environment, the verifier, +# or Ctrl-C. Harbor matches the exact class name, so every subclass is listed. +# Kept as results: AgentTimeoutError, ContextWindowExceededError, +# OutputTokenExceededError and AgentSafetyRefusalError. Names from Harbor 0.23; +# src/assetops_harbor/tests/test_run_sh.py checks them against the installed one. +retry_error_types=( + CancelledError + NonZeroAgentExitCodeError + ApiError + ApiRateLimitError + ApiUsageLimitError + ApiInternalServerError + ApiOverloadedError + ApiConnectionClosedError + ApiResponseStalledError + UnknownApiError + ApiProviderResourceNotFoundError + AgentAuthenticationError + ModelNotFoundError + NetworkConnectionError + AgentSetupTimeoutError + EnvironmentStartTimeoutError + HealthcheckError + VerifierTimeoutError + RewardFileNotFoundError + RewardFileEmptyError + VerifierOutputParseError + AddTestsDirError + DownloadVerifierDirError +) +retry_filters=() +for error_type in "${retry_error_types[@]}"; do + retry_filters+=(--filter-error-type "$error_type") +done + # Non-zero when a model's job could not start or resume, or was skipped. A # trial that fails inside a job does not count: Harbor records it and the loop # moves on. @@ -236,47 +386,60 @@ for model_config in "${model_configs[@]}"; do # Keyed by the job's full path, so the same job name under another # LEADERBOARD_DIR gets its own copy. tasks_dir="$tasks_root/$job_name-$(printf '%s' "$job_path" | cksum | cut -d' ' -f1)" - # The runtime image a job started on, beside the job rather than in it. - # Harbor's resume lock covers the task files but not the base they build - # FROM, so without this a resume with another -r would mix two images in - # one job. + # The runtime image and code tar a job started on, beside the job rather than + # in it. Harbor's resume lock covers the task files but not the base they + # build FROM, so without the first a resume with another -r would mix two + # images in one job; the second makes a resume load the job's own code image. image_record="$jobs_dir/$job_name.runtime-image" + code_record="$jobs_dir/$job_name.code-tar" + + if ! check_model "$model_id"; then + echo "Skipping $model_id" >&2 + status=1 + continue + fi - if ! router_reachable "$model_id"; then - echo "Skipping $model_id: its router is unreachable" >&2 + if ! lock_job "$job_path.lock"; then + printf 'Skipping %s: another run.sh (PID %s) is working on it.\n' \ + "$job_path" "$(cat "$job_path.lock/pid" 2>/dev/null || echo unknown)" >&2 + printf 'If none is, remove %s.\n' "$job_path.lock" >&2 status=1 continue fi + job_lock="$job_path.lock" echo "Running $model_id with reasoning effort ${reasoning_effort:-default} -> $job_path" if [[ -f "$job_path/config.json" ]]; then + job_code_tar="$code_tar" + [[ -f "$code_record" ]] && job_code_tar="$(cat "$code_record")" if [[ -f "$image_record" ]] && [[ "$(cut -f1 "$image_record")" != "$runtime_id" ]]; then printf '%s started on runtime image %s, not %s (%s).\n' \ "$job_path" "$(cut -f2 "$image_record")" "$runtime_image" "${runtime_id:0:12}" >&2 printf 'Pass that image as -r to finish it, or move the job aside to rerun %s.\n' \ "$model_id" >&2 status=1 - continue - fi - if [[ ! -d "$tasks_dir" ]]; then + elif [[ ! -d "$tasks_dir" ]]; then printf 'The tasks %s started with are gone (%s).\n' "$job_path" "$tasks_dir" >&2 printf 'Move the job aside to rerun %s from scratch.\n' "$model_id" >&2 status=1 - continue - fi - # Drop trials whose agent crashed (e.g. the model was unreachable) so - # they run again; scored trials are kept, as run.sh's --skip-existing - # kept scenarios that already had a trajectory. - # Harbor refuses to resume once the overlays differ from the job's lock. - # The tasks cannot differ: the job builds from its own copy. Say so rather - # than skip the model silently. - if ! uv run --env-file "$env_file" harbor jobs resume -p "$job_path" \ - --filter-error-type NonZeroAgentExitCodeError; then + elif [[ ! -s "$job_code_tar" ]]; then + printf 'The code image %s started with is gone (%s).\n' "$job_path" "$job_code_tar" >&2 + printf 'Move the job aside to rerun %s from scratch.\n' "$model_id" >&2 + status=1 + # Drop the trials in retry_error_types so they run again; scored trials + # are kept, as run.sh's --skip-existing kept scenarios that already had a + # trajectory. Harbor refuses to resume once the overlays differ from the + # job's lock. The tasks cannot differ: the job builds from its own copy. + # Say so rather than skip the model silently. + elif ! AOB_CODE_TAR="$job_code_tar" uv run --env-file "$env_file" \ + harbor jobs resume -p "$job_path" "${retry_filters[@]}"; then printf 'Could not resume %s. If its overlays changed since it started,\n' "$job_path" >&2 printf 'move it aside to rerun %s from scratch.\n' "$model_id" >&2 status=1 fi + rm -rf "$job_lock" + job_lock="" continue fi @@ -292,8 +455,8 @@ for model_config in "${model_configs[@]}"; do --skip-missing \ --overwrite >/dev/null - mkdir -p "$jobs_dir" printf '%s\t%s\n' "$runtime_id" "$runtime_image" >"$image_record" + printf '%s\n' "$code_tar" >"$code_record" effort_args=() if [[ -n "${reasoning_effort:-}" ]]; then @@ -304,7 +467,7 @@ for model_config in "${model_configs[@]}"; do # moves on to the next model as run.sh's --continue-on-error did. Non-zero # means the job itself could not run, e.g. a rejected config or StirrupAgent # refusing its credentials, which aborts the job as the first trial starts. - if ! uv run --env-file "$env_file" harbor run -y \ + if ! AOB_CODE_TAR="$code_tar" uv run --env-file "$env_file" harbor run -y \ -p "$tasks_dir" \ --agent assetops_harbor.stirrup:StirrupAgent \ --model "$model_id" \ @@ -321,6 +484,8 @@ for model_config in "${model_configs[@]}"; do printf 'Harbor could not run %s; see the error above.\n' "$job_path" >&2 status=1 fi + rm -rf "$job_lock" + job_lock="" done exit "$status" diff --git a/src/assetops_harbor/stirrup.py b/src/assetops_harbor/stirrup.py index 8bee60a1b..fcadce9d3 100644 --- a/src/assetops_harbor/stirrup.py +++ b/src/assetops_harbor/stirrup.py @@ -39,6 +39,7 @@ from __future__ import annotations import json +import os import shlex from datetime import UTC, datetime from pathlib import Path @@ -85,6 +86,11 @@ # "explicit value wins" would hold for the CLI and not for Harbor. SETTING_ENV_VARS: tuple[str, ...] = (FMSR_MODEL_ENV,) +# Names the one env file _load_dotenv reads, in place of the nearest .env. +# benchmarks/harbor/run.sh sets it to its ENV_FILE, so a run with another file +# cannot pick up variables from the repo's .env. +ENV_FILE_ENV = "AOB_ENV_FILE" + # Forwarded from the Harbor process into the agent container when present. # Harbor scopes them to the agent phase, so the verifier and build steps never # see them. `--ae KEY=VALUE` still takes precedence over the host environment. @@ -548,8 +554,12 @@ def _load_dotenv() -> None: file itself out of every image. The verifier resolves its ${VAR:-} templates after the agent is constructed, so AOB_JUDGE_MODEL and the judge keys are picked up from .env as well. + + AOB_ENV_FILE, when set, is read instead of the nearest .env, so the file a + caller chose is the only one: its gaps are not filled from the repo's .env. """ - load_dotenv(find_dotenv(usecwd=True), override=False) + path = os.environ.get(ENV_FILE_ENV) or find_dotenv(usecwd=True) + load_dotenv(path, override=False) def _as_bool(value: Any) -> bool: diff --git a/src/assetops_harbor/tests/test_run_sh.py b/src/assetops_harbor/tests/test_run_sh.py new file mode 100644 index 000000000..f0fa14ae6 --- /dev/null +++ b/src/assetops_harbor/tests/test_run_sh.py @@ -0,0 +1,89 @@ +"""Keep benchmarks/harbor/run.sh in step with Harbor and llm.routers. + +run.sh names Harbor exception classes as strings for `harbor jobs resume +--filter-error-type`, which matches the exact class name and ignores a name +that matches nothing. A typo, a rename, or a new subclass would silently stop +those trials from running again, so these tests read the names out of the +script and check them against the installed Harbor. +""" + +from __future__ import annotations + +import asyncio +import inspect +import re +from pathlib import Path + +import pytest + +pytest.importorskip( + "harbor.agents.installed.base", + reason="Harbor is an optional extra: uv sync --dev --extra harbor", +) + +from harbor.agents.installed import base as installed_base +from harbor.environments import base as environments_base +from harbor.trial import errors as trial_errors +from harbor.verifier import verifier + +from assetops_harbor.stirrup import ROUTER_CREDENTIALS + +RUN_SH = Path(__file__).resolve().parents[3] / "benchmarks/harbor/run.sh" + +# Failures that are the model's own work, so a resume keeps them as results. +KEPT = { + "AgentTimeoutError", + "ContextWindowExceededError", + "OutputTokenExceededError", + "AgentSafetyRefusalError", +} + + +def _retry_error_types() -> list[str]: + text = RUN_SH.read_text(encoding="utf-8") + block = re.search(r"^retry_error_types=\((.*?)^\)", text, re.MULTILINE | re.DOTALL) + assert block, "retry_error_types=( ... ) not found in run.sh" + return block.group(1).split() + + +def _exception_classes() -> dict[str, type]: + classes = {"CancelledError": asyncio.CancelledError} + for module in (installed_base, environments_base, trial_errors, verifier): + for name, obj in vars(module).items(): + if inspect.isclass(obj) and issubclass(obj, BaseException): + classes[name] = obj + return classes + + +def test_every_retried_error_type_exists_in_harbor() -> None: + known = _exception_classes() + unknown = [name for name in _retry_error_types() if name not in known] + assert not unknown, f"not Harbor exception classes: {unknown}" + + +def test_every_agent_error_is_retried_or_kept() -> None: + """A new NonZeroAgentExitCodeError subclass must be placed on one side.""" + listed = set(_retry_error_types()) + base = installed_base.NonZeroAgentExitCodeError + subclasses = { + name + for name, obj in vars(installed_base).items() + if inspect.isclass(obj) and issubclass(obj, base) + } + unplaced = subclasses - listed - KEPT + assert not unplaced, f"add to retry_error_types in run.sh or to KEPT: {unplaced}" + assert not listed & KEPT + + +def test_the_router_probe_matches_the_agent() -> None: + """check_model's ROUTERS is a copy; ROUTER_CREDENTIALS is tested against llm.""" + text = RUN_SH.read_text(encoding="utf-8") + block = re.search(r"^ROUTERS = \{(.*?)^\}", text, re.MULTILINE | re.DOTALL) + assert block, "ROUTERS = { ... } not found in run.sh" + routers = { + prefix: (base, key) + for prefix, base, key in re.findall( + r'"([^"]+)": \("([^"]+)", "([^"]+)"\)', block.group(1) + ) + } + assert routers == ROUTER_CREDENTIALS diff --git a/src/assetops_harbor/tests/test_stirrup.py b/src/assetops_harbor/tests/test_stirrup.py index c8d5d0034..83860a2e9 100644 --- a/src/assetops_harbor/tests/test_stirrup.py +++ b/src/assetops_harbor/tests/test_stirrup.py @@ -25,6 +25,7 @@ from harbor.models.trajectories import Trajectory from assetops_harbor.stirrup import ( + ENV_FILE_ENV, FMSR_MODEL_ENV, ROUTER_CREDENTIALS, SETTING_ENV_VARS, @@ -325,6 +326,28 @@ def test_exported_variables_win_over_dotenv( assert forwarded["TOKENROUTER_BASE_URL"] == "https://example.invalid/v1" +def test_aob_env_file_replaces_the_dotenv_search( + tmp_path: Path, no_router_credentials: None, monkeypatch: pytest.MonkeyPatch +) -> None: + """run.sh's ENV_FILE is the only file read; the cwd's .env fills no gaps.""" + (tmp_path / ".env").write_text( + "TOKENROUTER_BASE_URL=https://repo.invalid/v1\nTOKENROUTER_API_KEY=repo\n", + encoding="utf-8", + ) + chosen = tmp_path / "chosen.env" + chosen.write_text( + "LITELLM_BASE_URL=https://chosen.invalid\nLITELLM_API_KEY=chosen\n", + encoding="utf-8", + ) + monkeypatch.setenv(ENV_FILE_ENV, str(chosen)) + agent = StirrupAgent( + logs_dir=tmp_path / "agent", model_name="litellm_proxy/azure/gpt-5.6-sol" + ) + forwarded = agent._credential_env() + assert forwarded["LITELLM_API_KEY"] == "chosen" + assert "TOKENROUTER_API_KEY" not in forwarded + + def test_unprefixed_models_need_no_router_creds(tmp_path: Path) -> None: StirrupAgent(logs_dir=tmp_path / "agent", model_name="watsonx/llama-4") From d9eb32d0357b7a6b8f52d75589551fb7fe67374d Mon Sep 17 00:00:00 2001 From: Shuxin Lin Date: Wed, 30 Sep 2026 19:36:04 -0400 Subject: [PATCH 3/3] fix(harbor): refuse to resume a run.sh job on another suite Found reviewing the per-job task copies. A resume reuses the job's own manifests but mounts shared/ from the current -s, so a resume with another suite would pair one suite's manifests with another suite's data. Before the copies, Harbor's task lock refused that resume; now run.sh records the suite beside the job (.suite) and refuses it, as it does for another runtime image. Co-Authored-By: Claude Opus 5.5 Signed-off-by: Shuxin Lin --- benchmarks/harbor/README.md | 4 +++- benchmarks/harbor/run.sh | 19 +++++++++++++++---- 2 files changed, 18 insertions(+), 5 deletions(-) diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index 9c5fd63f5..a32b0eda4 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -142,7 +142,9 @@ The script: Re-running resumes an existing job and finishes only its incomplete trials, the equivalent of `--skip-existing`. A resume reuses the job's own tasks and settings: later changes to the template, the suite's scenario files or the -generator apply to new jobs only, and so does a changed `-n`. The profile and +generator apply to new jobs only, and so does a changed `-n`. A job started on +another `-s` is not resumed: its tasks hold that suite's manifests, while +`shared/` would come from the new one. The profile and effort are part of the job name so that a second effort, or another profile in the same leaderboard directory, starts its own job instead of resuming the first. Harbor still refuses to resume a job whose overlays have changed since it diff --git a/benchmarks/harbor/run.sh b/benchmarks/harbor/run.sh index d6d6f6d4f..ff10b404b 100755 --- a/benchmarks/harbor/run.sh +++ b/benchmarks/harbor/run.sh @@ -386,12 +386,16 @@ for model_config in "${model_configs[@]}"; do # Keyed by the job's full path, so the same job name under another # LEADERBOARD_DIR gets its own copy. tasks_dir="$tasks_root/$job_name-$(printf '%s' "$job_path" | cksum | cut -d' ' -f1)" - # The runtime image and code tar a job started on, beside the job rather than - # in it. Harbor's resume lock covers the task files but not the base they - # build FROM, so without the first a resume with another -r would mix two - # images in one job; the second makes a resume load the job's own code image. + # What a job started on, beside the job rather than in it. Harbor's resume + # lock covers the task files but not the base they build FROM, so without the + # runtime image record a resume with another -r would mix two images in one + # job. The code tar record makes a resume load the job's own code image. The + # suite record matters because a resume reuses the job's manifests but mounts + # shared/ from the current -s: another suite would pair one suite's manifests + # with another's data. image_record="$jobs_dir/$job_name.runtime-image" code_record="$jobs_dir/$job_name.code-tar" + suite_record="$jobs_dir/$job_name.suite" if ! check_model "$model_id"; then echo "Skipping $model_id" >&2 @@ -419,6 +423,12 @@ for model_config in "${model_configs[@]}"; do printf 'Pass that image as -r to finish it, or move the job aside to rerun %s.\n' \ "$model_id" >&2 status=1 + elif [[ -f "$suite_record" ]] && [[ "$(cat "$suite_record")" != "$scenario_dir" ]]; then + printf '%s started on the suite in %s, not %s.\n' \ + "$job_path" "$(cat "$suite_record")" "$scenario_dir" >&2 + printf 'Pass that directory as -s to finish it, or move the job aside to rerun %s.\n' \ + "$model_id" >&2 + status=1 elif [[ ! -d "$tasks_dir" ]]; then printf 'The tasks %s started with are gone (%s).\n' "$job_path" "$tasks_dir" >&2 printf 'Move the job aside to rerun %s from scratch.\n' "$model_id" >&2 @@ -457,6 +467,7 @@ for model_config in "${model_configs[@]}"; do printf '%s\t%s\n' "$runtime_id" "$runtime_image" >"$image_record" printf '%s\n' "$code_tar" >"$code_record" + printf '%s\n' "$scenario_dir" >"$suite_record" effort_args=() if [[ -n "${reasoning_effort:-}" ]]; then