Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 6 additions & 4 deletions benchmarks/harbor/CODE-SANDBOX.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,8 +70,9 @@ well below what you use for the tools-only arm.
## The easy path: run.sh

`benchmarks/harbor/run.sh` already does all of it. It builds the code image,
saves the tar, exports both variables, generates the dataset and passes the
overlay with the four required `--ak` flags.
saves the tar (one per image id, in `~/.cache/assetopsbench`), points each job
at its own tar, generates the dataset and passes the overlay with the four
required `--ak` flags.

```bash
./benchmarks/harbor/run.sh \
Expand Down Expand Up @@ -197,8 +198,9 @@ one this overlay is known to work on.

**"no AOB_CODE_TAR; the daemon will pull ..."** in the loader log. Expected when
you went the registry route. If you meant to use a tar, the path in
`AOB_CODE_TAR` is wrong or the file is empty. It defaults to
`$HOME/assetops-code.tar` under `run.sh`.
`AOB_CODE_TAR` is wrong or the file is empty. Under `run.sh` it is the job's
own tar in `~/.cache/assetopsbench`, whose path is in `<job>.code-tar` beside
the job.

**Runs are much slower than the tools-only arm.** Each trial pays for a
container and an image load. Lower `--n-concurrent`.
Expand Down
73 changes: 58 additions & 15 deletions benchmarks/harbor/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@ benchmarks/harbor/
overlays/private-data.yaml Opt-in read-only mount of a private suite's shared/ (AOB_PRIVATE_DIR)
metric.py Category rollups
datasets/assetopsbench-<profile>/ Generated by generate_tasks.py (gitignored), one per profile
datasets/jobs/<job>-<key>/ run.sh's per-job copy of the tasks (gitignored), same layout
dataset.toml Harbor dataset manifest
metric.py Copied from benchmarks/harbor/metric.py
wosr-1/ One task = one scenario
Expand Down Expand Up @@ -93,7 +94,19 @@ bash benchmarks/harbor/run.sh \
`-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), `-n`
the number of concurrent trials, and `-r` the runtime image (otherwise
`AOB_RUNTIME_IMAGE` from the shell, then from `.env`, then
`assetopsbench/runtime:dev`). The script:
`assetopsbench/runtime:dev`). Credentials come from `ENV_FILE`, the repo's
`.env` by default, and only from that file: with another file, the repo's
`.env` fills none of its gaps. Relative paths in `-s`, `-l`, `-p`, `ENV_FILE`
and `AOB_CODE_TAR_DIR` are relative to the directory you run the script from.

Before each model, the script checks that it can be served, and skips it
otherwise. The model's router and `FMSR_MODEL_ID`'s must answer and must not
reject their key (a 401 or 403 from `/models`). A model with no `litellm_proxy/`
or `tokenrouter/` prefix also needs `FMSR_MODEL_ID` set to one, because the FMSR
server would otherwise reject it and every fmsr scenario would run without
`generate_failure_modes`.

The script:

1. resolves the runtime image, which every task image builds FROM. A published
image, one whose local copy came from that registry or that is not local at
Expand All @@ -108,21 +121,50 @@ the number of concurrent trials, and `-r` the runtime image (otherwise
`overlays/private-data.yaml` mounts its `shared/` at
`/opt/suite/scenarios_data/shared`. Only `init_data.py` reads it, as in
`scenario_suite_runner`;
3. builds the code sandbox image and saves it to `~/assetops-code.tar` for the
per-trial Docker-in-Docker daemon (`overlays/code-sandbox.yaml`);
4. generates one task per scenario with `--scenario-root` and
`--skip-missing`, skipping profile entries the suite lacks;
5. runs one Harbor job per model at
`<leaderboard>/harbor-jobs/stirrup_agent__<model>`, with both overlays and
with credentials loaded from `.env` by `uv run --env-file` into the Harbor
process only.
3. builds the code sandbox image for the per-trial Docker-in-Docker daemon
(`overlays/code-sandbox.yaml`) on every run, which the build cache makes
cheap, so a change to `Dockerfile.code` is picked up. It saves the image as
a tar named after its id, in `AOB_CODE_TAR_DIR` (default
`~/.cache/assetopsbench`), once per image, and never rewrites it. Each job
records its tar, and a resume loads that one, so a job keeps one code image
from start to finish. `AOB_CODE_TAR` from your shell is not used. Old tars
stay until you delete them, and `~/assetops-code.tar` is no longer used;
4. runs one Harbor job per profile, model and reasoning effort at
`<leaderboard>/harbor-jobs/stirrup_agent__<profile>__<model>[__<effort>]`,
with both overlays and with credentials loaded from `.env` by
`uv run --env-file` into the Harbor process only;
5. gives each new job its own copy of the tasks, generated with
`--scenario-root` and `--skip-missing` (skipping profile entries the suite
lacks) into `benchmarks/harbor/datasets/jobs/`. The copies hold the
answers, so they stay in the repo's gitignored `datasets/`, not in the
leaderboard directory.

Re-running resumes an existing job and finishes only its incomplete trials,
the equivalent of `--skip-existing`. Harbor refuses to resume a job whose tasks
or overlays have changed since it started, such as one from before `run.sh`
mounted only `shared/`; the script says so, and moving that job aside reruns
the model from scratch. The script exits non-zero when any model's job could
not start or resume, so a wrapper sees it. Each trial with the code sandbox runs a
the equivalent of `--skip-existing`. A resume reuses the job's own tasks and
settings: later changes to the template, the suite's scenario files or the
generator apply to new jobs only, and so does a changed `-n`. A job started on
another `-s` is not resumed: its tasks hold that suite's manifests, while
`shared/` would come from the new one. The profile and
effort are part of the job name so that a second effort, or another profile in
the same leaderboard directory, starts its own job instead of resuming the
first. Harbor still refuses to resume a job whose overlays have changed since it
started; the script says so, and moving that job aside reruns the model from
scratch.

A resume runs again every trial that failed for a reason other than the model's
own work: its API or the network, the environment, the verifier, or Ctrl-C.
Harbor matches exact class names, so `run.sh` lists each one
(`retry_error_types`), and `src/assetops_harbor/tests/test_run_sh.py` checks the
list against the installed Harbor. A timeout, an exceeded context window or
output limit, and a safety refusal are kept as results.

Several `run.sh` processes can share a leaderboard directory. Only one works on
a given job at a time: the other skips it and names the lock
(`<job>.lock`). A lock whose process has died is taken over.

The script exits non-zero when any model's job could not start or resume,
including a model skipped by the checks above, so a wrapper sees it. A trial
that fails inside a job does not count. Each trial with the code sandbox runs a
privileged `dind` sidecar, so keep `-n` around 4 on a laptop-sized Docker VM.

## Generating other profiles by hand
Expand Down Expand Up @@ -307,7 +349,8 @@ the runner's behaviour and the SDK default it compensates for.
The agent forwards credentials from the Harbor process into the container for
the agent phase only. Putting them in the repo's `.env` is enough when you run
from the repo root: `StirrupAgent` loads the nearest `.env` above the working
directory before it checks for credentials.
directory before it checks for credentials. `AOB_ENV_FILE`, when set, names the
one file it loads instead; `run.sh` sets it to its `ENV_FILE`.

```bash
uv run harbor run -p benchmarks/harbor/datasets/assetopsbench-open \
Expand Down
Loading
Loading