From e91aa075278ab9288e3fda505957e6df3fc9305f Mon Sep 17 00:00:00 2001 From: Shuxin Lin Date: Wed, 30 Sep 2026 15:22:30 -0400 Subject: [PATCH 1/3] feat(harbor): pick the runtime image at run time with AOB_RUNTIME_IMAGE Every task image builds FROM the runtime image, and the only way to choose it was the Dockerfile's ARG default, so using a published image meant retagging it as the local assetopsbench/runtime:dev. The template's docker-compose.yaml now passes AOB_RUNTIME_IMAGE to that ARG as a build arg. Harbor runs `docker compose build` with the shell's environment and every compose file, so `AOB_RUNTIME_IMAGE= harbor run ...` builds FROM that image, with no task regeneration; unset, it keeps the local tag. run.sh takes it as -r (or AOB_RUNTIME_IMAGE). A registry reference is pulled first, because a local copy satisfies FROM and the build would otherwise run on a stale one; a bare name must already exist locally, as before. It exports the variable so `harbor jobs resume` sees it too, and prints the image id it uses. The image changes only the base. Each task still adds its own manifest.json, and shared/ still comes from overlays/private-data.yaml. The template change alters every generated task, so Harbor will not resume a job started before it; run.sh already says so and suggests moving it aside. Verified with Harbor's oracle agent, AOB_RUNTIME_IMAGE set to a labelled copy of runtime:dev: the 3 open tasks, and tsfm-1019 and fmsr-913 from the private suite with the private-data overlay. All 5 trials scored 1.000 with no errors. A nonexistent AOB_RUNTIME_IMAGE failed the build pulling exactly that name. The tsfm-1019 task image built on the labelled base carries the label. With the overlay's mount it sees scenario_1019/manifest.json, and a read-only shared/ where all five manifest paths resolve. run.sh's image step pulls a quay.io reference and rejects a missing local tag. 22 assetops_harbor tests pass. Co-Authored-By: Claude Opus 5.5 Signed-off-by: Shuxin Lin --- benchmarks/harbor/QUICKSTART.md | 16 ++++---- benchmarks/harbor/README.md | 28 +++++++++---- benchmarks/harbor/run.sh | 41 ++++++++++++++++--- benchmarks/harbor/scripts/publish-images.sh | 10 ++--- .../template/environment/docker-compose.yaml | 14 ++++++- 5 files changed, 80 insertions(+), 29 deletions(-) diff --git a/benchmarks/harbor/QUICKSTART.md b/benchmarks/harbor/QUICKSTART.md index f6780d5ca..d96e79668 100644 --- a/benchmarks/harbor/QUICKSTART.md +++ b/benchmarks/harbor/QUICKSTART.md @@ -23,11 +23,12 @@ uv sync --dev --extra harbor ## 2. Get the runtime image -Every task image layers its scenario onto one shared runtime image. Pull it: +Every task image layers its scenario onto one shared runtime image. Pull the +published one (currently linux/arm64 only) and point the tasks at it: ```bash -docker pull assetopsbench/runtime:latest -docker tag assetopsbench/runtime:latest assetopsbench/runtime:dev +docker pull quay.io/assetopsbench/runtime:dev +export AOB_RUNTIME_IMAGE=quay.io/assetopsbench/runtime:dev ``` Or build it yourself, which takes a few minutes and needs no registry: @@ -40,10 +41,11 @@ The script builds `assetopsbench/runtime:dev` from `git archive HEAD`, not from your working tree, so untracked files such as local results never reach the container the agent runs in. Commit a change first to include it. -Either way the local tag `assetopsbench/runtime:dev` is what the task -Dockerfiles reference, through the `AOB_RUNTIME_IMAGE` build arg in -`benchmarks/harbor/template/environment/Dockerfile`. Change that default if you -publish under a different namespace. +Each task's `environment/docker-compose.yaml` passes `AOB_RUNTIME_IMAGE` to +its Dockerfile as a build arg, so `harbor run` builds FROM whatever that +variable names when it runs, and from the local `assetopsbench/runtime:dev` +when it is unset. Changing the image needs no task regeneration. `run.sh` takes +it as `-r`. ## 3. Generate the tasks diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index 20e266eac..244e3059b 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -41,8 +41,10 @@ command from the repo root. ```bash # 1. Build the runtime base (once per AssetOpsBench commit). The script builds # from `git archive HEAD`, so untracked files and uncommitted changes stay -# out of the image. The tag is local and is the default AOB_RUNTIME_IMAGE in -# template/environment/Dockerfile; nothing is pulled or pushed. +# out of the image. The tag is local and is the default AOB_RUNTIME_IMAGE +# that each task's docker-compose.yaml passes to its Dockerfile; nothing is +# pulled or pushed. To use a published image instead, pull it and +# `export AOB_RUNTIME_IMAGE=` before `harbor run`. bash benchmarks/harbor/scripts/build-runtime-image.sh # 2. Generate one task per scenario in the open profile. The defaults point at @@ -83,18 +85,26 @@ bash benchmarks/harbor/run.sh \ ``` `-m` may repeat; without it the script runs `benchmarks/run.sh`'s model list. -`-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), and `-n` -the number of concurrent trials. The script: - -1. exports `AOB_PRIVATE_DIR` as the `-s` directory, so +`-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), `-n` +the number of concurrent trials, and `-r` the runtime image (default +`assetopsbench/runtime:dev`, or `AOB_RUNTIME_IMAGE`). The script: + +1. exports `AOB_RUNTIME_IMAGE` as the `-r` image, which every task image builds + FROM. A registry reference such as `quay.io/assetopsbench/runtime:dev` is + pulled first, so the run uses the published image rather than a stale local + copy; a bare name like the default must already exist locally. The image + changes only the base: each task still adds its own `manifest.json`, and + `shared/` still comes from the mount below. A resumed job finishes on the + image given now; +2. exports `AOB_PRIVATE_DIR` as the `-s` directory, so `overlays/private-data.yaml` mounts its `shared/` at `/opt/suite/scenarios_data/shared`. Only `init_data.py` reads it, as in `scenario_suite_runner`; -2. builds the code sandbox image and saves it to `~/assetops-code.tar` for the +3. builds the code sandbox image and saves it to `~/assetops-code.tar` for the per-trial Docker-in-Docker daemon (`overlays/code-sandbox.yaml`); -3. generates one task per scenario with `--scenario-root` and +4. generates one task per scenario with `--scenario-root` and `--skip-missing`, skipping profile entries the suite lacks; -4. runs one Harbor job per model at +5. runs one Harbor job per model at `/harbor-jobs/stirrup_agent__`, with both overlays and with credentials loaded from `.env` by `uv run --env-file` into the Harbor process only. diff --git a/benchmarks/harbor/run.sh b/benchmarks/harbor/run.sh index 16d5dbf47..2cc4fe3a9 100755 --- a/benchmarks/harbor/run.sh +++ b/benchmarks/harbor/run.sh @@ -7,12 +7,17 @@ # shared database. # # bash benchmarks/harbor/run.sh -s SCENARIO_DIR -l LEADERBOARD_DIR \ -# [-n N_CONCURRENT] [-p PROFILE] [-m "MODEL_ID REASONING_EFFORT"]... +# [-n N_CONCURRENT] [-p PROFILE] [-r RUNTIME_IMAGE] \ +# [-m "MODEL_ID REASONING_EFFORT"]... # -# Prerequisites: Docker running, `uv sync --extra harbor`, and the runtime image +# Prerequisites: Docker running, `uv sync --extra harbor`, and the runtime image, +# either built locally # # bash benchmarks/harbor/scripts/build-runtime-image.sh # +# or published, passed as -r (or AOB_RUNTIME_IMAGE), e.g. +# -r quay.io/assetopsbench/runtime:dev. +# # Credentials are read from ENV_FILE (default .env) by `uv run --env-file`, into # the Harbor process only; StirrupAgent forwards them to the agent phase. They # never enter an image. @@ -24,7 +29,7 @@ set -euo pipefail usage() { - printf 'Usage: %s -s SCENARIO_DIR -l LEADERBOARD_DIR [-n N_CONCURRENT] [-p PROFILE] [-m "MODEL_ID EFFORT"]...\n' "$0" >&2 + printf 'Usage: %s -s SCENARIO_DIR -l LEADERBOARD_DIR [-n N_CONCURRENT] [-p PROFILE] [-r RUNTIME_IMAGE] [-m "MODEL_ID EFFORT"]...\n' "$0" >&2 } scenario_dir="${SCENARIO_DIR:-}" @@ -32,14 +37,16 @@ leaderboard_dir="${LEADERBOARD_DIR:-}" n_concurrent="${N_CONCURRENT:-4}" profile="${PROFILE:-benchmarks/scenario_suite/all.yaml}" env_file="${ENV_FILE:-.env}" +runtime_image="${AOB_RUNTIME_IMAGE:-assetopsbench/runtime:dev}" model_configs=() -while getopts ':s:l:n:p:m:' option; do +while getopts ':s:l:n:p:r:m:' option; do case "$option" in s) scenario_dir="$OPTARG" ;; l) leaderboard_dir="$OPTARG" ;; n) n_concurrent="$OPTARG" ;; p) profile="$OPTARG" ;; + r) runtime_image="$OPTARG" ;; m) model_configs+=("$OPTARG") ;; :) printf 'Option -%s requires an argument.\n' "$OPTARG" >&2; usage; exit 2 ;; \?) printf 'Unknown option: -%s\n' "$OPTARG" >&2; usage; exit 2 ;; @@ -77,16 +84,38 @@ if [[ ! -f "$env_file" ]]; then exit 2 fi -runtime_image=assetopsbench/runtime:dev code_image=assetops-code:dev code_tar="${AOB_CODE_TAR:-$HOME/assetops-code.tar}" dataset_dir=benchmarks/harbor/datasets/assetopsbench-suite -if ! docker image inspect "$runtime_image" >/dev/null 2>&1; then +# Every task image builds FROM the runtime image, through the AOB_RUNTIME_IMAGE +# build arg in the task's docker-compose.yaml. A reference whose first component +# names a registry host (quay.io/..., localhost:5000/...) is pulled, so the run +# gets the published image rather than a stale local copy; the build would not +# pull it, because a local copy satisfies FROM. A bare name such as the default +# is a local build from build-runtime-image.sh, published nowhere under that +# name, so it has to exist already. +registry="${runtime_image%%/*}" +if [[ "$runtime_image" == */* && ( "$registry" == *.* || "$registry" == *:* || "$registry" == localhost ) ]]; then + if ! docker pull "$runtime_image"; then + if ! docker image inspect "$runtime_image" >/dev/null 2>&1; then + printf 'Could not pull runtime image %s\n' "$runtime_image" >&2 + exit 1 + fi + printf 'warning: could not pull %s; using the local copy, which may be stale\n' \ + "$runtime_image" >&2 + fi +elif ! docker image inspect "$runtime_image" >/dev/null 2>&1; then printf 'Runtime image %s not found. Build it first:\n' "$runtime_image" >&2 printf ' bash benchmarks/harbor/scripts/build-runtime-image.sh\n' >&2 exit 1 fi +# Compose reads this on `harbor run` and again on `harbor jobs resume`, so a +# resumed job finishes its trials on the image given now, not the one it +# started with. +export AOB_RUNTIME_IMAGE="$runtime_image" +printf 'Runtime image: %s (%s)\n' "$runtime_image" \ + "$(docker image inspect --format '{{.Id}}' "$runtime_image" | cut -c8-19)" # The suite's shared/ data reaches each trial through # overlays/private-data.yaml, a read-only bind mount of this directory's shared/. diff --git a/benchmarks/harbor/scripts/publish-images.sh b/benchmarks/harbor/scripts/publish-images.sh index 25a793189..1056cf37b 100755 --- a/benchmarks/harbor/scripts/publish-images.sh +++ b/benchmarks/harbor/scripts/publish-images.sh @@ -62,13 +62,11 @@ cat < Date: Wed, 30 Sep 2026 15:29:52 -0400 Subject: [PATCH 2/3] feat(harbor): tag each runtime build with its commit build-runtime-image.sh tagged only assetopsbench/runtime:dev, which every build overwrites, so a run could not name the build it used after a rebuild. With no -t given it now also tags assetopsbench/runtime:. That tag does not move, so `run.sh -r assetopsbench/runtime:` (or AOB_RUNTIME_IMAGE) pins a run to one build while :dev stays the default the tasks build FROM. publish-images.sh adds /runtime: beside :TAG and :latest. The commit is HEAD's, which is exactly what the image holds, because the build comes from `git archive HEAD` whatever the working tree contains. Co-Authored-By: Claude Opus 5.5 Signed-off-by: Shuxin Lin --- benchmarks/harbor/QUICKSTART.md | 5 +++-- benchmarks/harbor/README.md | 5 +++-- benchmarks/harbor/scripts/build-runtime-image.sh | 11 +++++++---- benchmarks/harbor/scripts/publish-images.sh | 8 +++++++- 4 files changed, 20 insertions(+), 9 deletions(-) diff --git a/benchmarks/harbor/QUICKSTART.md b/benchmarks/harbor/QUICKSTART.md index d96e79668..5c34b1a85 100644 --- a/benchmarks/harbor/QUICKSTART.md +++ b/benchmarks/harbor/QUICKSTART.md @@ -37,8 +37,9 @@ Or build it yourself, which takes a few minutes and needs no registry: bash benchmarks/harbor/scripts/build-runtime-image.sh ``` -The script builds `assetopsbench/runtime:dev` from `git archive HEAD`, not -from your working tree, so untracked files such as local results never reach +The script builds `assetopsbench/runtime:dev`, also tagged +`assetopsbench/runtime:`, from `git archive HEAD`, not from your +working tree, so untracked files such as local results never reach the container the agent runs in. Commit a change first to include it. Each task's `environment/docker-compose.yaml` passes `AOB_RUNTIME_IMAGE` to diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index 244e3059b..21bfc6e03 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -43,8 +43,9 @@ command from the repo root. # from `git archive HEAD`, so untracked files and uncommitted changes stay # out of the image. The tag is local and is the default AOB_RUNTIME_IMAGE # that each task's docker-compose.yaml passes to its Dockerfile; nothing is -# pulled or pushed. To use a published image instead, pull it and -# `export AOB_RUNTIME_IMAGE=` before `harbor run`. +# pulled or pushed. It also tags assetopsbench/runtime:, which +# later builds do not move, to pin a run to this build. To use a published +# or pinned image, `export AOB_RUNTIME_IMAGE=` before `harbor run`. bash benchmarks/harbor/scripts/build-runtime-image.sh # 2. Generate one task per scenario in the open profile. The defaults point at diff --git a/benchmarks/harbor/scripts/build-runtime-image.sh b/benchmarks/harbor/scripts/build-runtime-image.sh index 88d7b6576..e92a12af1 100755 --- a/benchmarks/harbor/scripts/build-runtime-image.sh +++ b/benchmarks/harbor/scripts/build-runtime-image.sh @@ -3,11 +3,14 @@ # # bash benchmarks/harbor/scripts/build-runtime-image.sh # bash benchmarks/harbor/scripts/build-runtime-image.sh --no-cache -# bash benchmarks/harbor/scripts/build-runtime-image.sh -t assetopsbench/runtime:abc1234 +# bash benchmarks/harbor/scripts/build-runtime-image.sh -t myorg/runtime:test # # Arguments go to `docker buildx build` unchanged. Unless they name a tag (-t) -# the image is tagged assetopsbench/runtime:dev, the tag the task template -# builds FROM, and unless they name an output (--push, --output, --load) it is +# the image is tagged twice: assetopsbench/runtime:dev, the tag the task +# template builds FROM by default, and assetopsbench/runtime:, the +# short hash of the HEAD it was built from. :dev moves with every build; the +# commit tag does not, so `run.sh -r assetopsbench/runtime:` pins a run +# to one build. Unless they name an output (--push, --output, --load) it is # loaded into the local image store. publish-images.sh passes --platform, its # own tags and --push. # @@ -54,7 +57,7 @@ for arg in "$@"; do esac done defaults=() -$has_tag || defaults+=(-t assetopsbench/runtime:dev) +$has_tag || defaults+=(-t assetopsbench/runtime:dev -t "assetopsbench/runtime:${commit:0:7}") $has_output || defaults+=(--load) printf 'Building the runtime image from %s (git archive)\n' "${commit:0:7}" >&2 diff --git a/benchmarks/harbor/scripts/publish-images.sh b/benchmarks/harbor/scripts/publish-images.sh index 1056cf37b..6ede0e0ce 100755 --- a/benchmarks/harbor/scripts/publish-images.sh +++ b/benchmarks/harbor/scripts/publish-images.sh @@ -42,10 +42,16 @@ if ! docker buildx inspect >/dev/null 2>&1; then exit 1 fi -echo "==> ${NAMESPACE}/runtime:${TAG} for ${PLATFORMS}" +# The commit tag never moves, unlike TAG and latest, so a published run can +# name the exact build it used. build-runtime-image.sh builds HEAD, so this is +# the commit in the image whatever the working tree holds. +COMMIT="$(git rev-parse HEAD | cut -c1-7)" + +echo "==> ${NAMESPACE}/runtime:${TAG} (${COMMIT}) for ${PLATFORMS}" bash benchmarks/harbor/scripts/build-runtime-image.sh \ --platform "${PLATFORMS}" \ -t "${NAMESPACE}/runtime:${TAG}" \ + -t "${NAMESPACE}/runtime:${COMMIT}" \ -t "${NAMESPACE}/runtime:latest" \ --push From 1ad4eea97d56103865c8d75c0928f5e7c90589b6 Mon Sep 17 00:00:00 2001 From: Shuxin Lin Date: Wed, 30 Sep 2026 15:51:16 -0400 Subject: [PATCH 3/3] fix(harbor): pin the runtime image for a run and guard resumes Review fixes for the runtime-image selection. run.sh: - Resolves the image as -r, then the shell's AOB_RUNTIME_IMAGE, then ENV_FILE's, then the local default. It used to take only the shell's value and export a default over it, so an AOB_RUNTIME_IMAGE in .env was ignored; `uv run --env-file` never overrides a variable the shell already has. - Pulls whenever the image is not local, or its local copy came from that same repository. The old rule pulled only references naming a registry host, so a Docker Hub image such as publish-images.sh's own /runtime:latest was never refreshed. A local build (no RepoDigest for its name) is never pulled over. - Pins the resolved image for the whole run under a tag private to the process (aob-runtime-pin:-, removed on exit) and exports that. Every trial resolves FROM when it builds, so exporting the movable :dev let a rebuild or another run's pull switch later trials' base mid-run. FROM cannot name a bare image id, which is why this uses a tag. - Records the image beside each job (.runtime-image) and refuses to resume a job on a different one. Harbor's resume lock hashes the task files, not the base they build FROM, so a resume with another -r mixed two images in one job. - Exits non-zero when a model's job could not start or resume. - Runs the image step after the cheap input checks, and strips a sha256: prefix only if present when printing the id. build-runtime-image.sh: treats -tNAME and -t=NAME as naming a tag, so they no longer also get the default :dev and : tags. Its comment no longer claims the commit tag never moves (rebuilding the same commit moves it), and it says how to remove old commit tags, which now keep their images alive. publish-images.sh: the NOTE points users at the commit tag it just pushed, with :latest as the moving option. Docs: QUICKSTART says tasks generated before the build arg existed ignore the variable and need regenerating once, that the variable must be set per shell or in .env, and its pull-error entry now covers an unset AOB_RUNTIME_IMAGE. README describes the resolution order, pin, per-job record and exit status. Verified end to end with run.sh, a one-scenario profile and a fake model id, so trials fail fast without calling a provider: - The first run printed the resolved id. The pin tag existed during the run and was gone after it, the job's .runtime-image record was written, and the trial built, ran and was scored (reward 0). - A second run on the same image resumed the job, dropped the crashed trial and reran it, and exited 0. - A third run with -r set to a different image refused to resume, named the job's original image, and exited 1. - Helper checks: image_repo strips tags and digests; runtime:dev and :3bb4f12 count as local builds; quay.io/... and couchdb:3.5 count as registry images. -r and the shell beat .env, which beats the default. - `uv run pytest src/ -k "not integration"`: 714 passed, 19 failed. The base branch has the same 19 failures, in the scorer and trace-exporter tests and in the iot invalid-site tests, which need a CouchDB or no .env. Co-Authored-By: Claude Opus 5.5 Signed-off-by: Shuxin Lin --- benchmarks/harbor/QUICKSTART.md | 15 ++- benchmarks/harbor/README.md | 31 +++-- benchmarks/harbor/run.sh | 115 +++++++++++++----- .../harbor/scripts/build-runtime-image.sh | 14 ++- benchmarks/harbor/scripts/publish-images.sh | 14 ++- 5 files changed, 128 insertions(+), 61 deletions(-) diff --git a/benchmarks/harbor/QUICKSTART.md b/benchmarks/harbor/QUICKSTART.md index 5c34b1a85..27db5f5f5 100644 --- a/benchmarks/harbor/QUICKSTART.md +++ b/benchmarks/harbor/QUICKSTART.md @@ -45,8 +45,11 @@ the container the agent runs in. Commit a change first to include it. Each task's `environment/docker-compose.yaml` passes `AOB_RUNTIME_IMAGE` to its Dockerfile as a build arg, so `harbor run` builds FROM whatever that variable names when it runs, and from the local `assetopsbench/runtime:dev` -when it is unset. Changing the image needs no task regeneration. `run.sh` takes -it as `-r`. +when it is unset. The variable lives in the shell, so set it again in a new +one, or put it in `.env` and run Harbor as `uv run --env-file .env harbor run`. +Changing the image needs no task regeneration, but tasks generated before the +build arg existed ignore the variable; regenerate those once (step 3). +`run.sh` takes the image as `-r`. ## 3. Generate the tasks @@ -170,9 +173,11 @@ reach CouchDB. Confirm your checkout includes the `env` entry in `StirrupAgentRunner._build_mcp_config`, without which the MCP SDK hands each server a six-variable environment that omits `COUCHDB_URL`. -**Trials fail instantly with a pull error** — the runtime image is missing, or -it is tagged under a name the task Dockerfile does not reference. Redo step 2 -and check `docker images assetopsbench/runtime`. +**Trials fail instantly with a pull error** — the build could not find its +base. Either `AOB_RUNTIME_IMAGE` is unset in this shell, so the build fell back +to the local `assetopsbench/runtime:dev`, which does not exist; or it names an +image that is not local (check `echo $AOB_RUNTIME_IMAGE` and +`docker images`). Redo step 2 in this shell. **`No module named 'google.protobuf'` during a run** — the image was built without the `otel` dependency group, which the file trace exporter needs. diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index 21bfc6e03..a46a7e1a4 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -44,8 +44,9 @@ command from the repo root. # out of the image. The tag is local and is the default AOB_RUNTIME_IMAGE # that each task's docker-compose.yaml passes to its Dockerfile; nothing is # pulled or pushed. It also tags assetopsbench/runtime:, which -# later builds do not move, to pin a run to this build. To use a published -# or pinned image, `export AOB_RUNTIME_IMAGE=` before `harbor run`. +# later commits' builds do not move. To use a published image or a commit +# tag, `export AOB_RUNTIME_IMAGE=` before `harbor run`; tasks +# generated before that build arg existed ignore it, so regenerate them. bash benchmarks/harbor/scripts/build-runtime-image.sh # 2. Generate one task per scenario in the open profile. The defaults point at @@ -87,16 +88,19 @@ bash benchmarks/harbor/run.sh \ `-m` may repeat; without it the script runs `benchmarks/run.sh`'s model list. `-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), `-n` -the number of concurrent trials, and `-r` the runtime image (default -`assetopsbench/runtime:dev`, or `AOB_RUNTIME_IMAGE`). The script: - -1. exports `AOB_RUNTIME_IMAGE` as the `-r` image, which every task image builds - FROM. A registry reference such as `quay.io/assetopsbench/runtime:dev` is - pulled first, so the run uses the published image rather than a stale local - copy; a bare name like the default must already exist locally. The image - changes only the base: each task still adds its own `manifest.json`, and - `shared/` still comes from the mount below. A resumed job finishes on the - image given now; +the number of concurrent trials, and `-r` the runtime image (otherwise +`AOB_RUNTIME_IMAGE` from the shell, then from `.env`, then +`assetopsbench/runtime:dev`). The script: + +1. resolves the runtime image, which every task image builds FROM. A published + image, one whose local copy came from that registry or that is not local at + all, is pulled first, so the run uses the current publish rather than a + stale copy; a local build such as the default is used as is. It then pins + that exact image for the whole run, so a rebuild or pull of the same tag + during the run does not switch later trials' base, and records it beside + each job: a job started on another image is not resumed. The image changes + only the base: each task still adds its own `manifest.json`, and `shared/` + still comes from the mount below; 2. exports `AOB_PRIVATE_DIR` as the `-s` directory, so `overlays/private-data.yaml` mounts its `shared/` at `/opt/suite/scenarios_data/shared`. Only `init_data.py` reads it, as in @@ -114,7 +118,8 @@ Re-running resumes an existing job and finishes only its incomplete trials, the equivalent of `--skip-existing`. Harbor refuses to resume a job whose tasks or overlays have changed since it started, such as one from before `run.sh` mounted only `shared/`; the script says so, and moving that job aside reruns -the model from scratch. Each trial with the code sandbox runs a +the model from scratch. The script exits non-zero when any model's job could +not start or resume, so a wrapper sees it. Each trial with the code sandbox runs a privileged `dind` sidecar, so keep `-n` around 4 on a laptop-sized Docker VM. ## Generating other profiles by hand diff --git a/benchmarks/harbor/run.sh b/benchmarks/harbor/run.sh index 2cc4fe3a9..17f2c63a0 100755 --- a/benchmarks/harbor/run.sh +++ b/benchmarks/harbor/run.sh @@ -15,8 +15,8 @@ # # bash benchmarks/harbor/scripts/build-runtime-image.sh # -# or published, passed as -r (or AOB_RUNTIME_IMAGE), e.g. -# -r quay.io/assetopsbench/runtime:dev. +# or published, passed as -r (or AOB_RUNTIME_IMAGE in the shell or ENV_FILE), +# e.g. -r quay.io/assetopsbench/runtime:dev. # # Credentials are read from ENV_FILE (default .env) by `uv run --env-file`, into # the Harbor process only; StirrupAgent forwards them to the agent phase. They @@ -37,7 +37,7 @@ leaderboard_dir="${LEADERBOARD_DIR:-}" n_concurrent="${N_CONCURRENT:-4}" profile="${PROFILE:-benchmarks/scenario_suite/all.yaml}" env_file="${ENV_FILE:-.env}" -runtime_image="${AOB_RUNTIME_IMAGE:-assetopsbench/runtime:dev}" +runtime_image="${AOB_RUNTIME_IMAGE:-}" model_configs=() while getopts ':s:l:n:p:r:m:' option; do @@ -88,35 +88,6 @@ code_image=assetops-code:dev code_tar="${AOB_CODE_TAR:-$HOME/assetops-code.tar}" dataset_dir=benchmarks/harbor/datasets/assetopsbench-suite -# Every task image builds FROM the runtime image, through the AOB_RUNTIME_IMAGE -# build arg in the task's docker-compose.yaml. A reference whose first component -# names a registry host (quay.io/..., localhost:5000/...) is pulled, so the run -# gets the published image rather than a stale local copy; the build would not -# pull it, because a local copy satisfies FROM. A bare name such as the default -# is a local build from build-runtime-image.sh, published nowhere under that -# name, so it has to exist already. -registry="${runtime_image%%/*}" -if [[ "$runtime_image" == */* && ( "$registry" == *.* || "$registry" == *:* || "$registry" == localhost ) ]]; then - if ! docker pull "$runtime_image"; then - if ! docker image inspect "$runtime_image" >/dev/null 2>&1; then - printf 'Could not pull runtime image %s\n' "$runtime_image" >&2 - exit 1 - fi - printf 'warning: could not pull %s; using the local copy, which may be stale\n' \ - "$runtime_image" >&2 - fi -elif ! docker image inspect "$runtime_image" >/dev/null 2>&1; then - printf 'Runtime image %s not found. Build it first:\n' "$runtime_image" >&2 - printf ' bash benchmarks/harbor/scripts/build-runtime-image.sh\n' >&2 - exit 1 -fi -# Compose reads this on `harbor run` and again on `harbor jobs resume`, so a -# resumed job finishes its trials on the image given now, not the one it -# started with. -export AOB_RUNTIME_IMAGE="$runtime_image" -printf 'Runtime image: %s (%s)\n' "$runtime_image" \ - "$(docker image inspect --format '{{.Id}}' "$runtime_image" | cut -c8-19)" - # The suite's shared/ data reaches each trial through # overlays/private-data.yaml, a read-only bind mount of this directory's shared/. # Compose reads the variable on `harbor run` and again on `harbor jobs resume`. @@ -128,6 +99,63 @@ if [[ ! -d "$scenario_dir/shared" ]]; then fi export AOB_PRIVATE_DIR="$scenario_dir" +# Every task image builds FROM the runtime image, through the AOB_RUNTIME_IMAGE +# build arg in the task's docker-compose.yaml. -r wins, then the shell's +# AOB_RUNTIME_IMAGE, then ENV_FILE's, then the local default: the order +# `uv run --env-file` gives Harbor, where the environment beats the file. +if [[ -z "$runtime_image" ]]; then + runtime_image="$(uv run --env-file "$env_file" python -c \ + 'import os; print(os.environ.get("AOB_RUNTIME_IMAGE", ""))')" +fi +runtime_image="${runtime_image:-assetopsbench/runtime:dev}" + +# The reference without its tag or digest, spelled as .RepoDigests spells it. +image_repo() { + local ref="${1%@*}" + if [[ "${ref##*/}" == *:* ]]; then ref="${ref%:*}"; fi + ref="${ref#docker.io/}" + printf '%s' "${ref#library/}" +} + +# True when the local copy of $1 came from (or went to) that same repository, +# i.e. it is a published image rather than a local build. +from_registry() { + local repo digest + repo="$(image_repo "$1")" + while read -r digest; do + [[ "${digest%@*}" == "$repo" ]] && return 0 + done < <(docker image inspect --format '{{range .RepoDigests}}{{println .}}{{end}}' "$1") + return 1 +} + +# A local copy satisfies FROM, so the build never refreshes a published image; +# pull it here instead, on Docker Hub or any other registry. A local build that +# was never pushed under this name (the default assetopsbench/runtime:dev) is +# used as is, and a pull never replaces it. +if ! docker image inspect "$runtime_image" >/dev/null 2>&1; then + if ! docker pull "$runtime_image"; then + printf 'Runtime image %s is not local and could not be pulled. Build it with\n' "$runtime_image" >&2 + printf ' bash benchmarks/harbor/scripts/build-runtime-image.sh\n' >&2 + exit 1 + fi +elif from_registry "$runtime_image" && ! docker pull "$runtime_image"; then + printf 'warning: could not pull %s; using the local copy, which may be stale\n' \ + "$runtime_image" >&2 +fi + +# Pin the base for the whole run. A tag such as :dev can move while the run is +# going (a rebuild, or another run's pull), and every trial resolves FROM when +# it builds, so later trials would silently switch base. FROM cannot name an +# image id, so tag this one under a name private to this process and remove it +# on exit. Compose reads the variable on `harbor run` and on `harbor jobs resume`. +runtime_id="$(docker image inspect --format '{{.Id}}' "$runtime_image")" +runtime_id="${runtime_id#sha256:}" +runtime_pin="aob-runtime-pin:${runtime_id:0:12}-$$" +docker tag "$runtime_image" "$runtime_pin" +trap 'docker rmi "$runtime_pin" >/dev/null 2>&1 || true' EXIT +export AOB_RUNTIME_IMAGE="$runtime_pin" +printf 'Runtime image: %s (%s)\n' "$runtime_image" "${runtime_id:0:12}" + # The code sandbox image, as a tar each trial's Docker-in-Docker daemon loads # (benchmarks/harbor/overlays/code-sandbox.yaml). if [[ ! -s "$code_tar" ]]; then @@ -175,6 +203,10 @@ except Exception as exc: PY } +# Non-zero when a model's job could not run or resume. A trial that fails +# inside a job does not count: Harbor records it and the loop moves on. +status=0 + for model_config in "${model_configs[@]}"; do read -r model_id reasoning_effort <<< "$model_config" [[ -z "${model_id:-}" ]] && continue @@ -182,6 +214,11 @@ for model_config in "${model_configs[@]}"; do model_slug="$(printf '%s' "$model_id" | tr -c 'A-Za-z0-9._-' '-' | tr -s '-')" job_name="stirrup_agent__${model_slug%-}" job_path="$jobs_dir/$job_name" + # The runtime image a job started on, beside the job rather than in it. + # Harbor's resume lock covers the task files but not the base they build + # FROM, so without this a resume with another -r would mix two images in + # one job. + image_record="$jobs_dir/$job_name.runtime-image" if ! router_reachable "$model_id"; then echo "Skipping $model_id: its router is unreachable" >&2 @@ -191,6 +228,14 @@ for model_config in "${model_configs[@]}"; do echo "Running $model_id with reasoning effort ${reasoning_effort:-default} -> $job_path" if [[ -f "$job_path/config.json" ]]; then + if [[ -f "$image_record" ]] && [[ "$(cut -f1 "$image_record")" != "$runtime_id" ]]; then + printf '%s started on runtime image %s, not %s (%s).\n' \ + "$job_path" "$(cut -f2 "$image_record")" "$runtime_image" "${runtime_id:0:12}" >&2 + printf 'Pass that image as -r to finish it, or move the job aside to rerun %s.\n' \ + "$model_id" >&2 + status=1 + continue + fi # Drop trials whose agent crashed (e.g. the model was unreachable) so # they run again; scored trials are kept, as run.sh's --skip-existing # kept scenarios that already had a trajectory. @@ -201,10 +246,14 @@ for model_config in "${model_configs[@]}"; do --filter-error-type NonZeroAgentExitCodeError; then printf 'Could not resume %s. If its tasks or overlays changed since it\n' "$job_path" >&2 printf 'started, move it aside to rerun %s from scratch.\n' "$model_id" >&2 + status=1 fi continue fi + mkdir -p "$jobs_dir" + printf '%s\t%s\n' "$runtime_id" "$runtime_image" >"$image_record" + effort_args=() if [[ -n "${reasoning_effort:-}" ]]; then effort_args=(--ak "reasoning_effort=$reasoning_effort") @@ -227,3 +276,5 @@ for model_config in "${model_configs[@]}"; do --job-name "$job_name" \ -o "$jobs_dir" || true done + +exit "$status" diff --git a/benchmarks/harbor/scripts/build-runtime-image.sh b/benchmarks/harbor/scripts/build-runtime-image.sh index e92a12af1..1746f1d2a 100755 --- a/benchmarks/harbor/scripts/build-runtime-image.sh +++ b/benchmarks/harbor/scripts/build-runtime-image.sh @@ -9,10 +9,14 @@ # the image is tagged twice: assetopsbench/runtime:dev, the tag the task # template builds FROM by default, and assetopsbench/runtime:, the # short hash of the HEAD it was built from. :dev moves with every build; the -# commit tag does not, so `run.sh -r assetopsbench/runtime:` pins a run -# to one build. Unless they name an output (--push, --output, --load) it is -# loaded into the local image store. publish-images.sh passes --platform, its -# own tags and --push. +# commit tag moves only when that same commit is rebuilt into a different image +# (e.g. --no-cache, or after the build cache is pruned), so +# `run.sh -r assetopsbench/runtime:` names the code a run used. Each +# commit tag keeps its multi-GB image alive; list them with +# `docker images assetopsbench/runtime` and `docker rmi` the ones you no longer +# need. Unless they name an output (--push, --output, --load) it is loaded into +# the local image store. publish-images.sh passes --platform, its own tags and +# --push. # # Why an archive: base-image/Dockerfile does `COPY . .`, so a build from the # working tree takes whatever sits there, untracked and git-ignored files @@ -52,7 +56,7 @@ has_tag=false has_output=false for arg in "$@"; do case "$arg" in - -t | --tag | --tag=*) has_tag=true ;; + -t* | --tag | --tag=*) has_tag=true ;; --push | --load | --output | --output=* | -o) has_output=true ;; esac done diff --git a/benchmarks/harbor/scripts/publish-images.sh b/benchmarks/harbor/scripts/publish-images.sh index 6ede0e0ce..2f10e945e 100755 --- a/benchmarks/harbor/scripts/publish-images.sh +++ b/benchmarks/harbor/scripts/publish-images.sh @@ -65,14 +65,16 @@ docker buildx build \ cat <