diff --git a/benchmarks/harbor/QUICKSTART.md b/benchmarks/harbor/QUICKSTART.md index f6780d5ca..27db5f5f5 100644 --- a/benchmarks/harbor/QUICKSTART.md +++ b/benchmarks/harbor/QUICKSTART.md @@ -23,11 +23,12 @@ uv sync --dev --extra harbor ## 2. Get the runtime image -Every task image layers its scenario onto one shared runtime image. Pull it: +Every task image layers its scenario onto one shared runtime image. Pull the +published one (currently linux/arm64 only) and point the tasks at it: ```bash -docker pull assetopsbench/runtime:latest -docker tag assetopsbench/runtime:latest assetopsbench/runtime:dev +docker pull quay.io/assetopsbench/runtime:dev +export AOB_RUNTIME_IMAGE=quay.io/assetopsbench/runtime:dev ``` Or build it yourself, which takes a few minutes and needs no registry: @@ -36,14 +37,19 @@ Or build it yourself, which takes a few minutes and needs no registry: bash benchmarks/harbor/scripts/build-runtime-image.sh ``` -The script builds `assetopsbench/runtime:dev` from `git archive HEAD`, not -from your working tree, so untracked files such as local results never reach +The script builds `assetopsbench/runtime:dev`, also tagged +`assetopsbench/runtime:`, from `git archive HEAD`, not from your +working tree, so untracked files such as local results never reach the container the agent runs in. Commit a change first to include it. -Either way the local tag `assetopsbench/runtime:dev` is what the task -Dockerfiles reference, through the `AOB_RUNTIME_IMAGE` build arg in -`benchmarks/harbor/template/environment/Dockerfile`. Change that default if you -publish under a different namespace. +Each task's `environment/docker-compose.yaml` passes `AOB_RUNTIME_IMAGE` to +its Dockerfile as a build arg, so `harbor run` builds FROM whatever that +variable names when it runs, and from the local `assetopsbench/runtime:dev` +when it is unset. The variable lives in the shell, so set it again in a new +one, or put it in `.env` and run Harbor as `uv run --env-file .env harbor run`. +Changing the image needs no task regeneration, but tasks generated before the +build arg existed ignore the variable; regenerate those once (step 3). +`run.sh` takes the image as `-r`. ## 3. Generate the tasks @@ -167,9 +173,11 @@ reach CouchDB. Confirm your checkout includes the `env` entry in `StirrupAgentRunner._build_mcp_config`, without which the MCP SDK hands each server a six-variable environment that omits `COUCHDB_URL`. -**Trials fail instantly with a pull error** — the runtime image is missing, or -it is tagged under a name the task Dockerfile does not reference. Redo step 2 -and check `docker images assetopsbench/runtime`. +**Trials fail instantly with a pull error** — the build could not find its +base. Either `AOB_RUNTIME_IMAGE` is unset in this shell, so the build fell back +to the local `assetopsbench/runtime:dev`, which does not exist; or it names an +image that is not local (check `echo $AOB_RUNTIME_IMAGE` and +`docker images`). Redo step 2 in this shell. **`No module named 'google.protobuf'` during a run** — the image was built without the `otel` dependency group, which the file trace exporter needs. diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index 20e266eac..a46a7e1a4 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -41,8 +41,12 @@ command from the repo root. ```bash # 1. Build the runtime base (once per AssetOpsBench commit). The script builds # from `git archive HEAD`, so untracked files and uncommitted changes stay -# out of the image. The tag is local and is the default AOB_RUNTIME_IMAGE in -# template/environment/Dockerfile; nothing is pulled or pushed. +# out of the image. The tag is local and is the default AOB_RUNTIME_IMAGE +# that each task's docker-compose.yaml passes to its Dockerfile; nothing is +# pulled or pushed. It also tags assetopsbench/runtime:, which +# later commits' builds do not move. To use a published image or a commit +# tag, `export AOB_RUNTIME_IMAGE=` before `harbor run`; tasks +# generated before that build arg existed ignore it, so regenerate them. bash benchmarks/harbor/scripts/build-runtime-image.sh # 2. Generate one task per scenario in the open profile. The defaults point at @@ -83,18 +87,29 @@ bash benchmarks/harbor/run.sh \ ``` `-m` may repeat; without it the script runs `benchmarks/run.sh`'s model list. -`-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), and `-n` -the number of concurrent trials. The script: - -1. exports `AOB_PRIVATE_DIR` as the `-s` directory, so +`-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), `-n` +the number of concurrent trials, and `-r` the runtime image (otherwise +`AOB_RUNTIME_IMAGE` from the shell, then from `.env`, then +`assetopsbench/runtime:dev`). The script: + +1. resolves the runtime image, which every task image builds FROM. A published + image, one whose local copy came from that registry or that is not local at + all, is pulled first, so the run uses the current publish rather than a + stale copy; a local build such as the default is used as is. It then pins + that exact image for the whole run, so a rebuild or pull of the same tag + during the run does not switch later trials' base, and records it beside + each job: a job started on another image is not resumed. The image changes + only the base: each task still adds its own `manifest.json`, and `shared/` + still comes from the mount below; +2. exports `AOB_PRIVATE_DIR` as the `-s` directory, so `overlays/private-data.yaml` mounts its `shared/` at `/opt/suite/scenarios_data/shared`. Only `init_data.py` reads it, as in `scenario_suite_runner`; -2. builds the code sandbox image and saves it to `~/assetops-code.tar` for the +3. builds the code sandbox image and saves it to `~/assetops-code.tar` for the per-trial Docker-in-Docker daemon (`overlays/code-sandbox.yaml`); -3. generates one task per scenario with `--scenario-root` and +4. generates one task per scenario with `--scenario-root` and `--skip-missing`, skipping profile entries the suite lacks; -4. runs one Harbor job per model at +5. runs one Harbor job per model at `/harbor-jobs/stirrup_agent__`, with both overlays and with credentials loaded from `.env` by `uv run --env-file` into the Harbor process only. @@ -103,7 +118,8 @@ Re-running resumes an existing job and finishes only its incomplete trials, the equivalent of `--skip-existing`. Harbor refuses to resume a job whose tasks or overlays have changed since it started, such as one from before `run.sh` mounted only `shared/`; the script says so, and moving that job aside reruns -the model from scratch. Each trial with the code sandbox runs a +the model from scratch. The script exits non-zero when any model's job could +not start or resume, so a wrapper sees it. Each trial with the code sandbox runs a privileged `dind` sidecar, so keep `-n` around 4 on a laptop-sized Docker VM. ## Generating other profiles by hand diff --git a/benchmarks/harbor/run.sh b/benchmarks/harbor/run.sh index 16d5dbf47..17f2c63a0 100755 --- a/benchmarks/harbor/run.sh +++ b/benchmarks/harbor/run.sh @@ -7,12 +7,17 @@ # shared database. # # bash benchmarks/harbor/run.sh -s SCENARIO_DIR -l LEADERBOARD_DIR \ -# [-n N_CONCURRENT] [-p PROFILE] [-m "MODEL_ID REASONING_EFFORT"]... +# [-n N_CONCURRENT] [-p PROFILE] [-r RUNTIME_IMAGE] \ +# [-m "MODEL_ID REASONING_EFFORT"]... # -# Prerequisites: Docker running, `uv sync --extra harbor`, and the runtime image +# Prerequisites: Docker running, `uv sync --extra harbor`, and the runtime image, +# either built locally # # bash benchmarks/harbor/scripts/build-runtime-image.sh # +# or published, passed as -r (or AOB_RUNTIME_IMAGE in the shell or ENV_FILE), +# e.g. -r quay.io/assetopsbench/runtime:dev. +# # Credentials are read from ENV_FILE (default .env) by `uv run --env-file`, into # the Harbor process only; StirrupAgent forwards them to the agent phase. They # never enter an image. @@ -24,7 +29,7 @@ set -euo pipefail usage() { - printf 'Usage: %s -s SCENARIO_DIR -l LEADERBOARD_DIR [-n N_CONCURRENT] [-p PROFILE] [-m "MODEL_ID EFFORT"]...\n' "$0" >&2 + printf 'Usage: %s -s SCENARIO_DIR -l LEADERBOARD_DIR [-n N_CONCURRENT] [-p PROFILE] [-r RUNTIME_IMAGE] [-m "MODEL_ID EFFORT"]...\n' "$0" >&2 } scenario_dir="${SCENARIO_DIR:-}" @@ -32,14 +37,16 @@ leaderboard_dir="${LEADERBOARD_DIR:-}" n_concurrent="${N_CONCURRENT:-4}" profile="${PROFILE:-benchmarks/scenario_suite/all.yaml}" env_file="${ENV_FILE:-.env}" +runtime_image="${AOB_RUNTIME_IMAGE:-}" model_configs=() -while getopts ':s:l:n:p:m:' option; do +while getopts ':s:l:n:p:r:m:' option; do case "$option" in s) scenario_dir="$OPTARG" ;; l) leaderboard_dir="$OPTARG" ;; n) n_concurrent="$OPTARG" ;; p) profile="$OPTARG" ;; + r) runtime_image="$OPTARG" ;; m) model_configs+=("$OPTARG") ;; :) printf 'Option -%s requires an argument.\n' "$OPTARG" >&2; usage; exit 2 ;; \?) printf 'Unknown option: -%s\n' "$OPTARG" >&2; usage; exit 2 ;; @@ -77,17 +84,10 @@ if [[ ! -f "$env_file" ]]; then exit 2 fi -runtime_image=assetopsbench/runtime:dev code_image=assetops-code:dev code_tar="${AOB_CODE_TAR:-$HOME/assetops-code.tar}" dataset_dir=benchmarks/harbor/datasets/assetopsbench-suite -if ! docker image inspect "$runtime_image" >/dev/null 2>&1; then - printf 'Runtime image %s not found. Build it first:\n' "$runtime_image" >&2 - printf ' bash benchmarks/harbor/scripts/build-runtime-image.sh\n' >&2 - exit 1 -fi - # The suite's shared/ data reaches each trial through # overlays/private-data.yaml, a read-only bind mount of this directory's shared/. # Compose reads the variable on `harbor run` and again on `harbor jobs resume`. @@ -99,6 +99,63 @@ if [[ ! -d "$scenario_dir/shared" ]]; then fi export AOB_PRIVATE_DIR="$scenario_dir" +# Every task image builds FROM the runtime image, through the AOB_RUNTIME_IMAGE +# build arg in the task's docker-compose.yaml. -r wins, then the shell's +# AOB_RUNTIME_IMAGE, then ENV_FILE's, then the local default: the order +# `uv run --env-file` gives Harbor, where the environment beats the file. +if [[ -z "$runtime_image" ]]; then + runtime_image="$(uv run --env-file "$env_file" python -c \ + 'import os; print(os.environ.get("AOB_RUNTIME_IMAGE", ""))')" +fi +runtime_image="${runtime_image:-assetopsbench/runtime:dev}" + +# The reference without its tag or digest, spelled as .RepoDigests spells it. +image_repo() { + local ref="${1%@*}" + if [[ "${ref##*/}" == *:* ]]; then ref="${ref%:*}"; fi + ref="${ref#docker.io/}" + printf '%s' "${ref#library/}" +} + +# True when the local copy of $1 came from (or went to) that same repository, +# i.e. it is a published image rather than a local build. +from_registry() { + local repo digest + repo="$(image_repo "$1")" + while read -r digest; do + [[ "${digest%@*}" == "$repo" ]] && return 0 + done < <(docker image inspect --format '{{range .RepoDigests}}{{println .}}{{end}}' "$1") + return 1 +} + +# A local copy satisfies FROM, so the build never refreshes a published image; +# pull it here instead, on Docker Hub or any other registry. A local build that +# was never pushed under this name (the default assetopsbench/runtime:dev) is +# used as is, and a pull never replaces it. +if ! docker image inspect "$runtime_image" >/dev/null 2>&1; then + if ! docker pull "$runtime_image"; then + printf 'Runtime image %s is not local and could not be pulled. Build it with\n' "$runtime_image" >&2 + printf ' bash benchmarks/harbor/scripts/build-runtime-image.sh\n' >&2 + exit 1 + fi +elif from_registry "$runtime_image" && ! docker pull "$runtime_image"; then + printf 'warning: could not pull %s; using the local copy, which may be stale\n' \ + "$runtime_image" >&2 +fi + +# Pin the base for the whole run. A tag such as :dev can move while the run is +# going (a rebuild, or another run's pull), and every trial resolves FROM when +# it builds, so later trials would silently switch base. FROM cannot name an +# image id, so tag this one under a name private to this process and remove it +# on exit. Compose reads the variable on `harbor run` and on `harbor jobs resume`. +runtime_id="$(docker image inspect --format '{{.Id}}' "$runtime_image")" +runtime_id="${runtime_id#sha256:}" +runtime_pin="aob-runtime-pin:${runtime_id:0:12}-$$" +docker tag "$runtime_image" "$runtime_pin" +trap 'docker rmi "$runtime_pin" >/dev/null 2>&1 || true' EXIT +export AOB_RUNTIME_IMAGE="$runtime_pin" +printf 'Runtime image: %s (%s)\n' "$runtime_image" "${runtime_id:0:12}" + # The code sandbox image, as a tar each trial's Docker-in-Docker daemon loads # (benchmarks/harbor/overlays/code-sandbox.yaml). if [[ ! -s "$code_tar" ]]; then @@ -146,6 +203,10 @@ except Exception as exc: PY } +# Non-zero when a model's job could not run or resume. A trial that fails +# inside a job does not count: Harbor records it and the loop moves on. +status=0 + for model_config in "${model_configs[@]}"; do read -r model_id reasoning_effort <<< "$model_config" [[ -z "${model_id:-}" ]] && continue @@ -153,6 +214,11 @@ for model_config in "${model_configs[@]}"; do model_slug="$(printf '%s' "$model_id" | tr -c 'A-Za-z0-9._-' '-' | tr -s '-')" job_name="stirrup_agent__${model_slug%-}" job_path="$jobs_dir/$job_name" + # The runtime image a job started on, beside the job rather than in it. + # Harbor's resume lock covers the task files but not the base they build + # FROM, so without this a resume with another -r would mix two images in + # one job. + image_record="$jobs_dir/$job_name.runtime-image" if ! router_reachable "$model_id"; then echo "Skipping $model_id: its router is unreachable" >&2 @@ -162,6 +228,14 @@ for model_config in "${model_configs[@]}"; do echo "Running $model_id with reasoning effort ${reasoning_effort:-default} -> $job_path" if [[ -f "$job_path/config.json" ]]; then + if [[ -f "$image_record" ]] && [[ "$(cut -f1 "$image_record")" != "$runtime_id" ]]; then + printf '%s started on runtime image %s, not %s (%s).\n' \ + "$job_path" "$(cut -f2 "$image_record")" "$runtime_image" "${runtime_id:0:12}" >&2 + printf 'Pass that image as -r to finish it, or move the job aside to rerun %s.\n' \ + "$model_id" >&2 + status=1 + continue + fi # Drop trials whose agent crashed (e.g. the model was unreachable) so # they run again; scored trials are kept, as run.sh's --skip-existing # kept scenarios that already had a trajectory. @@ -172,10 +246,14 @@ for model_config in "${model_configs[@]}"; do --filter-error-type NonZeroAgentExitCodeError; then printf 'Could not resume %s. If its tasks or overlays changed since it\n' "$job_path" >&2 printf 'started, move it aside to rerun %s from scratch.\n' "$model_id" >&2 + status=1 fi continue fi + mkdir -p "$jobs_dir" + printf '%s\t%s\n' "$runtime_id" "$runtime_image" >"$image_record" + effort_args=() if [[ -n "${reasoning_effort:-}" ]]; then effort_args=(--ak "reasoning_effort=$reasoning_effort") @@ -198,3 +276,5 @@ for model_config in "${model_configs[@]}"; do --job-name "$job_name" \ -o "$jobs_dir" || true done + +exit "$status" diff --git a/benchmarks/harbor/scripts/build-runtime-image.sh b/benchmarks/harbor/scripts/build-runtime-image.sh index 88d7b6576..1746f1d2a 100755 --- a/benchmarks/harbor/scripts/build-runtime-image.sh +++ b/benchmarks/harbor/scripts/build-runtime-image.sh @@ -3,13 +3,20 @@ # # bash benchmarks/harbor/scripts/build-runtime-image.sh # bash benchmarks/harbor/scripts/build-runtime-image.sh --no-cache -# bash benchmarks/harbor/scripts/build-runtime-image.sh -t assetopsbench/runtime:abc1234 +# bash benchmarks/harbor/scripts/build-runtime-image.sh -t myorg/runtime:test # # Arguments go to `docker buildx build` unchanged. Unless they name a tag (-t) -# the image is tagged assetopsbench/runtime:dev, the tag the task template -# builds FROM, and unless they name an output (--push, --output, --load) it is -# loaded into the local image store. publish-images.sh passes --platform, its -# own tags and --push. +# the image is tagged twice: assetopsbench/runtime:dev, the tag the task +# template builds FROM by default, and assetopsbench/runtime:, the +# short hash of the HEAD it was built from. :dev moves with every build; the +# commit tag moves only when that same commit is rebuilt into a different image +# (e.g. --no-cache, or after the build cache is pruned), so +# `run.sh -r assetopsbench/runtime:` names the code a run used. Each +# commit tag keeps its multi-GB image alive; list them with +# `docker images assetopsbench/runtime` and `docker rmi` the ones you no longer +# need. Unless they name an output (--push, --output, --load) it is loaded into +# the local image store. publish-images.sh passes --platform, its own tags and +# --push. # # Why an archive: base-image/Dockerfile does `COPY . .`, so a build from the # working tree takes whatever sits there, untracked and git-ignored files @@ -49,12 +56,12 @@ has_tag=false has_output=false for arg in "$@"; do case "$arg" in - -t | --tag | --tag=*) has_tag=true ;; + -t* | --tag | --tag=*) has_tag=true ;; --push | --load | --output | --output=* | -o) has_output=true ;; esac done defaults=() -$has_tag || defaults+=(-t assetopsbench/runtime:dev) +$has_tag || defaults+=(-t assetopsbench/runtime:dev -t "assetopsbench/runtime:${commit:0:7}") $has_output || defaults+=(--load) printf 'Building the runtime image from %s (git archive)\n' "${commit:0:7}" >&2 diff --git a/benchmarks/harbor/scripts/publish-images.sh b/benchmarks/harbor/scripts/publish-images.sh index 25a793189..2f10e945e 100755 --- a/benchmarks/harbor/scripts/publish-images.sh +++ b/benchmarks/harbor/scripts/publish-images.sh @@ -42,10 +42,16 @@ if ! docker buildx inspect >/dev/null 2>&1; then exit 1 fi -echo "==> ${NAMESPACE}/runtime:${TAG} for ${PLATFORMS}" +# The commit tag never moves, unlike TAG and latest, so a published run can +# name the exact build it used. build-runtime-image.sh builds HEAD, so this is +# the commit in the image whatever the working tree holds. +COMMIT="$(git rev-parse HEAD | cut -c1-7)" + +echo "==> ${NAMESPACE}/runtime:${TAG} (${COMMIT}) for ${PLATFORMS}" bash benchmarks/harbor/scripts/build-runtime-image.sh \ --platform "${PLATFORMS}" \ -t "${NAMESPACE}/runtime:${TAG}" \ + -t "${NAMESPACE}/runtime:${COMMIT}" \ -t "${NAMESPACE}/runtime:latest" \ --push @@ -59,16 +65,16 @@ docker buildx build \ cat <