Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 20 additions & 12 deletions benchmarks/harbor/QUICKSTART.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,11 +23,12 @@ uv sync --dev --extra harbor

## 2. Get the runtime image

Every task image layers its scenario onto one shared runtime image. Pull it:
Every task image layers its scenario onto one shared runtime image. Pull the
published one (currently linux/arm64 only) and point the tasks at it:

```bash
docker pull assetopsbench/runtime:latest
docker tag assetopsbench/runtime:latest assetopsbench/runtime:dev
docker pull quay.io/assetopsbench/runtime:dev
export AOB_RUNTIME_IMAGE=quay.io/assetopsbench/runtime:dev
```

Or build it yourself, which takes a few minutes and needs no registry:
Expand All @@ -36,14 +37,19 @@ Or build it yourself, which takes a few minutes and needs no registry:
bash benchmarks/harbor/scripts/build-runtime-image.sh
```

The script builds `assetopsbench/runtime:dev` from `git archive HEAD`, not
from your working tree, so untracked files such as local results never reach
The script builds `assetopsbench/runtime:dev`, also tagged
`assetopsbench/runtime:<commit>`, from `git archive HEAD`, not from your
working tree, so untracked files such as local results never reach
the container the agent runs in. Commit a change first to include it.

Either way the local tag `assetopsbench/runtime:dev` is what the task
Dockerfiles reference, through the `AOB_RUNTIME_IMAGE` build arg in
`benchmarks/harbor/template/environment/Dockerfile`. Change that default if you
publish under a different namespace.
Each task's `environment/docker-compose.yaml` passes `AOB_RUNTIME_IMAGE` to
its Dockerfile as a build arg, so `harbor run` builds FROM whatever that
variable names when it runs, and from the local `assetopsbench/runtime:dev`
when it is unset. The variable lives in the shell, so set it again in a new
one, or put it in `.env` and run Harbor as `uv run --env-file .env harbor run`.
Changing the image needs no task regeneration, but tasks generated before the
build arg existed ignore the variable; regenerate those once (step 3).
`run.sh` takes the image as `-r`.

## 3. Generate the tasks

Expand Down Expand Up @@ -167,9 +173,11 @@ reach CouchDB. Confirm your checkout includes the `env` entry in
`StirrupAgentRunner._build_mcp_config`, without which the MCP SDK hands each
server a six-variable environment that omits `COUCHDB_URL`.

**Trials fail instantly with a pull error** — the runtime image is missing, or
it is tagged under a name the task Dockerfile does not reference. Redo step 2
and check `docker images assetopsbench/runtime`.
**Trials fail instantly with a pull error** — the build could not find its
base. Either `AOB_RUNTIME_IMAGE` is unset in this shell, so the build fell back
to the local `assetopsbench/runtime:dev`, which does not exist; or it names an
image that is not local (check `echo $AOB_RUNTIME_IMAGE` and
`docker images`). Redo step 2 in this shell.

**`No module named 'google.protobuf'` during a run** — the image was built
without the `otel` dependency group, which the file trace exporter needs.
Expand Down
36 changes: 26 additions & 10 deletions benchmarks/harbor/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,8 +41,12 @@ command from the repo root.
```bash
# 1. Build the runtime base (once per AssetOpsBench commit). The script builds
# from `git archive HEAD`, so untracked files and uncommitted changes stay
# out of the image. The tag is local and is the default AOB_RUNTIME_IMAGE in
# template/environment/Dockerfile; nothing is pulled or pushed.
# out of the image. The tag is local and is the default AOB_RUNTIME_IMAGE
# that each task's docker-compose.yaml passes to its Dockerfile; nothing is
# pulled or pushed. It also tags assetopsbench/runtime:<commit>, which
# later commits' builds do not move. To use a published image or a commit
# tag, `export AOB_RUNTIME_IMAGE=<image>` before `harbor run`; tasks
# generated before that build arg existed ignore it, so regenerate them.
bash benchmarks/harbor/scripts/build-runtime-image.sh

# 2. Generate one task per scenario in the open profile. The defaults point at
Expand Down Expand Up @@ -83,18 +87,29 @@ bash benchmarks/harbor/run.sh \
```

`-m` may repeat; without it the script runs `benchmarks/run.sh`'s model list.
`-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), and `-n`
the number of concurrent trials. The script:

1. exports `AOB_PRIVATE_DIR` as the `-s` directory, so
`-p` picks the profile (default `benchmarks/scenario_suite/all.yaml`), `-n`
the number of concurrent trials, and `-r` the runtime image (otherwise
`AOB_RUNTIME_IMAGE` from the shell, then from `.env`, then
`assetopsbench/runtime:dev`). The script:

1. resolves the runtime image, which every task image builds FROM. A published
image, one whose local copy came from that registry or that is not local at
all, is pulled first, so the run uses the current publish rather than a
stale copy; a local build such as the default is used as is. It then pins
that exact image for the whole run, so a rebuild or pull of the same tag
during the run does not switch later trials' base, and records it beside
each job: a job started on another image is not resumed. The image changes
only the base: each task still adds its own `manifest.json`, and `shared/`
still comes from the mount below;
2. exports `AOB_PRIVATE_DIR` as the `-s` directory, so
`overlays/private-data.yaml` mounts its `shared/` at
`/opt/suite/scenarios_data/shared`. Only `init_data.py` reads it, as in
`scenario_suite_runner`;
2. builds the code sandbox image and saves it to `~/assetops-code.tar` for the
3. builds the code sandbox image and saves it to `~/assetops-code.tar` for the
per-trial Docker-in-Docker daemon (`overlays/code-sandbox.yaml`);
3. generates one task per scenario with `--scenario-root` and
4. generates one task per scenario with `--scenario-root` and
`--skip-missing`, skipping profile entries the suite lacks;
4. runs one Harbor job per model at
5. runs one Harbor job per model at
`<leaderboard>/harbor-jobs/stirrup_agent__<model>`, with both overlays and
with credentials loaded from `.env` by `uv run --env-file` into the Harbor
process only.
Expand All @@ -103,7 +118,8 @@ Re-running resumes an existing job and finishes only its incomplete trials,
the equivalent of `--skip-existing`. Harbor refuses to resume a job whose tasks
or overlays have changed since it started, such as one from before `run.sh`
mounted only `shared/`; the script says so, and moving that job aside reruns
the model from scratch. Each trial with the code sandbox runs a
the model from scratch. The script exits non-zero when any model's job could
not start or resume, so a wrapper sees it. Each trial with the code sandbox runs a
privileged `dind` sidecar, so keep `-n` around 4 on a laptop-sized Docker VM.

## Generating other profiles by hand
Expand Down
102 changes: 91 additions & 11 deletions benchmarks/harbor/run.sh
Original file line number Diff line number Diff line change
Expand Up @@ -7,12 +7,17 @@
# shared database.
#
# bash benchmarks/harbor/run.sh -s SCENARIO_DIR -l LEADERBOARD_DIR \
# [-n N_CONCURRENT] [-p PROFILE] [-m "MODEL_ID REASONING_EFFORT"]...
# [-n N_CONCURRENT] [-p PROFILE] [-r RUNTIME_IMAGE] \
# [-m "MODEL_ID REASONING_EFFORT"]...
#
# Prerequisites: Docker running, `uv sync --extra harbor`, and the runtime image
# Prerequisites: Docker running, `uv sync --extra harbor`, and the runtime image,
# either built locally
#
# bash benchmarks/harbor/scripts/build-runtime-image.sh
#
# or published, passed as -r (or AOB_RUNTIME_IMAGE in the shell or ENV_FILE),
# e.g. -r quay.io/assetopsbench/runtime:dev.
#
# Credentials are read from ENV_FILE (default .env) by `uv run --env-file`, into
# the Harbor process only; StirrupAgent forwards them to the agent phase. They
# never enter an image.
Expand All @@ -24,22 +29,24 @@
set -euo pipefail

usage() {
printf 'Usage: %s -s SCENARIO_DIR -l LEADERBOARD_DIR [-n N_CONCURRENT] [-p PROFILE] [-m "MODEL_ID EFFORT"]...\n' "$0" >&2
printf 'Usage: %s -s SCENARIO_DIR -l LEADERBOARD_DIR [-n N_CONCURRENT] [-p PROFILE] [-r RUNTIME_IMAGE] [-m "MODEL_ID EFFORT"]...\n' "$0" >&2
}

scenario_dir="${SCENARIO_DIR:-}"
leaderboard_dir="${LEADERBOARD_DIR:-}"
n_concurrent="${N_CONCURRENT:-4}"
profile="${PROFILE:-benchmarks/scenario_suite/all.yaml}"
env_file="${ENV_FILE:-.env}"
runtime_image="${AOB_RUNTIME_IMAGE:-}"
model_configs=()

while getopts ':s:l:n:p:m:' option; do
while getopts ':s:l:n:p:r:m:' option; do
case "$option" in
s) scenario_dir="$OPTARG" ;;
l) leaderboard_dir="$OPTARG" ;;
n) n_concurrent="$OPTARG" ;;
p) profile="$OPTARG" ;;
r) runtime_image="$OPTARG" ;;
m) model_configs+=("$OPTARG") ;;
:) printf 'Option -%s requires an argument.\n' "$OPTARG" >&2; usage; exit 2 ;;
\?) printf 'Unknown option: -%s\n' "$OPTARG" >&2; usage; exit 2 ;;
Expand Down Expand Up @@ -77,17 +84,10 @@ if [[ ! -f "$env_file" ]]; then
exit 2
fi

runtime_image=assetopsbench/runtime:dev
code_image=assetops-code:dev
code_tar="${AOB_CODE_TAR:-$HOME/assetops-code.tar}"
dataset_dir=benchmarks/harbor/datasets/assetopsbench-suite

if ! docker image inspect "$runtime_image" >/dev/null 2>&1; then
printf 'Runtime image %s not found. Build it first:\n' "$runtime_image" >&2
printf ' bash benchmarks/harbor/scripts/build-runtime-image.sh\n' >&2
exit 1
fi

# The suite's shared/ data reaches each trial through
# overlays/private-data.yaml, a read-only bind mount of this directory's shared/.
# Compose reads the variable on `harbor run` and again on `harbor jobs resume`.
Expand All @@ -99,6 +99,63 @@ if [[ ! -d "$scenario_dir/shared" ]]; then
fi
export AOB_PRIVATE_DIR="$scenario_dir"

# Every task image builds FROM the runtime image, through the AOB_RUNTIME_IMAGE
# build arg in the task's docker-compose.yaml. -r wins, then the shell's
# AOB_RUNTIME_IMAGE, then ENV_FILE's, then the local default: the order
# `uv run --env-file` gives Harbor, where the environment beats the file.
if [[ -z "$runtime_image" ]]; then
runtime_image="$(uv run --env-file "$env_file" python -c \
'import os; print(os.environ.get("AOB_RUNTIME_IMAGE", ""))')"
fi
runtime_image="${runtime_image:-assetopsbench/runtime:dev}"

# The reference without its tag or digest, spelled as .RepoDigests spells it.
image_repo() {
local ref="${1%@*}"
if [[ "${ref##*/}" == *:* ]]; then ref="${ref%:*}"; fi
ref="${ref#docker.io/}"
printf '%s' "${ref#library/}"
}

# True when the local copy of $1 came from (or went to) that same repository,
# i.e. it is a published image rather than a local build.
from_registry() {
local repo digest
repo="$(image_repo "$1")"
while read -r digest; do
[[ "${digest%@*}" == "$repo" ]] && return 0
done < <(docker image inspect --format '{{range .RepoDigests}}{{println .}}{{end}}' "$1")
return 1
}

# A local copy satisfies FROM, so the build never refreshes a published image;
# pull it here instead, on Docker Hub or any other registry. A local build that
# was never pushed under this name (the default assetopsbench/runtime:dev) is
# used as is, and a pull never replaces it.
if ! docker image inspect "$runtime_image" >/dev/null 2>&1; then
if ! docker pull "$runtime_image"; then
printf 'Runtime image %s is not local and could not be pulled. Build it with\n' "$runtime_image" >&2
printf ' bash benchmarks/harbor/scripts/build-runtime-image.sh\n' >&2
exit 1
fi
elif from_registry "$runtime_image" && ! docker pull "$runtime_image"; then
printf 'warning: could not pull %s; using the local copy, which may be stale\n' \
"$runtime_image" >&2
fi

# Pin the base for the whole run. A tag such as :dev can move while the run is
# going (a rebuild, or another run's pull), and every trial resolves FROM when
# it builds, so later trials would silently switch base. FROM cannot name an
# image id, so tag this one under a name private to this process and remove it
# on exit. Compose reads the variable on `harbor run` and on `harbor jobs resume`.
runtime_id="$(docker image inspect --format '{{.Id}}' "$runtime_image")"
runtime_id="${runtime_id#sha256:}"
runtime_pin="aob-runtime-pin:${runtime_id:0:12}-$$"
docker tag "$runtime_image" "$runtime_pin"
trap 'docker rmi "$runtime_pin" >/dev/null 2>&1 || true' EXIT
export AOB_RUNTIME_IMAGE="$runtime_pin"
printf 'Runtime image: %s (%s)\n' "$runtime_image" "${runtime_id:0:12}"

# The code sandbox image, as a tar each trial's Docker-in-Docker daemon loads
# (benchmarks/harbor/overlays/code-sandbox.yaml).
if [[ ! -s "$code_tar" ]]; then
Expand Down Expand Up @@ -146,13 +203,22 @@ except Exception as exc:
PY
}

# Non-zero when a model's job could not run or resume. A trial that fails
# inside a job does not count: Harbor records it and the loop moves on.
status=0

for model_config in "${model_configs[@]}"; do
read -r model_id reasoning_effort <<< "$model_config"
[[ -z "${model_id:-}" ]] && continue

model_slug="$(printf '%s' "$model_id" | tr -c 'A-Za-z0-9._-' '-' | tr -s '-')"
job_name="stirrup_agent__${model_slug%-}"
job_path="$jobs_dir/$job_name"
# The runtime image a job started on, beside the job rather than in it.
# Harbor's resume lock covers the task files but not the base they build
# FROM, so without this a resume with another -r would mix two images in
# one job.
image_record="$jobs_dir/$job_name.runtime-image"

if ! router_reachable "$model_id"; then
echo "Skipping $model_id: its router is unreachable" >&2
Expand All @@ -162,6 +228,14 @@ for model_config in "${model_configs[@]}"; do
echo "Running $model_id with reasoning effort ${reasoning_effort:-default} -> $job_path"

if [[ -f "$job_path/config.json" ]]; then
if [[ -f "$image_record" ]] && [[ "$(cut -f1 "$image_record")" != "$runtime_id" ]]; then
printf '%s started on runtime image %s, not %s (%s).\n' \
"$job_path" "$(cut -f2 "$image_record")" "$runtime_image" "${runtime_id:0:12}" >&2
printf 'Pass that image as -r to finish it, or move the job aside to rerun %s.\n' \
"$model_id" >&2
status=1
continue
fi
# Drop trials whose agent crashed (e.g. the model was unreachable) so
# they run again; scored trials are kept, as run.sh's --skip-existing
# kept scenarios that already had a trajectory.
Expand All @@ -172,10 +246,14 @@ for model_config in "${model_configs[@]}"; do
--filter-error-type NonZeroAgentExitCodeError; then
printf 'Could not resume %s. If its tasks or overlays changed since it\n' "$job_path" >&2
printf 'started, move it aside to rerun %s from scratch.\n' "$model_id" >&2
status=1
fi
continue
fi

mkdir -p "$jobs_dir"
printf '%s\t%s\n' "$runtime_id" "$runtime_image" >"$image_record"

effort_args=()
if [[ -n "${reasoning_effort:-}" ]]; then
effort_args=(--ak "reasoning_effort=$reasoning_effort")
Expand All @@ -198,3 +276,5 @@ for model_config in "${model_configs[@]}"; do
--job-name "$job_name" \
-o "$jobs_dir" || true
done

exit "$status"
21 changes: 14 additions & 7 deletions benchmarks/harbor/scripts/build-runtime-image.sh
Original file line number Diff line number Diff line change
Expand Up @@ -3,13 +3,20 @@
#
# bash benchmarks/harbor/scripts/build-runtime-image.sh
# bash benchmarks/harbor/scripts/build-runtime-image.sh --no-cache
# bash benchmarks/harbor/scripts/build-runtime-image.sh -t assetopsbench/runtime:abc1234
# bash benchmarks/harbor/scripts/build-runtime-image.sh -t myorg/runtime:test
#
# Arguments go to `docker buildx build` unchanged. Unless they name a tag (-t)
# the image is tagged assetopsbench/runtime:dev, the tag the task template
# builds FROM, and unless they name an output (--push, --output, --load) it is
# loaded into the local image store. publish-images.sh passes --platform, its
# own tags and --push.
# the image is tagged twice: assetopsbench/runtime:dev, the tag the task
# template builds FROM by default, and assetopsbench/runtime:<commit>, the
# short hash of the HEAD it was built from. :dev moves with every build; the
# commit tag moves only when that same commit is rebuilt into a different image
# (e.g. --no-cache, or after the build cache is pruned), so
# `run.sh -r assetopsbench/runtime:<commit>` names the code a run used. Each
# commit tag keeps its multi-GB image alive; list them with
# `docker images assetopsbench/runtime` and `docker rmi` the ones you no longer
# need. Unless they name an output (--push, --output, --load) it is loaded into
# the local image store. publish-images.sh passes --platform, its own tags and
# --push.
#
# Why an archive: base-image/Dockerfile does `COPY . .`, so a build from the
# working tree takes whatever sits there, untracked and git-ignored files
Expand Down Expand Up @@ -49,12 +56,12 @@ has_tag=false
has_output=false
for arg in "$@"; do
case "$arg" in
-t | --tag | --tag=*) has_tag=true ;;
-t* | --tag | --tag=*) has_tag=true ;;
--push | --load | --output | --output=* | -o) has_output=true ;;
esac
done
defaults=()
$has_tag || defaults+=(-t assetopsbench/runtime:dev)
$has_tag || defaults+=(-t assetopsbench/runtime:dev -t "assetopsbench/runtime:${commit:0:7}")
$has_output || defaults+=(--load)

printf 'Building the runtime image from %s (git archive)\n' "${commit:0:7}" >&2
Expand Down
24 changes: 15 additions & 9 deletions benchmarks/harbor/scripts/publish-images.sh
Original file line number Diff line number Diff line change
Expand Up @@ -42,10 +42,16 @@ if ! docker buildx inspect >/dev/null 2>&1; then
exit 1
fi

echo "==> ${NAMESPACE}/runtime:${TAG} for ${PLATFORMS}"
# The commit tag never moves, unlike TAG and latest, so a published run can
# name the exact build it used. build-runtime-image.sh builds HEAD, so this is
# the commit in the image whatever the working tree holds.
COMMIT="$(git rev-parse HEAD | cut -c1-7)"

echo "==> ${NAMESPACE}/runtime:${TAG} (${COMMIT}) for ${PLATFORMS}"
bash benchmarks/harbor/scripts/build-runtime-image.sh \
--platform "${PLATFORMS}" \
-t "${NAMESPACE}/runtime:${TAG}" \
-t "${NAMESPACE}/runtime:${COMMIT}" \
-t "${NAMESPACE}/runtime:latest" \
--push

Expand All @@ -59,16 +65,16 @@ docker buildx build \

cat <<NOTE

Published. Users now pull instead of building:
Published. Users now pull instead of building. To run on exactly this build:

docker pull ${NAMESPACE}/runtime:latest
docker tag ${NAMESPACE}/runtime:latest assetopsbench/runtime:dev
docker pull ${NAMESPACE}/runtime:${COMMIT}
export AOB_RUNTIME_IMAGE=${NAMESPACE}/runtime:${COMMIT}

The second line matters. benchmarks/harbor/template/environment/Dockerfile
references the local tag assetopsbench/runtime:dev through its
AOB_RUNTIME_IMAGE build arg, so the pulled image has to carry that tag. If you
publish under a different namespace, change that default instead of asking
every user to retag.
or use ${NAMESPACE}/runtime:latest to follow the newest publish. Each task's
docker-compose.yaml passes AOB_RUNTIME_IMAGE to its Dockerfile as a build arg,
so harbor run builds FROM that image; unset, it falls back to the local tag
assetopsbench/runtime:dev. benchmarks/harbor/run.sh takes it as -r, and pulls
it first.

For the code track, point the overlay at the published image rather than a tar:

Expand Down
Loading
Loading