Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion benchmarks/harbor/CODE-SANDBOX.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,7 +78,7 @@ overlay with the four required `--ak` flags.
-s /path/to/scenarios_data \
-l /path/to/leaderboard \
-n 2 \
-m "watsonx/meta-llama/llama-4-maverick-17b-128e-instruct-fp8"
-m "litellm_proxy/azure/gpt-5.6-sol max"
```

`-n 2` rather than the default 4, for the reason above. Repeat `-m` to sweep
Expand Down
21 changes: 13 additions & 8 deletions benchmarks/harbor/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,13 +63,16 @@ uv run harbor run -p benchmarks/harbor/datasets/assetopsbench-open \
--agent oracle --n-concurrent 2

# 4. Run Stirrup, scenarios in parallel. assetops_harbor is installed as part
# of the project, so no PYTHONPATH is needed.
# of the project, so no PYTHONPATH is needed. Use a litellm_proxy/ or
# tokenrouter/ model: the FMSR server runs generate_failure_modes on the
# agent's model and accepts only those two routers, so a watsonx/ agent runs
# without that tool unless FMSR_MODEL_ID names a router model.
uv run harbor run \
-p benchmarks/harbor/datasets/assetopsbench-open \
--agent assetops_harbor.stirrup:StirrupAgent \
--model watsonx/meta-llama/llama-4-maverick-17b-128e-instruct-fp8 \
--model litellm_proxy/azure/gpt-5.6-sol \
--ak code_enabled=false \
--n-concurrent 16
--n-concurrent 4
```

## Running a full scenario suite (the `benchmarks/run.sh` equivalent)
Expand Down Expand Up @@ -271,16 +274,18 @@ memory (see [Resource limits](#resource-limits)).

For the local code backend instead, drop `code-sandbox.yaml` and
`AOB_CODE_TAR`, and pass only `--ak code_enabled=true --ak code_backend=local`.
Code then runs inside `main`, next to CouchDB and the runtime image's copy of
the repo, which includes the open scenarios' ground truth.
Code then runs inside `main`, as root, next to CouchDB and the model
credentials the agent forwards. The runtime image carries no answer files, but
the verifier later runs in that same container.

A private run loaded its data when no trial's `agent/*.stdout.txt` contains
`Database does not exist`.

## Why the MCP servers need an explicit env

`StirrupAgentRunner._build_mcp_config` sets `env` on every stdio server. It has
to. `mcp.client.stdio` applies `get_default_environment()` when
`StirrupAgentRunner._build_mcp_config` sets `env` on every stdio server, from
`mcp_server_env` in `src/agent/runner.py`, which the plan-execute executor uses
too. It has to. `mcp.client.stdio` applies `get_default_environment()` when
`StdioServerParameters.env` is None, and that inherits only HOME, LOGNAME, PATH,
SHELL, TERM and USER. Without it, no server sees `COUCHDB_URL` and each falls
back to `http://localhost:5984`.
Expand Down Expand Up @@ -340,7 +345,7 @@ to a registry, and only the names in `CREDENTIAL_ENV_VARS` reach the container.
| `--ak temperature=0.2` | `--temperature 0.2` | omitted |
| `--ak reasoning_effort=high` | `--reasoning-effort high` | omitted |

`code_backend` defaults to `local` rather than `stirrup-agent`'s own default of `docker`. The docker backend spawns a sibling container from `STIRRUP_CODE_IMAGE`, and a Harbor task container has no Docker daemon, so that default would fail every run. The adapter rejects `docker` with that explanation unless you pass `allow_docker_backend=true` and wire a socket into the task's compose file. `local` is the right backend under Harbor anyway: the container is already a per-trial sandbox, so the isolation the docker backend buys on a laptop is redundant here.
`code_backend` defaults to `local` rather than `stirrup-agent`'s own default of `docker`. The docker backend spawns a sibling container from `STIRRUP_CODE_IMAGE`, and a plain Harbor task container has no Docker daemon, so that default would fail every run. The adapter rejects `docker` with that explanation unless you pass `allow_docker_backend=true`, which is for runs that add `overlays/code-sandbox.yaml`: that overlay gives each trial its own Docker-in-Docker daemon, and `workspace_dir=/workspace-share` is then required too (see [CODE-SANDBOX.md](CODE-SANDBOX.md)). The two backends are not equivalent. `local` runs agent code in `main`, where it can reach CouchDB by name and read the forwarded credentials. Under the sandbox, code containers get neither; the overlay's header says what that does and does not block.

## Three things the existing CLI forced

Expand Down
6 changes: 3 additions & 3 deletions benchmarks/harbor/adapter/generate_tasks.py
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
"""Generate Harbor task directories from AssetOpsBench scenarios.

One adapter, two datasets. Point it at the open profile and the in-repo
scenario data to produce the public set; point it at the full corpus to produce
the restricted set. The task template is shared, which is what keeps the two
from drifting apart.
scenario data to produce the public set; point it at the private suite to
produce the mini, lite and all sets. The task template is shared, which is what
keeps the two from drifting apart.

python benchmarks/harbor/adapter/generate_tasks.py --overwrite

Expand Down
12 changes: 7 additions & 5 deletions benchmarks/harbor/base-image/Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,9 @@
# the whole environment/ directory, so a fat per-task context defeats the cache
# and rebuilds the world for every scenario.
#
# bash benchmarks/harbor/scripts/build-runtime-image.sh \
# -t assetopsbench/runtime:$(git rev-parse --short HEAD)
# bash benchmarks/harbor/scripts/build-runtime-image.sh
#
# tags it assetopsbench/runtime:dev and assetopsbench/runtime:<commit>.
#
# The script builds from `git archive HEAD`, so `COPY . .` below sees committed
# files only, and passes that commit as AOB_COMMIT. A plain `docker build` of the
Expand Down Expand Up @@ -52,7 +53,7 @@ COPY pyproject.toml uv.lock ./
# it fails at export time on a missing google.protobuf and drops every batch.
RUN uv sync --frozen --group otel --extra tsfm --no-install-project

# Repo, including the shared scenario corpus under
# Repo, including the shared scenario data under
# src/couchdb/scenarios_data/shared (7.7 MB in the open profile). Shared here
# means pulled once, not copied into every task context.
COPY . .
Expand All @@ -64,8 +65,9 @@ RUN uv sync --frozen --group otel --extra tsfm
# preload_models.py --print-repos > benchmarks/harbor/base-image/models.txt
#
# The list rather than the catalog, because the catalog is private and a COPY
# would write it into a layer for good. check_models_list.sh fails CI when the
# two drift.
# would write it into a layer for good. Nothing checks that the two agree:
# rerun the command above whenever the catalog's Hub entries change, or the new
# models are not preloaded.
RUN uv run python benchmarks/harbor/scripts/preload_models.py \
--from-list benchmarks/harbor/base-image/models.txt --download

Expand Down
12 changes: 6 additions & 6 deletions benchmarks/harbor/scripts/publish-images.sh
Original file line number Diff line number Diff line change
Expand Up @@ -2,16 +2,16 @@
# Build and publish the images the Harbor tasks depend on, for both
# architectures, so nobody has to build them locally.
#
# ./benchmarks/harbor/scripts/publish-images.sh assetopsbench v0.1.0
# ./benchmarks/harbor/scripts/publish-images.sh quay.io/assetopsbench v0.1.0
#
# Run from the repository root. Requires `docker login` and a buildx builder
# that can do multi-platform builds:
# Run from the repository root. Requires `docker login` to that registry and a
# buildx builder that can do multi-platform builds:
#
# docker buildx create --name aob --use --bootstrap
#
# Two images:
# <namespace>/runtime the repo, its uv environment and the shared corpus.
# Every task image layers its scenario onto this.
# <namespace>/runtime the repo, its uv environment and the shared scenario
# data. Every task image layers its scenario onto this.
# <namespace>/code the sandbox for the code track (numpy, pandas, scipy).
# Only needed by the code-sandbox overlay.
#
Expand All @@ -23,7 +23,7 @@
# are not published; commit them first.
set -euo pipefail

NAMESPACE="${1:?usage: publish-images.sh <dockerhub-namespace> [tag]}"
NAMESPACE="${1:?usage: publish-images.sh <namespace, e.g. quay.io/assetopsbench> [tag]}"
TAG="${2:-dev}"
PLATFORMS="${PLATFORMS:-linux/amd64,linux/arm64}"

Expand Down
4 changes: 2 additions & 2 deletions src/agent/stirrup_agent/runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -250,8 +250,8 @@ def _build_mcp_config(self):
# mcphub already passes the parent environment through for the
# other runners; this keeps Stirrup consistent with it.
# mcp_server_env also pins FMSR_MODEL_ID so the FMSR server's
# generate_failure_modes uses this run's model rather than its
# standalone watsonx default.
# generate_failure_modes uses this run's model; the server has
# no default of its own.
"env": env,
}
return MCPConfig.model_validate({"mcpServers": servers})
Expand Down
58 changes: 31 additions & 27 deletions src/assetops_harbor/stirrup.py
Original file line number Diff line number Diff line change
@@ -1,36 +1,39 @@
"""Stirrup as a Harbor agent.

Written against AssetOpsBench main (81265cb). Harbor's agent factory imports
any ``module.path:ClassName`` passed to ``--agent`` directly, bypassing its
built-in name enum, so this class needs no upstream registration and lives in
the AssetOpsBench repo:
Harbor's agent factory imports any ``module.path:ClassName`` passed to
``--agent`` directly, bypassing its built-in name enum, so this class needs no
upstream registration and lives in the AssetOpsBench repo. It is installed with
the project, so no PYTHONPATH is needed:

PYTHONPATH=agent harbor run -p datasets/assetopsbench-open \
uv run harbor run -p benchmarks/harbor/datasets/assetopsbench-open \
--agent assetops_harbor.stirrup:StirrupAgent \
--model watsonx/meta-llama/llama-4-maverick-17b-128e-instruct-fp8 \
--ak code_enabled=false \
--n-concurrent 16
--model litellm_proxy/azure/gpt-5.6-sol \
--ak code_enabled=true --ak code_backend=local \
--n-concurrent 4

The agent process runs INSIDE the task container, which is what keeps this
phase small: the six MCP servers stay stdio children of Stirrup exactly as
``mcphub.DEFAULT_SERVERS`` launches them today (``uv run <name>-mcp-server``),
and they inherit this trial's COUCHDB_URL through mcphub's
``{**os.environ, **(env or {})}`` merge. Nothing about the transport changes.
phase small: the six MCP servers stay stdio children of Stirrup
(``uv run <name>-mcp-server``), and they get this trial's COUCHDB_URL through
the explicit ``env`` that ``agent.runner.mcp_server_env`` builds. Nothing about
the transport changes.

Three facts about main shape this file:
Three facts about ``stirrup-agent`` shape this file:

* ``stirrup-agent`` takes the question as a REQUIRED POSITIONAL argument
* It takes the question as a REQUIRED POSITIONAL argument
(``_cli_common.add_common_args``). There is no stdin path, so the question
is uploaded to a file and passed as ``"$(cat ...)"``, which survives
multi-line prose without the argv quoting hazards of inlining it.
* There is no ``--topology`` flag. The arms main actually exposes are
* There is no ``--topology`` flag. The arms it exposes are
``--code-enabled`` / ``--no-code``, ``--code-backend``, ``--max-turns``,
``--temperature`` and ``--reasoning-effort``.
* ``--code-backend`` defaults to ``docker``, which spawns a sibling container
from ``STIRRUP_CODE_IMAGE``. There is no Docker daemon inside a Harbor task
container, so that default would fail every run. ``local`` is the right
backend here: the Harbor container is already a per-trial sandbox, so the
isolation the docker backend buys on a laptop is redundant.
from ``STIRRUP_CODE_IMAGE``. A plain Harbor task container has no Docker
daemon, so this adapter defaults to ``local`` and accepts ``docker`` only
with ``allow_docker_backend=true``, for runs that add
``benchmarks/harbor/overlays/code-sandbox.yaml``. That overlay gives each
trial its own Docker-in-Docker daemon, whose code containers get neither
CouchDB's hostname nor the forwarded credentials. ``local`` runs that code
in ``main``, next to both.
"""

from __future__ import annotations
Expand Down Expand Up @@ -132,10 +135,10 @@ def __init__(
if code_backend == "docker" and not allow_docker_backend:
raise ValueError(
"code_backend='docker' spawns a sibling container from "
"STIRRUP_CODE_IMAGE, and a Harbor task container has no Docker "
"daemon. Use code_backend='local' (the Harbor container is "
"already a per-trial sandbox), or mount a Docker socket into "
"the task's compose file and pass allow_docker_backend=true."
"STIRRUP_CODE_IMAGE, and a plain Harbor task container has no "
"Docker daemon. Use code_backend='local', or add "
"--extra-docker-compose benchmarks/harbor/overlays/code-sandbox.yaml "
"and pass allow_docker_backend=true."
)

# Harbor records every agent kwarg in the trial's config.json, so the
Expand Down Expand Up @@ -540,10 +543,11 @@ def _load_dotenv() -> None:

Runs host-side, in the Harbor process. override=False keeps exported shell
variables ahead of the file, and --ae stays ahead of both because _get_env
checks the agent's extra env first. Only CREDENTIAL_ENV_VARS reach the agent
container, and .dockerignore keeps the file itself out of every image. The
verifier resolves its ${VAR:-} templates after the agent is constructed, so
AOB_JUDGE_MODEL and the judge keys are picked up from .env as well.
checks the agent's extra env first. Only CREDENTIAL_ENV_VARS and
SETTING_ENV_VARS reach the agent container, and .dockerignore keeps the
file itself out of every image. The verifier resolves its ${VAR:-}
templates after the agent is constructed, so AOB_JUDGE_MODEL and the judge
keys are picked up from .env as well.
"""
load_dotenv(find_dotenv(usecwd=True), override=False)

Expand Down
Loading