diff --git a/benchmarks/harbor/CODE-SANDBOX.md b/benchmarks/harbor/CODE-SANDBOX.md index a48c31bea..4e103a3d2 100644 --- a/benchmarks/harbor/CODE-SANDBOX.md +++ b/benchmarks/harbor/CODE-SANDBOX.md @@ -78,7 +78,7 @@ overlay with the four required `--ak` flags. -s /path/to/scenarios_data \ -l /path/to/leaderboard \ -n 2 \ - -m "watsonx/meta-llama/llama-4-maverick-17b-128e-instruct-fp8" + -m "litellm_proxy/azure/gpt-5.6-sol max" ``` `-n 2` rather than the default 4, for the reason above. Repeat `-m` to sweep diff --git a/benchmarks/harbor/README.md b/benchmarks/harbor/README.md index a46a7e1a4..ad2dec26a 100644 --- a/benchmarks/harbor/README.md +++ b/benchmarks/harbor/README.md @@ -63,13 +63,16 @@ uv run harbor run -p benchmarks/harbor/datasets/assetopsbench-open \ --agent oracle --n-concurrent 2 # 4. Run Stirrup, scenarios in parallel. assetops_harbor is installed as part -# of the project, so no PYTHONPATH is needed. +# of the project, so no PYTHONPATH is needed. Use a litellm_proxy/ or +# tokenrouter/ model: the FMSR server runs generate_failure_modes on the +# agent's model and accepts only those two routers, so a watsonx/ agent runs +# without that tool unless FMSR_MODEL_ID names a router model. uv run harbor run \ -p benchmarks/harbor/datasets/assetopsbench-open \ --agent assetops_harbor.stirrup:StirrupAgent \ - --model watsonx/meta-llama/llama-4-maverick-17b-128e-instruct-fp8 \ + --model litellm_proxy/azure/gpt-5.6-sol \ --ak code_enabled=false \ - --n-concurrent 16 + --n-concurrent 4 ``` ## Running a full scenario suite (the `benchmarks/run.sh` equivalent) @@ -271,16 +274,18 @@ memory (see [Resource limits](#resource-limits)). For the local code backend instead, drop `code-sandbox.yaml` and `AOB_CODE_TAR`, and pass only `--ak code_enabled=true --ak code_backend=local`. -Code then runs inside `main`, next to CouchDB and the runtime image's copy of -the repo, which includes the open scenarios' ground truth. +Code then runs inside `main`, as root, next to CouchDB and the model +credentials the agent forwards. The runtime image carries no answer files, but +the verifier later runs in that same container. A private run loaded its data when no trial's `agent/*.stdout.txt` contains `Database does not exist`. ## Why the MCP servers need an explicit env -`StirrupAgentRunner._build_mcp_config` sets `env` on every stdio server. It has -to. `mcp.client.stdio` applies `get_default_environment()` when +`StirrupAgentRunner._build_mcp_config` sets `env` on every stdio server, from +`mcp_server_env` in `src/agent/runner.py`, which the plan-execute executor uses +too. It has to. `mcp.client.stdio` applies `get_default_environment()` when `StdioServerParameters.env` is None, and that inherits only HOME, LOGNAME, PATH, SHELL, TERM and USER. Without it, no server sees `COUCHDB_URL` and each falls back to `http://localhost:5984`. @@ -340,7 +345,7 @@ to a registry, and only the names in `CREDENTIAL_ENV_VARS` reach the container. | `--ak temperature=0.2` | `--temperature 0.2` | omitted | | `--ak reasoning_effort=high` | `--reasoning-effort high` | omitted | -`code_backend` defaults to `local` rather than `stirrup-agent`'s own default of `docker`. The docker backend spawns a sibling container from `STIRRUP_CODE_IMAGE`, and a Harbor task container has no Docker daemon, so that default would fail every run. The adapter rejects `docker` with that explanation unless you pass `allow_docker_backend=true` and wire a socket into the task's compose file. `local` is the right backend under Harbor anyway: the container is already a per-trial sandbox, so the isolation the docker backend buys on a laptop is redundant here. +`code_backend` defaults to `local` rather than `stirrup-agent`'s own default of `docker`. The docker backend spawns a sibling container from `STIRRUP_CODE_IMAGE`, and a plain Harbor task container has no Docker daemon, so that default would fail every run. The adapter rejects `docker` with that explanation unless you pass `allow_docker_backend=true`, which is for runs that add `overlays/code-sandbox.yaml`: that overlay gives each trial its own Docker-in-Docker daemon, and `workspace_dir=/workspace-share` is then required too (see [CODE-SANDBOX.md](CODE-SANDBOX.md)). The two backends are not equivalent. `local` runs agent code in `main`, where it can reach CouchDB by name and read the forwarded credentials. Under the sandbox, code containers get neither; the overlay's header says what that does and does not block. ## Three things the existing CLI forced diff --git a/benchmarks/harbor/adapter/generate_tasks.py b/benchmarks/harbor/adapter/generate_tasks.py index 0ec5ee089..d7408d155 100644 --- a/benchmarks/harbor/adapter/generate_tasks.py +++ b/benchmarks/harbor/adapter/generate_tasks.py @@ -1,9 +1,9 @@ """Generate Harbor task directories from AssetOpsBench scenarios. One adapter, two datasets. Point it at the open profile and the in-repo -scenario data to produce the public set; point it at the full corpus to produce -the restricted set. The task template is shared, which is what keeps the two -from drifting apart. +scenario data to produce the public set; point it at the private suite to +produce the mini, lite and all sets. The task template is shared, which is what +keeps the two from drifting apart. python benchmarks/harbor/adapter/generate_tasks.py --overwrite diff --git a/benchmarks/harbor/base-image/Dockerfile b/benchmarks/harbor/base-image/Dockerfile index 25b450f25..04dc2f301 100644 --- a/benchmarks/harbor/base-image/Dockerfile +++ b/benchmarks/harbor/base-image/Dockerfile @@ -6,8 +6,9 @@ # the whole environment/ directory, so a fat per-task context defeats the cache # and rebuilds the world for every scenario. # -# bash benchmarks/harbor/scripts/build-runtime-image.sh \ -# -t assetopsbench/runtime:$(git rev-parse --short HEAD) +# bash benchmarks/harbor/scripts/build-runtime-image.sh +# +# tags it assetopsbench/runtime:dev and assetopsbench/runtime:. # # The script builds from `git archive HEAD`, so `COPY . .` below sees committed # files only, and passes that commit as AOB_COMMIT. A plain `docker build` of the @@ -52,7 +53,7 @@ COPY pyproject.toml uv.lock ./ # it fails at export time on a missing google.protobuf and drops every batch. RUN uv sync --frozen --group otel --extra tsfm --no-install-project -# Repo, including the shared scenario corpus under +# Repo, including the shared scenario data under # src/couchdb/scenarios_data/shared (7.7 MB in the open profile). Shared here # means pulled once, not copied into every task context. COPY . . @@ -64,8 +65,9 @@ RUN uv sync --frozen --group otel --extra tsfm # preload_models.py --print-repos > benchmarks/harbor/base-image/models.txt # # The list rather than the catalog, because the catalog is private and a COPY -# would write it into a layer for good. check_models_list.sh fails CI when the -# two drift. +# would write it into a layer for good. Nothing checks that the two agree: +# rerun the command above whenever the catalog's Hub entries change, or the new +# models are not preloaded. RUN uv run python benchmarks/harbor/scripts/preload_models.py \ --from-list benchmarks/harbor/base-image/models.txt --download diff --git a/benchmarks/harbor/scripts/publish-images.sh b/benchmarks/harbor/scripts/publish-images.sh index 2f10e945e..609bb073b 100755 --- a/benchmarks/harbor/scripts/publish-images.sh +++ b/benchmarks/harbor/scripts/publish-images.sh @@ -2,16 +2,16 @@ # Build and publish the images the Harbor tasks depend on, for both # architectures, so nobody has to build them locally. # -# ./benchmarks/harbor/scripts/publish-images.sh assetopsbench v0.1.0 +# ./benchmarks/harbor/scripts/publish-images.sh quay.io/assetopsbench v0.1.0 # -# Run from the repository root. Requires `docker login` and a buildx builder -# that can do multi-platform builds: +# Run from the repository root. Requires `docker login` to that registry and a +# buildx builder that can do multi-platform builds: # # docker buildx create --name aob --use --bootstrap # # Two images: -# /runtime the repo, its uv environment and the shared corpus. -# Every task image layers its scenario onto this. +# /runtime the repo, its uv environment and the shared scenario +# data. Every task image layers its scenario onto this. # /code the sandbox for the code track (numpy, pandas, scipy). # Only needed by the code-sandbox overlay. # @@ -23,7 +23,7 @@ # are not published; commit them first. set -euo pipefail -NAMESPACE="${1:?usage: publish-images.sh [tag]}" +NAMESPACE="${1:?usage: publish-images.sh [tag]}" TAG="${2:-dev}" PLATFORMS="${PLATFORMS:-linux/amd64,linux/arm64}" diff --git a/src/agent/stirrup_agent/runner.py b/src/agent/stirrup_agent/runner.py index 63ce3373b..2061eb002 100644 --- a/src/agent/stirrup_agent/runner.py +++ b/src/agent/stirrup_agent/runner.py @@ -250,8 +250,8 @@ def _build_mcp_config(self): # mcphub already passes the parent environment through for the # other runners; this keeps Stirrup consistent with it. # mcp_server_env also pins FMSR_MODEL_ID so the FMSR server's - # generate_failure_modes uses this run's model rather than its - # standalone watsonx default. + # generate_failure_modes uses this run's model; the server has + # no default of its own. "env": env, } return MCPConfig.model_validate({"mcpServers": servers}) diff --git a/src/assetops_harbor/stirrup.py b/src/assetops_harbor/stirrup.py index f5bb3f114..8bee60a1b 100644 --- a/src/assetops_harbor/stirrup.py +++ b/src/assetops_harbor/stirrup.py @@ -1,36 +1,39 @@ """Stirrup as a Harbor agent. -Written against AssetOpsBench main (81265cb). Harbor's agent factory imports -any ``module.path:ClassName`` passed to ``--agent`` directly, bypassing its -built-in name enum, so this class needs no upstream registration and lives in -the AssetOpsBench repo: +Harbor's agent factory imports any ``module.path:ClassName`` passed to +``--agent`` directly, bypassing its built-in name enum, so this class needs no +upstream registration and lives in the AssetOpsBench repo. It is installed with +the project, so no PYTHONPATH is needed: - PYTHONPATH=agent harbor run -p datasets/assetopsbench-open \ + uv run harbor run -p benchmarks/harbor/datasets/assetopsbench-open \ --agent assetops_harbor.stirrup:StirrupAgent \ - --model watsonx/meta-llama/llama-4-maverick-17b-128e-instruct-fp8 \ - --ak code_enabled=false \ - --n-concurrent 16 + --model litellm_proxy/azure/gpt-5.6-sol \ + --ak code_enabled=true --ak code_backend=local \ + --n-concurrent 4 The agent process runs INSIDE the task container, which is what keeps this -phase small: the six MCP servers stay stdio children of Stirrup exactly as -``mcphub.DEFAULT_SERVERS`` launches them today (``uv run -mcp-server``), -and they inherit this trial's COUCHDB_URL through mcphub's -``{**os.environ, **(env or {})}`` merge. Nothing about the transport changes. +phase small: the six MCP servers stay stdio children of Stirrup +(``uv run -mcp-server``), and they get this trial's COUCHDB_URL through +the explicit ``env`` that ``agent.runner.mcp_server_env`` builds. Nothing about +the transport changes. -Three facts about main shape this file: +Three facts about ``stirrup-agent`` shape this file: -* ``stirrup-agent`` takes the question as a REQUIRED POSITIONAL argument +* It takes the question as a REQUIRED POSITIONAL argument (``_cli_common.add_common_args``). There is no stdin path, so the question is uploaded to a file and passed as ``"$(cat ...)"``, which survives multi-line prose without the argv quoting hazards of inlining it. -* There is no ``--topology`` flag. The arms main actually exposes are +* There is no ``--topology`` flag. The arms it exposes are ``--code-enabled`` / ``--no-code``, ``--code-backend``, ``--max-turns``, ``--temperature`` and ``--reasoning-effort``. * ``--code-backend`` defaults to ``docker``, which spawns a sibling container - from ``STIRRUP_CODE_IMAGE``. There is no Docker daemon inside a Harbor task - container, so that default would fail every run. ``local`` is the right - backend here: the Harbor container is already a per-trial sandbox, so the - isolation the docker backend buys on a laptop is redundant. + from ``STIRRUP_CODE_IMAGE``. A plain Harbor task container has no Docker + daemon, so this adapter defaults to ``local`` and accepts ``docker`` only + with ``allow_docker_backend=true``, for runs that add + ``benchmarks/harbor/overlays/code-sandbox.yaml``. That overlay gives each + trial its own Docker-in-Docker daemon, whose code containers get neither + CouchDB's hostname nor the forwarded credentials. ``local`` runs that code + in ``main``, next to both. """ from __future__ import annotations @@ -132,10 +135,10 @@ def __init__( if code_backend == "docker" and not allow_docker_backend: raise ValueError( "code_backend='docker' spawns a sibling container from " - "STIRRUP_CODE_IMAGE, and a Harbor task container has no Docker " - "daemon. Use code_backend='local' (the Harbor container is " - "already a per-trial sandbox), or mount a Docker socket into " - "the task's compose file and pass allow_docker_backend=true." + "STIRRUP_CODE_IMAGE, and a plain Harbor task container has no " + "Docker daemon. Use code_backend='local', or add " + "--extra-docker-compose benchmarks/harbor/overlays/code-sandbox.yaml " + "and pass allow_docker_backend=true." ) # Harbor records every agent kwarg in the trial's config.json, so the @@ -540,10 +543,11 @@ def _load_dotenv() -> None: Runs host-side, in the Harbor process. override=False keeps exported shell variables ahead of the file, and --ae stays ahead of both because _get_env - checks the agent's extra env first. Only CREDENTIAL_ENV_VARS reach the agent - container, and .dockerignore keeps the file itself out of every image. The - verifier resolves its ${VAR:-} templates after the agent is constructed, so - AOB_JUDGE_MODEL and the judge keys are picked up from .env as well. + checks the agent's extra env first. Only CREDENTIAL_ENV_VARS and + SETTING_ENV_VARS reach the agent container, and .dockerignore keeps the + file itself out of every image. The verifier resolves its ${VAR:-} + templates after the agent is constructed, so AOB_JUDGE_MODEL and the judge + keys are picked up from .env as well. """ load_dotenv(find_dotenv(usecwd=True), override=False)