feat(harbor): run AssetOpsBench scenarios as Harbor tasks - #558
Conversation
A missing or unreachable database was reported the same way as a wrong key (e.g. 'unknown asset_id', 'work order not found'). On not-found/failure paths the iot, fmsr, wo and vibration servers now report that the database does not exist in this environment and that retrying with other arguments won't help. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
…heck Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
fix: distinct error message when a CouchDB database is missing
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
ruff EXE001: the file has a shebang but shipped as mode 100644. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7W9hWwuZU8Mkwo3ZCrR7m Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
What the code changes doPR #558 does two things. It adds an opt-in way to run AssetOpsBench scenarios in parallel using Harbor, where each trial gets its own CouchDB. It also brings along a batch of fixes to the servers and changes to the TSFM model catalog. It's large: 57 files, +7,004 / −172, much of it 1. Harbor integration (the main feature)The reason for it: the current runner resets one shared CouchDB before every scenario, so running scenarios at once would wipe a live scenario's data. The agent would then get empty results and still produce a plausible-looking answer that gets scored. Harbor starts a separate Compose project per trial, so each scenario gets its own database.
2. Clearer errors when a CouchDB database is missing (from #554)The iot, fmsr, wo and vibration servers now tell apart a missing database and a wrong key. Before, both produced errors like "unknown asset_id". Now a missing database returns "the data source does not exist … do not retry with other arguments", without naming the database. Tests cover this. 3. Stirrup runner passes the environment to MCP servers
4. TSFM model catalog and weights
5. Dependencies (
|
- Layout lists the real paths (src/assetops_harbor, template/, overlays/) and marks datasets/ as generated. - Run steps work from the repo root: correct Dockerfile path, adapter defaults instead of a --template pointing at a generated task, no PYTHONPATH=agent, no 'push'. - Drop the pin to main@81265cb; the branch targets aafeedback_changes. - Replace 'Still unverified' with the Docker-host results reported in #558. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
generate_tasks.py gains three options for corpora that do not ship in the repo, such as AssetOpsBenchScenarioGeneration/scenarios_data: - --runtime-image: the image each task builds FROM. - --data-dir: where the corpus lives inside that image. The per-task layer copies the scenario there and the healthcheck runs init_data.py with SCENARIOS_DATA_DIR pointed at it; the agent keeps the repo copy, as scenario_suite_runner does. - --skip-missing: warn and skip profile scenarios absent from the corpus (all.yaml lists wosr-62, which the corpus lacks). corpus-image/Dockerfile bakes the corpus (2.1 GB, mostly shared/iot) into one layer over the runtime image, so tasks share it instead of each carrying it in their build context. Also stop copying scenario 1's description and the wosr keyword into every generated task. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Runs a scenario corpus through Harbor with the same profile, the same
stirrup-agent and the same Docker code sandbox as benchmarks/run.sh, but
with scenarios in parallel, each trial on its own CouchDB.
- Builds assetopsbench/runtime:corpus from -s, and the code sandbox tar
for the per-trial Docker-in-Docker daemon.
- Regenerates the dataset from scratch (the generator never removes stale
task folders), skipping profile entries the corpus lacks.
- One Harbor job per model under <leaderboard>/harbor-jobs; re-running
resumes it, the equivalent of --skip-existing.
- Credentials reach the Harbor process through uv run --env-file only.
Uses [[ -z "${arr[*]+set}" ]] and ${arr[@]+...} so empty arrays work
under set -u in macOS's bash 3.2.
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
corpus-image/ -> suite-image/, assetopsbench/runtime:corpus -> assetopsbench/runtime:suite, /opt/corpus/scenarios_data -> /opt/suite/scenarios_data, and the generated dataset assetopsbench-corpus -> assetopsbench-suite, with matching wording in run.sh, the generator's help and the README. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Completes 533bc5a, which only moved corpus-image/ to suite-image/: assetopsbench/runtime:corpus -> assetopsbench/runtime:suite, /opt/corpus/scenarios_data -> /opt/suite/scenarios_data, the dataset assetopsbench-corpus -> assetopsbench-suite, and the wording in run.sh, the generator's help and the README. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
base-image/Dockerfile does `COPY . .`, so building from the working tree baked in whatever sat there. An untracked reports/results_table.csv with a ground_truth column for 56 private scenarios (8 of mini, 41 of lite) reached assetopsbench/runtime:dev that way, readable by any agent running code in `main`, and publish-images.sh would have pushed it. scripts/build-runtime-image.sh now builds from a one-commit clone of HEAD, so only committed files enter the context. The clone keeps a real .git, which StirrupAgent.get_version_command needs for `git rev-parse`; a git worktree's .git pointer would not resolve inside the container. publish-images.sh, run.sh and the docs use the script. .dockerignore stays as the second guard, and the only one for a plain `docker build .`: env files at any depth (a bare `.env` matched only the context root), the open scenarios' answer files, which the verifier never reads from the image, and the local folders reports/, logs/, src/tmp/, .claude/ and the kdd_tutorial work. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Review of #587: the one-commit .git the clean clone kept still held every committed blob, so `git show HEAD:.../groundtruth.txt` returned the open scenarios' answers that .dockerignore had dropped from the tree. Its reflog also recorded the builder's name and email, and its index and reflog changed on every clone, so `COPY . .` never hit the cache and each build re-downloaded the models. build-runtime-image.sh now builds from `git archive HEAD` and passes the commit as the AOB_COMMIT build arg. .dockerignore drops .git. A `commit` stage in base-image/Dockerfile writes /opt/aob/.aob-commit, which StirrupAgent.get_version_command now reads, and fails any build without AOB_COMMIT, so a plain `docker build .` of the working tree stops instead of baking in untracked files. The ARG lives only in that stage, so a new commit invalidates the final layer and nothing else. The script also keeps its default tag and --load when given other flags (e.g. --no-cache), warns about untracked files it leaves out, and checks for docker buildx up front. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Review of #586: - generate_tasks.py copies into each task only manifest.json and the files it names inside the scenario folder. Skipping the verifier's input names let any other answer file through (a .bak, a new scorer input, GroundTruth.txt on a case-insensitive disk). - overlays/private-data.yaml uses the long bind syntax with create_host_path: false. A mistyped or relative AOB_PRIVATE_DIR used to make Docker create an empty shared/ and load empty collections with the healthcheck passing; it now fails the trial at start with "bind source path does not exist". A path with ':' also works now. - run.sh reports a failed resume instead of hiding it behind `|| true`. Harbor refuses to resume a job whose tasks or overlays no longer match its lock.json, which is every job from the old suite-image run.sh. - The overlay no longer claims suite edits need no rebuild: a changed manifest.json needs the tasks regenerated. The README asks for an absolute AOB_PRIVATE_DIR. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
fix(harbor): run.sh mounts only the private suite's shared/
fix(harbor): build the runtime image from a clean clone of HEAD
Every task image builds FROM the runtime image, and the only way to choose it was the Dockerfile's ARG default, so using a published image meant retagging it as the local assetopsbench/runtime:dev. The template's docker-compose.yaml now passes AOB_RUNTIME_IMAGE to that ARG as a build arg. Harbor runs `docker compose build` with the shell's environment and every compose file, so `AOB_RUNTIME_IMAGE=<image> harbor run ...` builds FROM that image, with no task regeneration; unset, it keeps the local tag. run.sh takes it as -r (or AOB_RUNTIME_IMAGE). A registry reference is pulled first, because a local copy satisfies FROM and the build would otherwise run on a stale one; a bare name must already exist locally, as before. It exports the variable so `harbor jobs resume` sees it too, and prints the image id it uses. The image changes only the base. Each task still adds its own manifest.json, and shared/ still comes from overlays/private-data.yaml. The template change alters every generated task, so Harbor will not resume a job started before it; run.sh already says so and suggests moving it aside. Verified with Harbor's oracle agent, AOB_RUNTIME_IMAGE set to a labelled copy of runtime:dev: the 3 open tasks, and tsfm-1019 and fmsr-913 from the private suite with the private-data overlay. All 5 trials scored 1.000 with no errors. A nonexistent AOB_RUNTIME_IMAGE failed the build pulling exactly that name. The tsfm-1019 task image built on the labelled base carries the label. With the overlay's mount it sees scenario_1019/manifest.json, and a read-only shared/ where all five manifest paths resolve. run.sh's image step pulls a quay.io reference and rejects a missing local tag. 22 assetops_harbor tests pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
build-runtime-image.sh tagged only assetopsbench/runtime:dev, which every build overwrites, so a run could not name the build it used after a rebuild. With no -t given it now also tags assetopsbench/runtime:<short commit>. That tag does not move, so `run.sh -r assetopsbench/runtime:<commit>` (or AOB_RUNTIME_IMAGE) pins a run to one build while :dev stays the default the tasks build FROM. publish-images.sh adds <namespace>/runtime:<commit> beside :TAG and :latest. The commit is HEAD's, which is exactly what the image holds, because the build comes from `git archive HEAD` whatever the working tree contains. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Review fixes for the runtime-image selection. run.sh: - Resolves the image as -r, then the shell's AOB_RUNTIME_IMAGE, then ENV_FILE's, then the local default. It used to take only the shell's value and export a default over it, so an AOB_RUNTIME_IMAGE in .env was ignored; `uv run --env-file` never overrides a variable the shell already has. - Pulls whenever the image is not local, or its local copy came from that same repository. The old rule pulled only references naming a registry host, so a Docker Hub image such as publish-images.sh's own <ns>/runtime:latest was never refreshed. A local build (no RepoDigest for its name) is never pulled over. - Pins the resolved image for the whole run under a tag private to the process (aob-runtime-pin:<id>-<pid>, removed on exit) and exports that. Every trial resolves FROM when it builds, so exporting the movable :dev let a rebuild or another run's pull switch later trials' base mid-run. FROM cannot name a bare image id, which is why this uses a tag. - Records the image beside each job (<job>.runtime-image) and refuses to resume a job on a different one. Harbor's resume lock hashes the task files, not the base they build FROM, so a resume with another -r mixed two images in one job. - Exits non-zero when a model's job could not start or resume. - Runs the image step after the cheap input checks, and strips a sha256: prefix only if present when printing the id. build-runtime-image.sh: treats -tNAME and -t=NAME as naming a tag, so they no longer also get the default :dev and :<commit> tags. Its comment no longer claims the commit tag never moves (rebuilding the same commit moves it), and it says how to remove old commit tags, which now keep their images alive. publish-images.sh: the NOTE points users at the commit tag it just pushed, with :latest as the moving option. Docs: QUICKSTART says tasks generated before the build arg existed ignore the variable and need regenerating once, that the variable must be set per shell or in .env, and its pull-error entry now covers an unset AOB_RUNTIME_IMAGE. README describes the resolution order, pin, per-job record and exit status. Verified end to end with run.sh, a one-scenario profile and a fake model id, so trials fail fast without calling a provider: - The first run printed the resolved id. The pin tag existed during the run and was gone after it, the job's .runtime-image record was written, and the trial built, ran and was scored (reward 0). - A second run on the same image resumed the job, dropped the crashed trial and reran it, and exited 0. - A third run with -r set to a different image refused to resume, named the job's original image, and exited 1. - Helper checks: image_repo strips tags and digests; runtime:dev and :3bb4f12 count as local builds; quay.io/... and couchdb:3.5 count as registry images. -r and the shell beat .env, which beats the default. - `uv run pytest src/ -k "not integration"`: 714 passed, 19 failed. The base branch has the same 19 failures, in the scorer and trace-exporter tests and in the iot invalid-site tests, which need a CouchDB or no .env. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
feat(harbor): pick the runtime image per run; tag builds by commit
Text the later Harbor changes left behind, found reviewing #558. No behaviour changes except the wording of one error message. - stirrup.py module docstring: drop the first-draft run line (PYTHONPATH=agent, a watsonx model, --n-concurrent 16) and the claim that MCP servers get the environment through mcphub; the Stirrup runner passes it through mcp_server_env. The docker backend is no longer "redundant": the code-sandbox overlay provides a daemon and keeps code off CouchDB's hostname and the credentials, which local does not. - StirrupAgent's docker-backend error pointed at mounting a Docker socket; it now names the code-sandbox overlay. It keeps "no Docker daemon", which a test matches. - _load_dotenv: SETTING_ENV_VARS (FMSR_MODEL_ID) reach the container too. - stirrup_agent/runner.py: FMSR has no "standalone watsonx default" since #583. - README step 4 and CODE-SANDBOX.md's run.sh example used a watsonx model, which FMSR now rejects, so those runs had no generate_failure_modes. They use litellm_proxy/azure/gpt-5.6-sol, and step 4 says why. - README: the local backend no longer sits next to the open scenarios' ground truth (#587 removed it from the image); say what it does sit next to. The explicit-env section names mcp_server_env and plan-execute. - base-image/Dockerfile: check_models_list.sh does not exist, so say nothing checks models.txt against the catalog; the build line matches the script's default tags. - generate_tasks.py, publish-images.sh: "corpus" is now "suite"/"scenario data"; publish-images.sh's example namespace is quay.io/assetopsbench. The task template's comments are stale in the same way but are left alone: editing them changes every task's content hash, and Harbor then refuses to resume jobs started on the current tasks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
docs(harbor): correct stale docs, docstrings and comments
Four run.sh problems that hold even with a stable runtime image. - A job was named after the model alone, and a resume runs the job's saved config. So `-m "X high" -m "X max"` resumed the high job for max and ran no max trials, and another -p in the same leaderboard directory hit Harbor's lock and could not resume. Jobs are now stirrup_agent__<profile>__<model>[__<effort>]. - Tasks were regenerated on every run, and Harbor refuses to resume a job whose tasks differ from its lock. Any change to the template (comments included), the suite's scenario files or the generator made every job unresumable. Each new job now gets its own copy of the tasks, generated once when it starts and reused on resume. The copies hold the answers, so they go to the repo's gitignored benchmarks/harbor/datasets/jobs/, keyed by the job's path, not into the leaderboard directory. - One shared task folder meant a second run.sh deleted tasks that a running job was still reading. Per-job copies remove the shared folder. - `harbor run ... || true` hid the only failures harbor run reports: it exits 0 when trials fail and non-zero when the job itself cannot run (a rejected config, or StirrupAgent's credential check aborting the job as its first trial starts). That, and a model skipped for an unreachable router, now set a non-zero exit. A resume whose task copy is gone says so. Jobs started under the old names are not resumed; a rerun starts new jobs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The rest of the run.sh problems found reviewing #558. - Resume reran only NonZeroAgentExitCodeError. Harbor matches exact class names, so its subclasses (rate limits, network, auth), Ctrl-C, and environment and verifier failures stayed failed. run.sh now lists every type that is not the model's own work; timeouts, context and output limits and safety refusals stay as results. test_run_sh.py checks the list against the installed Harbor and fails when a new agent error subclass is unplaced. - The code sandbox tar was built once and reused forever. It is now built each run (cached), saved once per image id under AOB_CODE_TAR_DIR (default ~/.cache/assetopsbench), and recorded per job; a resume loads the job's own tar, so a job never switches code image and a rebuild never rewrites a tar in use. - The router check only probed the base URL. It now rejects a key the router refuses (401/403 from /models), also checks FMSR_MODEL_ID's router, and skips a model with no router prefix unless FMSR_MODEL_ID names one, since FMSR would reject it and fmsr scenarios would run without generate_failure_modes. - ENV_FILE was not the only credential source: StirrupAgent also loaded the repo's .env and filled its gaps. StirrupAgent now reads AOB_ENV_FILE instead of searching when it is set, and run.sh sets it. - Two run.sh processes could work on one job. A per-job lock (mkdir, with the owner's PID so a dead owner's lock is taken over) makes the second skip it. - Relative -s, -l, -p, ENV_FILE and AOB_CODE_TAR_DIR resolved against the repo root. They now resolve against the caller's directory; the defaults are still the repo's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Found reviewing the per-job task copies. A resume reuses the job's own manifests but mounts shared/ from the current -s, so a resume with another suite would pair one suite's manifests with another suite's data. Before the copies, Harbor's task lock refused that resume; now run.sh records the suite beside the job (<job>.suite) and refuses it, as it does for another runtime image. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
fix(harbor): harden run.sh jobs, retries, checks and inputs
test_the_sdk_default_would_drop_couchdb_url set COUCHDB_URL in os.environ and popped it afterwards, deleting any value the developer already had, so later tests in the session ran without it. It now uses monkeypatch.setenv, which restores the original. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
…ments - pyproject.toml listed python-dotenv and granite-tsfm twice in dependencies, and numba's line in the dev group had trailing whitespace. uv.lock is unchanged (`uv lock --check` passes). - .gitignore listed benchmarks/harbor/datasets/ twice, and its comment named harbor/adapter/generate_tasks.py without the benchmarks/ prefix. - template/task.toml said mcphub carries COUCHDB_URL to the MCP servers "with no code change"; the runners pass it through mcp_server_env, because the MCP SDK's default environment drops it. - template/environment/Dockerfile described the per-task COPY as how "the restricted corpus" layers its scenarios; it now says what it copies (the manifest and the files it names, never answers) and that an external suite lands in /opt/suite/scenarios_data. Template edits change generated tasks' hashes. run.sh jobs are unaffected, since each job keeps its own copy of its tasks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The fine-tuned energy model sat in artifacts/output/tuned_models/, the root artifacts/README.md reserves for checkpoints an agent writes during a trial and says is never committed and ignored by git and Docker. But the model catalog serves it as a shipped model, so the ignore rules the README describes were never added: adding them would have dropped the model and broken its card. - Move the checkpoint (config.json, meta.json, model.safetensors) to artifacts/tsfm_models/ttm_energy_168_24 and repoint its card's three path fields (apply_catalog_fixes.py --move-energy --write). - .gitignore and .dockerignore now exclude artifacts/output/, as the README says. - generate_model_catalog.py no longer scans artifacts/output/tuned_models: agent output must not become catalog cards. - apply_catalog_fixes.py: the docstring says the repo has moved and the flag is for catalogs that have not, such as a private suite's; the reminder prints only when a catalog still points at the old path. - artifacts/README.md lists what tsfm_models/ holds, corrects the checkpoint sizes, and says save_to takes any path rather than implying agents write to output/ by themselves. test_catalog_checkpoints.py (12 passed) loads, fits and forecasts the model at the new path, and preload_models.py --check resolves it. A catalog outside the repo that still names the old path fails for this model once the runtime image is rebuilt. Update it with apply_catalog_fixes.py --catalog <path> --move-energy --write Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Cut the Harbor docs and code comments to what a reader needs now, and remove content that no longer holds. - README: drop the dated "verified" run logs, the smoke-test commands that repeated other sections, the CLI-history notes, the `git rm --cached jobs` step and the code-track section duplicated in CODE-SANDBOX.md; summarise run.sh as an options table and short steps. - QUICKSTART: drop the code-track section (now a link), the troubleshooting entries for fixes already in the code (MCP env, otel group) and a registry image name that is not published. - CODE-SANDBOX: replace the Harbor environment inventory with one line and drop the claim that main holds groundtruth.txt, which the image excludes. - Overlays, template, Dockerfile, run.sh, image scripts: shorten comments. code-sandbox.yaml said the registry route sets "both variables" but showed one. - stirrup.py, generate_tasks.py, metric.py, catalog scripts: shorten docstrings; drop "on main" / "yet" wording, the reference to a design doc that is not in the repo, and history about past export failures. - Fix README paths in pyproject.toml, assetops_harbor/__init__.py and the generated dataset.toml header (harbor/ -> benchmarks/harbor/). - stirrup_agent/runner.py: the MCP env comment no longer says mcphub covers the other runners; they all use mcp_server_env. No behaviour change. Generated tasks still validate against Harbor's TaskConfig; the unit suite shows the same 19 failures as before the change (evaluation, observability, iot), all outside these files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
code-image-loader echoes $STIRRUP_CODE_IMAGE when no AOB_CODE_TAR is given, but only `main` had that variable, so the message read "the daemon will pull instead". Give the loader the same STIRRUP_CODE_IMAGE as `main`. Checked by rendering the overlay with docker compose config and running the loader's no-tar branch: it now prints assetops-code:dev, or AOB_CODE_IMAGE when set. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
`uv sync` installs the dev group by default, so the runtime image carried pytest and the Jupyter/IPython stack, which nothing in a trial imports. They were 63 of the image's 291 downloaded packages; a timed-out download of one of them (debugpy, via ipykernel) failed a build. Both syncs now pass --no-dev. numba, pyod, anyio and the OpenTelemetry packages stay: other dependencies require them. UV_NO_SYNC=1 is set once the environment is final. Every process in a trial starts through `uv run`, which otherwise syncs the default groups first and would reinstall the dev group in every trial; it would also do so in the preload step of this build. Checked with uv 0.12.21, the version the image installs: after `sync --no-dev`, a plain `uv run` installs the dev group and `UV_NO_SYNC=1 uv run` does not. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The Hub model download (3.9 GB, about 6 minutes) ran after `COPY . .`, so every commit downloaded every model again, and each commit tag kept its own 4 GB copy of the weights. - Download the models in their own stage, which depends only on models.txt and preload_models.py (its --from-list path needs nothing but huggingface_hub, pinned to uv.lock's 1.33.0). The main stage copies /opt/hf below the dependency layer, since uv.lock changes more often than models.txt. - Give both `uv sync` calls a BuildKit cache mount for the uv cache. The cache used to stay in the image (2.2 GB, beside a 2.0 GB venv); now a lock change reuses downloaded wheels. Measured with the docker driver: a source edit rebuilds in 2 s (was a full model download), a pyproject.toml edit in 24 s, and a cold build in about 7 minutes. The image shrinks from 8.75 GB to 6.53 GB, and a new commit adds only its ~60 MB of repo layers. COPY --link was tried for the weights and did not avoid the copy on this driver, so it is not used. Checked on the built image: no uv cache, no dev group, all 20 models resolve offline, imports work, and Harbor's oracle scores 1.000 on the 3 open tasks with 0 exceptions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Replace "Running it", "Running a full scenario suite" and "Running a private profile by hand" with three sections: - Running a private profile end to end: build the image, pin the run to its commit tag, run.sh on mini, then harbor view and metric.py, followed by run.sh's options, steps, resume and locking (unchanged). - Running a public profile by hand: the open profile with plain Harbor commands, including saving the code sandbox tar, the oracle check and the Docker sandbox run; tools-only, results and resume as notes. - Running a private profile by hand: the same with --scenario-root, its own --output-dir and the private-data overlay, then the generator flags and "How the private data gets in" (unchanged). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Harbor names a local dataset after its folder and builds agent__model__dataset keys from it. Its CLI splits those keys on "__" to print the results table, and Job splits them to pick the dataset's metrics. run.sh named each job's tasks <job_name>-<checksum>, and job names contain "__", so a run.sh job failed once every trial had finished: too many values to unpack (expected 2) Harbor could not run <job>; see the error above. result.json and the trials were written, but with no dataset metrics, and run.sh reported the job as failed. Replace "__" with "--" in the folder name. A new test rebuilds the name from run.sh's own line and checks it against Harbor's key format. Reproduced with the oracle on the open tasks: a folder with "__" exits 1 with that error and empty metrics; with "--" it exits 0 and records them. A job started before this change keeps its old task folder, which run.sh no longer finds, so it refuses to resume it; move that job aside to rerun it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Description
This PR does two things:
Harbor gives each trial its own Compose project and names every container, network and volume after it, so each scenario gets its own CouchDB. That makes parallel scenario runs safe by design, rather than safe only if they're carefully sequenced.
Why this matters. The current path can't be run in parallel safely:
scenario_suite_runner.pyresets one shared CouchDB before each scenario, andloader._ensure_dbdrops each database before loading it.iot,workorderandvibrationdatabases partway through its run.Existing behaviour doesn't change.
scenario_suite_runner.pyis untouched, and Harbor is an opt-in extra (uv sync --extra harbor).Size: 73 files, +8,132 / −249. Most of that is
uv.lock, about 41 MB of TTM weights underartifacts/, and the catalog scripts underbenchmarks/harbor/scripts/. The PR targetsaafeedback_changes, notmain; see the reviewer notes.What the code changes do
1. Harbor integration (the main feature),
benchmarks/harbor/template/):task.tomlsets the trial's environment (COUCHDB_URL=http://couchdb:5984,SCENARIOS_DATA_DIR,AGENT_TRAJECTORY_DIR, …). Its healthcheck runsinit_data.py <id>, so the scenario's data is loaded before Harbor starts the agent.environment/docker-compose.yamladds a CouchDB sidecar that publishes no host port. It also passesAOB_RUNTIME_IMAGEto the task's Dockerfile, so the base image is chosen per run, with no need to regenerate tasks (feat(harbor): pick the runtime image per run; tag builds by commit #588).tests/test.shruns the existinguv run evaluate, andtests/to_reward.pyconverts its report into Harbor's{reward, passed}.adapter/generate_tasks.py) turns each scenario in a profile (open.yamlby default) into a Harbor task: the question, a thin Docker layer, the ground truth for the verifier, and an oracle solution that writes the known answer.manifest.jsonand the files that manifest names.groundtruth.txt,rubric.json, …) go only totests/andsolution/, which Harbor uploads after the agent phase (fix(harbor): run.sh mounts only the private suite's shared/ #586).base-image/):uv sync --no-dev --group otel --extra tsfm, and the 20 Hub models listed inmodels.txt(Chronos, Moirai, TimesFM, Toto, TSPulse, TTM) preloaded.UV_NO_SYNC=1stops each trial'suv runfrom reinstalling the dev group.102646e). The models download in a stage of their own that depends only onmodels.txtand the preload script, and the uv cache lives in a BuildKit cache mount, out of the image. A source edit rebuilds in about 2 s, apyproject.toml/uv.lockchange in about 25 s, and a cold build takes about 7 minutes. The image is 6.5 GB, and each new commit adds only its ~60 MB of repo layers.scripts/build-runtime-image.shbuilds it fromgit archive HEAD, so untracked files and.gitnever reach it. A plaindocker buildfails fast and points to the script (fix(harbor): build the runtime image from a clean clone of HEAD #587). Each build is taggedassetopsbench/runtime:devandassetopsbench/runtime:<commit>(feat(harbor): pick the runtime image per run; tag builds by commit #588).scripts/publish-images.shbuilds for both architectures and pushes:TAG,:<commit>and:latest.src/assetops_harbor/stirrup.py):stirrup-agentinside the task container, with Harbor's trial name as--run-id;--akkwargs onto Stirrup's settings (code_enabled,code_backend,max_turns,temperature,reasoning_effort,workspace_dir);CREDENTIAL_ENV_VARS, plusFMSR_MODEL_ID, into the agent phase only;.envfrom the working directory before checking credentials, and exported variables and--aetake precedence over it (feat(harbor): load .env for StirrupAgent credentials #575). WhenAOB_ENV_FILEis set, it reads that file instead, which is howrun.shmakesENV_FILEthe only source (fix(harbor): harden run.sh jobs, retries, checks and inputs #591);FMSR_MODEL_IDare missing (fix(fmsr): take the FMSR model from FMSR_MODEL_ID, with no built-in default #583);/opt/aob/.aob-commitas the agent version (fix(harbor): build the runtime image from a clean clone of HEAD #587).overlays/code-sandbox.yaml): an optional Docker-in-Docker sidecar, so code the agent writes can't see CouchDB's hostname or credentials.CODE-SANDBOX.mdcovers it (Doc for code execution #581).overlays/private-data.yaml), for the mini, lite and all profiles (feat(harbor): run mini/lite/all by mounting the private suite #577, fix(harbor): run.sh mounts only the private suite's shared/ #586). It mounts only the private suite'sshared/, read-only. A wrong or relativeAOB_PRIVATE_DIRfails the trial at start, instead of silently loading empty data.run.sh, the Harbor counterpart ofbenchmarks/run.sh.-snames the suite'sscenarios_data,-pthe profile and-rthe runtime image.<leaderboard>/harbor-jobs/stirrup_agent__<profile>__<model>[__<effort>]. Re-running it resumes the job.benchmarks/harbor/datasets/jobs/. Later changes to the template, the suite or the generator no longer block a resume.:devthat moves mid-run can't switch later trials to a different base.~/.cache/assetopsbench.FMSR_MODEL_ID's must answer and must accept their key. A model without a router prefix also needsFMSR_MODEL_ID.test_run_sh.pykeeps that list in step with Harbor.run.shprocesses from working on one job.ENV_FILEis the only credential source. Relative paths resolve against your current directory.metric.py: per-category averages over a finished job.README.md(how it works, with "Running a private profile end to end", "Running a public profile by hand" and "Running a private profile by hand"),QUICKSTART.md(the steps) andCODE-SANDBOX.md..dockerignoreexcludes.envfiles at any depth, the open scenarios' answer files, local-only folders (reports/,logs/,src/tmp/,.claude/, the kdd_tutorial work),artifacts/output/and.git(fix(harbor): build the runtime image from a clean clone of HEAD #587).2. Clearer errors when a CouchDB database is missing (#554)
The
iot,fmsr,woandvibrationservers now tell a missing database apart from a wrong key. Before, both produced errors like "unknown asset_id". Now a missing database returns "the data source does not exist in this environment; ... do not retry", without naming the database. Tests cover each server.3. Agent runners pass the environment to MCP servers
Without an explicit
env, the MCP SDK's stdio client passes on onlyHOME,PATHand a few other variables, so the servers never sawCOUCHDB_URLand silently fell back tolocalhost:5984. The newmcp_server_envinsrc/agent/runner.pyforwards the full environment, and the Stirrup, plan-execute, deep-agent and openai-agent runners now use it. This fixes a real bug and applies outside Harbor too.test_stirrup_mcp_env.pycovers the Stirrup case.4. FMSR model (#583)
generate_failure_modestakes its model only fromFMSR_MODEL_ID; there is no built-in default.--model-idunless it is set explicitly (fmsr_env_overridesinsrc/agent/runner.py).litellm_proxy/andtokenrouter/prefixes are accepted; watsonx support is dropped.INSTRUCTIONS.mdanddocs/running_benchmark.mddocument the new behaviour.5. TSFM model catalog and weights
shared/tsfm/model_catalog.jsondropsttm_96_28and now lists six models: four local TTM checkpoints, the fine-tuned energy modelttm_energy_168_24, andttm_1024_192_hub.artifacts/tsfm_models/as plain git files, not Git LFS (about 41 MB).artifacts/output/, for checkpoints an agent fine-tunes during a trial, is now ignored by git and Docker.benchmarks/harbor/scripts/generate, audit and fix the catalog, preload Hub models, fetch TTM checkpoints and smoke-test forecasts and tasks.src/servers/tsfm/tests/test_catalog_checkpoints.pychecks that every local checkpoint in the catalog resolves, fits and forecasts.6. Dependencies (
pyproject.toml)tsfmmoves from a dependency group to an optional extra and grows a lot:transformers[torch]>=5.3, gluonts, lightning, toto-models and others.harboroptional extra is added, andsrc/assetops_harboris added to the packages.sktimegoes up to>=1.2.0, andnumbaandpyodare added to thedevgroup.Type of Change
Industry Relevance
Third-party evaluators need concurrency and per-trial isolation before they can publish AssetOpsBench numbers, and this PR builds both into the runner. Each result also records the runtime image's commit, and a run can be pinned to one build with
runtime:<commit>.Related Issues
Running the benchmark
Run every command from the repo root. You need Docker running (at least 4 GB of memory) and model credentials in
.env(LITELLM_*and/orTOKENROUTER_*). Use alitellm_proxy/ortokenrouter/model, or setFMSR_MODEL_IDto one: FMSR'sgenerate_failure_modesaccepts only those two routers.benchmarks/harbor/QUICKSTART.mdhas the step-by-step version and common problems.End to end (as of
ba120aa)-sfolder:run.shwires the private suite from it. It mounts the suite'sshared/read-only into each trial and generates each job's tasks from it.-mto run the 8 default models, and-pto run theallprofile. Each trial runs a privilegeddindsidecar, so keep-naround 4 on a laptop-sized Docker VM.ttm_energy_168_24must point atartifacts/tsfm_models/before you run on a new image:uv run python benchmarks/harbor/scripts/apply_catalog_fixes.py --catalog <path-to>/scenarios_data/shared/catalog/model_catalog.json --move-energy --writequay.io/assetopsbench/runtime:devwas built before the model move and the--no-devchange, so build locally as in step 1.~/.cache/assetopsbench/.By hand, without
run.shThe same suites with plain Harbor commands. Both run Stirrup with the Docker code sandbox.
.env, whichStirrupAgentloads itself. It forwards them into the agent phase only, and fails at once if its model's router pair is missing.--ae KEY=VALUEoverrides them for a single run.AOB_RUNTIME_IMAGEis the runtime image the tasks build FROM. Set it to the$RUNTIMEfrom step 1; unset, it falls back to the localassetopsbench/runtime:dev.AOB_CODE_TARis the code-sandbox image each trial loads.run.shkeeps its own copy under~/.cache/assetopsbench/, so for a run by hand, save it once:docker build -t assetops-code:dev \ -f src/agent/stirrup_agent/Dockerfile.code src/agent/stirrup_agent docker save assetops-code:dev -o ~/assetops-code.tarOpen suite. There is no
--scenario-root, so the tasks use the repo's own scenario data (src/couchdb/scenarios_data):Mini suite.
--scenario-rootpoints at the private suite, andoverlays/private-data.yamlmounts itsshared/fromAOB_PRIVATE_DIR. Give an absolute path: Compose would resolve a relative one against the task's folder, and the trial would fail at start.Each private task's container sees only its own
scenario_<id>/manifest.jsonand the suite'sshared/, mounted read-only. It sees no other scenario's folder and no answer files.harbor runwith--agent oraclein place of the--extra-docker-compose,--agent,--modeland--akflags. Expect 3 trials, 0 exceptions and reward 1.000; anything less is a setup problem, not a model problem.--akflags, and pass--ak code_enabled=false. Avoidcode_backend=local: agent code then runs next to CouchDB and the forwarded credentials, and could bypass the MCP tools the benchmark measures.-i 'fmsr-*'(or another category) to theharbor runpart for a quick subset.jobs/<timestamp>/<task>__<id>/:result.jsonhas the rewards and token totals,agent/trajectory.jsonthe trajectory in ATIF form, andverifier/the reward and evaluation logs.uv run harbor view jobsopens them in Harbor's viewer.uv run harbor jobs resume -p jobs/<timestamp>. Harbor refuses if the tasks changed since the job started, for example after regenerating from an edited template or suite.Testing & Validation
Unit tests on this branch's head (
ba120aa):uv run pytest src/ -k "not integration"gives 718 passed, 19 failed and 3 skipped. None of the 19 is caused by this PR:aafeedback_changes:evaluation/tests/test_static_json_scorer.py(6) andobservability/tests/test_file_exporter.py(2).test_invalid_sitetests iniot/tests/test_tools.py. This PR adds a test class to that file but doesn't change these tests. They fail withConnection refusedonlocalhost:5984when a local.envsets CouchDB variables and no CouchDB is running. Run them without.env, or with CouchDB up.test_catalog_checkpoints.py(12 passed) loads, fits and forecasts every local checkpoint at its catalog path,ttm_energy_168_24included.Scenario validation: Harbor's oracle agent (which writes the known answer) scored 1.000 with no exceptions:
overlays/private-data.yaml;AOB_RUNTIME_IMAGEpointing at a different image;runtime:102646e(ba120aachanges only the README).A 100% oracle pass is Harbor's own gate that a task can be scored at all.
Agent run: Stirrup with
litellm_proxy/azure/gpt-5.6-sol, the local code backend and the open suite gave 3 trials, 0 exceptions and a mean reward of 0.952.groundtruth,/testsor/solution.-i 'fmsr-*'and the Docker sandbox, set up as under "By hand", gave 5 trials, 0 exceptions, 3/5 passed and a mean reward of 0.760. The catalog, IoT and FMSR tools returned data in every trial.Not yet run end to end: the agent runs above, and the private and alternate-image oracle runs, predate docs(harbor): correct stale docs, docstrings and comments #590, fix(harbor): harden run.sh jobs, retries, checks and inputs #591, the energy model's move and the runtime image changes in
2970b5dand102646e.run.sh's changes in fix(harbor): harden run.sh jobs, retries, checks and inputs #591 were tested with Harbor and Docker stubbed (see fix(harbor): harden run.sh jobs, retries, checks and inputs #591). The open suite with the Docker sandbox, as written under "By hand", has not been run.Per-trial isolation, observed during a concurrent run:
This shows two rows with different project prefixes and no host ports.
No answers in the images:
runtime:102646e) finds nogroundtruth*,rubric.json,reference_answer*,scenario_meta.jsonor.envfiles and no.git. None of the local-only folders.dockerignoreexcludes (reports/,logs/,src/tmp/,.claude/,artifacts/output/,jobs/, generated tasks) is present.all-profile tasks carry onlymanifest.jsonin their image.Data integrity: generated tasks and run output are gitignored, and no scenario data is duplicated into the repo.
Reviewer notes
Base branch:
aafeedback_changesis already fully merged intomain. The only commits onmainthat it lacks are fix: distinct error message when a CouchDB database is missing #554's, and this branch already contains them. Retargeting tomainwould leave the same change, minus fix: distinct error message when a CouchDB database is missing #554's files.Committed model weights: about 41 MB of
safetensors, all underartifacts/tsfm_models/, as plain git files rather than Git LFS. The largest file is 20 MB.ttm_energy_168_24moved. It sat inartifacts/output/tuned_models/, whichartifacts/README.mdreserves for agent output that is never committed. It now ships fromartifacts/tsfm_models/, and the repo's catalog points there.A catalog outside the repo that still names the old path fails for this model once the runtime image is rebuilt. The private suite's
shared/catalog/model_catalog.jsondoes. Update it with theapply_catalog_fixes.pycommand under "End to end" before the next private run on a new image.Ground truth in the image: fixed.
manifest.jsonand the files it names into their image (fix(harbor): run.sh mounts only the private suite's shared/ #586).git archive HEAD, with the open scenarios' answer files and.gitexcluded (fix(harbor): build the runtime image from a clean clone of HEAD #587).shared/(fix(harbor): run.sh mounts only the private suite's shared/ #586).run.sh's per-job task copies stay in the repo's gitignoreddatasets/, not beside the results (fix(harbor): harden run.sh jobs, retries, checks and inputs #591).Fixed since the first review:
COUCHDB_URL;038d900and2daca75, which also cut the Harbor docs down to what a reader needs now);pyproject.tomland a duplicate.gitignorerule;581a538);2970b5d);102646e). A source edit now rebuilds in seconds, and the image dropped from 8.75 GB to 6.5 GB.Known issues, not fixed in this PR, most severe first:
stirrup.pyforwards every provider credential it finds (AWS, Anthropic, OpenAI, Gemini, watsonx) into the agent container, whatever router the model uses. Undercode_backend=local, agent code can read them.run.shuses the Docker sandbox, whose code containers don't get them.template/task.tomlhas noTOKENROUTER_*variables. Atokenrouter/judge therefore can't authenticate, and everyllm_judgescenario scores 0 without an error.code_backend=local, the agent is root in the container the verifier runs in. It could write a run record to/logs/agentor edit/opt/aob. The Docker sandbox, whichrun.shuses, isn't affected.plan-executeandstirrup-agentstill default to a watsonx model, which FMSR rejects, so a run with no--model-idhas no FMSR generate tools.run.shrefuses such a model unlessFMSR_MODEL_IDnames a router model, and the docs no longer use one (docs(harbor): correct stale docs, docstrings and comments #590). The command-line tools are unchanged.run.shedges: two processes taking over a dead job's lock in the same instant could both proceed; Harbor still loads a.env.localfrom the repo root, if there is one; old code tars and task copies are never cleaned up.