Conversation
generate_tasks.py gains three options for corpora that do not ship in the repo, such as AssetOpsBenchScenarioGeneration/scenarios_data: - --runtime-image: the image each task builds FROM. - --data-dir: where the corpus lives inside that image. The per-task layer copies the scenario there and the healthcheck runs init_data.py with SCENARIOS_DATA_DIR pointed at it; the agent keeps the repo copy, as scenario_suite_runner does. - --skip-missing: warn and skip profile scenarios absent from the corpus (all.yaml lists wosr-62, which the corpus lacks). corpus-image/Dockerfile bakes the corpus (2.1 GB, mostly shared/iot) into one layer over the runtime image, so tasks share it instead of each carrying it in their build context. Also stop copying scenario 1's description and the wosr keyword into every generated task. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Runs a scenario corpus through Harbor with the same profile, the same
stirrup-agent and the same Docker code sandbox as benchmarks/run.sh, but
with scenarios in parallel, each trial on its own CouchDB.
- Builds assetopsbench/runtime:corpus from -s, and the code sandbox tar
for the per-trial Docker-in-Docker daemon.
- Regenerates the dataset from scratch (the generator never removes stale
task folders), skipping profile entries the corpus lacks.
- One Harbor job per model under <leaderboard>/harbor-jobs; re-running
resumes it, the equivalent of --skip-existing.
- Credentials reach the Harbor process through uv run --env-file only.
Uses [[ -z "${arr[*]+set}" ]] and ${arr[@]+...} so empty arrays work
under set -u in macOS's bash 3.2.
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
corpus-image/ -> suite-image/, assetopsbench/runtime:corpus -> assetopsbench/runtime:suite, /opt/corpus/scenarios_data -> /opt/suite/scenarios_data, and the generated dataset assetopsbench-corpus -> assetopsbench-suite, with matching wording in run.sh, the generator's help and the README. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Completes 533bc5a, which only moved corpus-image/ to suite-image/: assetopsbench/runtime:corpus -> assetopsbench/runtime:suite, /opt/corpus/scenarios_data -> /opt/suite/scenarios_data, the dataset assetopsbench-corpus -> assetopsbench-suite, and the wording in run.sh, the generator's help and the README. Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Collaborator
Author
|
Superseded by #560. Renaming the head branch to |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Runs a full scenario suite through Harbor the way
benchmarks/run.shruns itserially, with the same
allprofile, the samestirrup-agentand the sameDocker code sandbox, but with scenarios in parallel, each trial on its own
CouchDB.
Stacked on #558 (base:
feature/harbor-integration). #558 covers the openprofile, whose scenarios ship in the repo. An external suite such as
AssetOpsBenchScenarioGeneration/scenarios_datadid not fit its template forthree reasons:
shared/data is 2.1 GB (mostlyshared/iot), and every manifestresolves against it. The runtime image only carries the repo's small copy.
all.yamllists a scenario the suite lacks (wosr-62), and the generatoraborted on the first missing folder.
model, and resuming.
Changes
adapter/generate_tasks.py--runtime-image: the image each task buildsFROM.--data-dir: where the suite lives in that image. The per-task layer copiesthe scenario there, and the healthcheck runs
SCENARIOS_DATA_DIR=<dir> init_data.py <id>. The agent keeps the repo copy,as
scenario_suite_runnerdoes; it sets that variable for the data load only.--skip-missing: warn and skip profile entries that aren't in the suite.descriptionand thewosrkeyword.
suite-image/Dockerfile: bakes the suite into one layer overassetopsbench/runtime:dev, shared by all tasks. A bind mount would gothrough Rancher/Docker Desktop's file sharing, which is slow.
run.sh: the Harbor counterpart ofbenchmarks/run.sh. It uses the same-sand-lflags and the same model list, plus-nfor concurrency,-pfor the profile and a repeatable
-m "MODEL EFFORT". It builds the suite andsandbox images, regenerates the dataset, and runs one Harbor job per model at
<leaderboard>/harbor-jobs/stirrup_agent__<model>. Re-running resumes that job,the equivalent of
--skip-existing. Credentials reach Harbor only throughuv run --env-file .env.Type of Change
Industry Relevance
The published leaderboard runs all 232
all-profile scenarios per model, oneat a time against a shared CouchDB. With Harbor each scenario gets an isolated
database, so the same run parallelizes safely.
Related Issues
Testing & Validation
Run locally: Rancher Desktop with 8 CPUs and 24 GB, against
AssetOpsBenchScenarioGeneration/scenarios_datawith theallprofile.Task generation: 231 tasks from 232 profile entries (
wosr-62skipped).17 are scored
fmeaand 214static_json, from eachscenario_meta.json.Every file the manifests reference exists. The largest scenario loads 9.2 MB,
well inside the 300 s healthcheck.
Oracle, real containers: one scenario per category,
--n-concurrent 4.7/7 passed, 0 exceptions, reward 1.000, in 54 s.
Stirrup smoke via
run.sh:litellm_proxy/aws/claude-opus-5athigh,code sandbox on,
-n 4, one scenario per category. Every trial had its owncouchdbanddindcontainers and no host ports. Agents called the MCP tools(
iot,wo,fmsr,tsfm,utilities) and rancode_execin the Dockersandbox, and spilled tool results were readable through
/workspace-share.4/7 passed, mean reward 0.713, wall time 18 m 10 s.
No harness exceptions. The three misses are the model's answers, not
infrastructure failures. For example, on wosr-1 the agent's sandbox code read
all 1,076 work orders from the trial's CouchDB and counted 163 failures
against a ground truth of 211.
Harbor healthcheck semantics:
BaseEnvironment.run_healthcheckreturnson the first success, so
init_data.py(drop and reload) runs once per trialand never again during the agent's run.
Data Integrity: no suite data or credentials are committed. Generated
tasks, datasets and job output are gitignored or outside the repo.
Full 231-scenario run: not started yet.
Checklist
ruff checkandruff formatpass on the changed Python file.benchmarks/harbor/README.md).