Skip to content

feat(harbor): run a full scenario suite in parallel - #559

Closed
ShuxinLin wants to merge 5 commits into
feature/harbor-integrationfrom
feat/harbor-corpus-run
Closed

ShuxinLin wants to merge 5 commits into
feature/harbor-integrationfrom
feat/harbor-corpus-run

Conversation

@ShuxinLin

@ShuxinLin ShuxinLin commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Runs a full scenario suite through Harbor the way benchmarks/run.sh runs it
serially, with the same all profile, the same stirrup-agent and the same
Docker code sandbox, but with scenarios in parallel, each trial on its own
CouchDB.

Stacked on #558 (base: feature/harbor-integration). #558 covers the open
profile, whose scenarios ship in the repo. An external suite such as
AssetOpsBenchScenarioGeneration/scenarios_data did not fit its template for
three reasons:

  • Its shared/ data is 2.1 GB (mostly shared/iot), and every manifest
    resolves against it. The runtime image only carries the repo's small copy.
  • all.yaml lists a scenario the suite lacks (wosr-62), and the generator
    aborted on the first missing folder.
  • Nothing drove it end to end: building images, generating tasks, one job per
    model, and resuming.

Changes

  • adapter/generate_tasks.py
    • --runtime-image: the image each task builds FROM.
    • --data-dir: where the suite lives in that image. The per-task layer copies
      the scenario there, and the healthcheck runs
      SCENARIOS_DATA_DIR=<dir> init_data.py <id>. The agent keeps the repo copy,
      as scenario_suite_runner does; it sets that variable for the data load only.
    • --skip-missing: warn and skip profile entries that aren't in the suite.
    • Fix: every task used to get scenario 1's description and the wosr
      keyword.
  • suite-image/Dockerfile: bakes the suite into one layer over
    assetopsbench/runtime:dev, shared by all tasks. A bind mount would go
    through Rancher/Docker Desktop's file sharing, which is slow.
  • run.sh: the Harbor counterpart of benchmarks/run.sh. It uses the same
    -s and -l flags and the same model list, plus -n for concurrency, -p
    for the profile and a repeatable -m "MODEL EFFORT". It builds the suite and
    sandbox images, regenerates the dataset, and runs one Harbor job per model at
    <leaderboard>/harbor-jobs/stirrup_agent__<model>. Re-running resumes that job,
    the equivalent of --skip-existing. Credentials reach Harbor only through
    uv run --env-file .env.
  • README: a section on running a full suite.

Type of Change

  • Infrastructure / Tooling Improvement

Industry Relevance

The published leaderboard runs all 232 all-profile scenarios per model, one
at a time against a shared CouchDB. With Harbor each scenario gets an isolated
database, so the same run parallelizes safely.

Related Issues

Testing & Validation

Run locally: Rancher Desktop with 8 CPUs and 24 GB, against
AssetOpsBenchScenarioGeneration/scenarios_data with the all profile.

  • Task generation: 231 tasks from 232 profile entries (wosr-62 skipped).
    17 are scored fmea and 214 static_json, from each scenario_meta.json.
    Every file the manifests reference exists. The largest scenario loads 9.2 MB,
    well inside the 300 s healthcheck.

  • Oracle, real containers: one scenario per category, --n-concurrent 4.
    7/7 passed, 0 exceptions, reward 1.000, in 54 s.

  • Stirrup smoke via run.sh: litellm_proxy/aws/claude-opus-5 at high,
    code sandbox on, -n 4, one scenario per category. Every trial had its own
    couchdb and dind containers and no host ports. Agents called the MCP tools
    (iot, wo, fmsr, tsfm, utilities) and ran code_exec in the Docker
    sandbox, and spilled tool results were readable through /workspace-share.

    Task Reward Passed Input / output tokens
    car-151 1.0 ✅ 305k / 3.2k
    fcc-301 0.857 ❌ 313k / 10.3k
    fmea-9001 0.933 ✅ 203k / 14.8k
    fmsr-901 0.2 ❌ 1.04M / 15.2k
    health-401 1.0 ✅ 193k / 1.4k
    tsfm-1001 1.0 ✅ 623k / 10.6k
    wosr-1 0.0 ❌ 462k / 13.5k

    4/7 passed, mean reward 0.713, wall time 18 m 10 s.

    No harness exceptions. The three misses are the model's answers, not
    infrastructure failures. For example, on wosr-1 the agent's sandbox code read
    all 1,076 work orders from the trial's CouchDB and counted 163 failures
    against a ground truth of 211.

  • Harbor healthcheck semantics: BaseEnvironment.run_healthcheck returns
    on the first success, so init_data.py (drop and reload) runs once per trial
    and never again during the agent's run.

  • Data Integrity: no suite data or credentials are committed. Generated
    tasks, datasets and job output are gitignored or outside the repo.

  • Full 231-scenario run: not started yet.

Checklist

  • ruff check and ruff format pass on the changed Python file.
  • I have performed a self-review of my code.
  • I have updated the documentation (benchmarks/harbor/README.md).
  • I have signed off my commits (DCO).

generate_tasks.py gains three options for corpora that do not ship in the
repo, such as AssetOpsBenchScenarioGeneration/scenarios_data:

- --runtime-image: the image each task builds FROM.
- --data-dir: where the corpus lives inside that image. The per-task
  layer copies the scenario there and the healthcheck runs init_data.py
  with SCENARIOS_DATA_DIR pointed at it; the agent keeps the repo copy,
  as scenario_suite_runner does.
- --skip-missing: warn and skip profile scenarios absent from the corpus
  (all.yaml lists wosr-62, which the corpus lacks).

corpus-image/Dockerfile bakes the corpus (2.1 GB, mostly shared/iot) into
one layer over the runtime image, so tasks share it instead of each
carrying it in their build context.

Also stop copying scenario 1's description and the wosr keyword into
every generated task.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Runs a scenario corpus through Harbor with the same profile, the same
stirrup-agent and the same Docker code sandbox as benchmarks/run.sh, but
with scenarios in parallel, each trial on its own CouchDB.

- Builds assetopsbench/runtime:corpus from -s, and the code sandbox tar
  for the per-trial Docker-in-Docker daemon.
- Regenerates the dataset from scratch (the generator never removes stale
  task folders), skipping profile entries the corpus lacks.
- One Harbor job per model under <leaderboard>/harbor-jobs; re-running
  resumes it, the equivalent of --skip-existing.
- Credentials reach the Harbor process through uv run --env-file only.

Uses [[ -z "${arr[*]+set}" ]] and ${arr[@]+...} so empty arrays work
under set -u in macOS's bash 3.2.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
corpus-image/ -> suite-image/, assetopsbench/runtime:corpus ->
assetopsbench/runtime:suite, /opt/corpus/scenarios_data ->
/opt/suite/scenarios_data, and the generated dataset
assetopsbench-corpus -> assetopsbench-suite, with matching wording in
run.sh, the generator's help and the README.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Completes 533bc5a, which only moved corpus-image/ to suite-image/:
assetopsbench/runtime:corpus -> assetopsbench/runtime:suite,
/opt/corpus/scenarios_data -> /opt/suite/scenarios_data, the dataset
assetopsbench-corpus -> assetopsbench-suite, and the wording in run.sh,
the generator's help and the README.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
@ShuxinLin ShuxinLin closed this Sep 28, 2026
@ShuxinLin
ShuxinLin deleted the feat/harbor-corpus-run branch September 28, 2026 16:09
@ShuxinLin ShuxinLin changed the title feat(harbor): run a full scenario corpus in parallel feat(harbor): run a full scenario suite in parallel Sep 28, 2026
@ShuxinLin

Copy link
Copy Markdown
Collaborator Author

Superseded by #560. Renaming the head branch to feat/harbor-suite-run (corpus → suite) closed this PR; same commits, now under the new branch name.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant