Skip to content

fix(harbor): harden run.sh jobs, retries, checks and inputs - #591

Merged
ShuxinLin merged 3 commits into
feature/harbor-integrationfrom
fix/harbor-run-job-identity
Sep 30, 2026
Merged

ShuxinLin merged 3 commits into
feature/harbor-integrationfrom
fix/harbor-run-job-identity

Conversation

@ShuxinLin

@ShuxinLin ShuxinLin commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Fixes the benchmarks/harbor/run.sh problems found while reviewing #558. They hold even when the runtime image is stable. Each commit covers one group:

1. Job identity, task copies, exit status (2494f1e)

  • Job names. A job was named after the model alone, and a resume runs the job's saved settings. So -m "X high" -m "X max" resumed the high job for max, and a different -p in the same results folder couldn't resume. Jobs are now stirrup_agent__<profile>__<model>[__<effort>].
  • Task copies. Tasks were regenerated on every run, and Harbor refuses to resume once the tasks differ from its lock. Any template, suite or generator change made every job unresumable. Each new job now gets its own copy, generated once and reused on resume.
    • The copies hold every scenario's answers, so they go in the repo's gitignored benchmarks/harbor/datasets/jobs/, not into the results folder.
    • This also removes the shared task folder that a second run.sh could delete mid-run.
  • Exit status. harbor run … || true hid the only failures harbor run reports: a job that can't start at all. A skipped router didn't count either. Both now make the script exit non-zero.

2. Retries, code tar, key checks, env file, lock, paths (66eb514)

  • Resume retries. A resume reran only NonZeroAgentExitCodeError. Harbor matches exact class names, so trials that failed on rate limits, network or auth errors, Ctrl-C, or environment and verifier failures stayed failed.

    • run.sh now lists every error type that isn't the model's own work. Timeouts, context and output limits, and safety refusals stay as results.
    • test_run_sh.py checks the list against the installed Harbor, and fails when Harbor adds an agent error subclass the list doesn't place.
  • Code tar. It was built once into ~/assetops-code.tar and reused forever. Now:

    • it's built on every run, which the Docker cache makes cheap;
    • it's saved once per image id under AOB_CODE_TAR_DIR (default ~/.cache/assetopsbench);
    • each job records its tar, and a resume loads that one.

    So a job never switches code image, and a rebuild never rewrites a tar a job is using.

  • Router check. It only checked that the router's URL answered. Now:

    • a key the router rejects (401 or 403 from /models) skips the model;
    • FMSR_MODEL_ID's router is checked too;
    • a model without a router prefix is skipped unless FMSR_MODEL_ID names a router model. Otherwise the FMSR server rejects it and fmsr scenarios run without generate_failure_modes.
  • Env file. ENV_FILE wasn't the only credential source, because StirrupAgent also loaded the repo's .env and filled the gaps. When AOB_ENV_FILE is set, StirrupAgent now reads that file instead of searching, and run.sh sets it.

  • Lock. A per-job lock stops a second run.sh from working on the same job. The lock holds its owner's PID, so a lock left by a crashed run is taken over.

  • Paths. Relative -s, -l, -p, ENV_FILE and AOB_CODE_TAR_DIR now resolve against your current directory, not the repo root. The defaults are still the repo's.

3. Suite guard (d9eb32d), found while reviewing the task copies. A resume reuses the job's own manifests but mounts shared/ from the current -s. run.sh now records the suite beside each job and refuses a resume from another one, the same way it handles another runtime image.

Behaviour changes

  • Jobs started under the old stirrup_agent__<model> names aren't resumed. A rerun starts new jobs beside them.

  • The first run saves a new code tar (about 450 MB) under ~/.cache/assetopsbench. ~/assetops-code.tar is no longer used, and neither is AOB_CODE_TAR from your shell.

  • Documentation / Tutorial update

Testing

  • uv run pytest src/ -k "not integration" gives 718 passed, 19 failed and 3 skipped. The 19 failures are the same ones already on the branch, listed in feat(harbor): run AssetOpsBench scenarios as Harbor tasks #558. The 4 new tests are:
    • test_aob_env_file_replaces_the_dotenv_search, which fails on the old code;
    • three in test_run_sh.py, which catch both a misspelled name and a subclass left off the list.
  • Dry runs of run.sh, with Harbor and Docker stubbed. The task generator runs for real, and so does the router check, against a local server that accepts one key and answers 404 on the TokenRouter path:
Case Result
Two efforts plus a second model, run from another folder with relative -s -l -p ENV_FILE 3 jobs, 1 code tar, AOB_ENV_FILE and AOB_CODE_TAR passed to Harbor, exit 0
Rerun 3 resumes with all 23 filters and the recorded tar, exit 0
Code image changes the old job resumes on its old tar, a new job gets the new tar
Key rejected / router down skipped, exit 1
watsonx model: without FMSR_MODEL_ID / with a router FMSR_MODEL_ID / with a non-router one skipped / runs / skipped
Lock held by a live process / left by a dead one skipped with exit 1 / taken over, lock released
Two run.sh on one job at the same time the second skips it, Harbor runs once
The job's code tar is gone / another -s refused, exit 1
  • Real routers: the configured LiteLLM and TokenRouter endpoints both answer /models with 200 for the real key and 401 for a bogus one, which is what the check assumes.
  • Not run: end to end with real Harbor and Docker.

Checklist

  • I have signed off my commits (DCO).

ShuxinLin and others added 3 commits September 30, 2026 19:21
Four run.sh problems that hold even with a stable runtime image.

- A job was named after the model alone, and a resume runs the job's saved
  config. So `-m "X high" -m "X max"` resumed the high job for max and ran no
  max trials, and another -p in the same leaderboard directory hit Harbor's
  lock and could not resume. Jobs are now
  stirrup_agent__<profile>__<model>[__<effort>].
- Tasks were regenerated on every run, and Harbor refuses to resume a job
  whose tasks differ from its lock. Any change to the template (comments
  included), the suite's scenario files or the generator made every job
  unresumable. Each new job now gets its own copy of the tasks, generated once
  when it starts and reused on resume. The copies hold the answers, so they go
  to the repo's gitignored benchmarks/harbor/datasets/jobs/, keyed by the job's
  path, not into the leaderboard directory.
- One shared task folder meant a second run.sh deleted tasks that a running
  job was still reading. Per-job copies remove the shared folder.
- `harbor run ... || true` hid the only failures harbor run reports: it exits 0
  when trials fail and non-zero when the job itself cannot run (a rejected
  config, or StirrupAgent's credential check aborting the job as its first
  trial starts). That, and a model skipped for an unreachable router, now set a
  non-zero exit. A resume whose task copy is gone says so.

Jobs started under the old names are not resumed; a rerun starts new jobs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The rest of the run.sh problems found reviewing #558.

- Resume reran only NonZeroAgentExitCodeError. Harbor matches exact class
  names, so its subclasses (rate limits, network, auth), Ctrl-C, and
  environment and verifier failures stayed failed. run.sh now lists every type
  that is not the model's own work; timeouts, context and output limits and
  safety refusals stay as results. test_run_sh.py checks the list against the
  installed Harbor and fails when a new agent error subclass is unplaced.
- The code sandbox tar was built once and reused forever. It is now built each
  run (cached), saved once per image id under AOB_CODE_TAR_DIR (default
  ~/.cache/assetopsbench), and recorded per job; a resume loads the job's own
  tar, so a job never switches code image and a rebuild never rewrites a tar
  in use.
- The router check only probed the base URL. It now rejects a key the router
  refuses (401/403 from /models), also checks FMSR_MODEL_ID's router, and skips
  a model with no router prefix unless FMSR_MODEL_ID names one, since FMSR
  would reject it and fmsr scenarios would run without generate_failure_modes.
- ENV_FILE was not the only credential source: StirrupAgent also loaded the
  repo's .env and filled its gaps. StirrupAgent now reads AOB_ENV_FILE instead
  of searching when it is set, and run.sh sets it.
- Two run.sh processes could work on one job. A per-job lock (mkdir, with the
  owner's PID so a dead owner's lock is taken over) makes the second skip it.
- Relative -s, -l, -p, ENV_FILE and AOB_CODE_TAR_DIR resolved against the repo
  root. They now resolve against the caller's directory; the defaults are still
  the repo's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Found reviewing the per-job task copies. A resume reuses the job's own
manifests but mounts shared/ from the current -s, so a resume with another
suite would pair one suite's manifests with another suite's data. Before the
copies, Harbor's task lock refused that resume; now run.sh records the suite
beside the job (<job>.suite) and refuses it, as it does for another runtime
image.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
@ShuxinLin ShuxinLin changed the title fix(harbor): per-job names, tasks and exit status in run.sh fix(harbor): harden run.sh jobs, retries, checks and inputs Sep 30, 2026
@ShuxinLin
ShuxinLin merged commit cb058ab into feature/harbor-integration Sep 30, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant