fix(harbor): harden run.sh jobs, retries, checks and inputs - #591
Merged
ShuxinLin merged 3 commits intoSep 30, 2026
Merged
Conversation
Four run.sh problems that hold even with a stable runtime image. - A job was named after the model alone, and a resume runs the job's saved config. So `-m "X high" -m "X max"` resumed the high job for max and ran no max trials, and another -p in the same leaderboard directory hit Harbor's lock and could not resume. Jobs are now stirrup_agent__<profile>__<model>[__<effort>]. - Tasks were regenerated on every run, and Harbor refuses to resume a job whose tasks differ from its lock. Any change to the template (comments included), the suite's scenario files or the generator made every job unresumable. Each new job now gets its own copy of the tasks, generated once when it starts and reused on resume. The copies hold the answers, so they go to the repo's gitignored benchmarks/harbor/datasets/jobs/, keyed by the job's path, not into the leaderboard directory. - One shared task folder meant a second run.sh deleted tasks that a running job was still reading. Per-job copies remove the shared folder. - `harbor run ... || true` hid the only failures harbor run reports: it exits 0 when trials fail and non-zero when the job itself cannot run (a rejected config, or StirrupAgent's credential check aborting the job as its first trial starts). That, and a model skipped for an unreachable router, now set a non-zero exit. A resume whose task copy is gone says so. Jobs started under the old names are not resumed; a rerun starts new jobs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The rest of the run.sh problems found reviewing #558. - Resume reran only NonZeroAgentExitCodeError. Harbor matches exact class names, so its subclasses (rate limits, network, auth), Ctrl-C, and environment and verifier failures stayed failed. run.sh now lists every type that is not the model's own work; timeouts, context and output limits and safety refusals stay as results. test_run_sh.py checks the list against the installed Harbor and fails when a new agent error subclass is unplaced. - The code sandbox tar was built once and reused forever. It is now built each run (cached), saved once per image id under AOB_CODE_TAR_DIR (default ~/.cache/assetopsbench), and recorded per job; a resume loads the job's own tar, so a job never switches code image and a rebuild never rewrites a tar in use. - The router check only probed the base URL. It now rejects a key the router refuses (401/403 from /models), also checks FMSR_MODEL_ID's router, and skips a model with no router prefix unless FMSR_MODEL_ID names one, since FMSR would reject it and fmsr scenarios would run without generate_failure_modes. - ENV_FILE was not the only credential source: StirrupAgent also loaded the repo's .env and filled its gaps. StirrupAgent now reads AOB_ENV_FILE instead of searching when it is set, and run.sh sets it. - Two run.sh processes could work on one job. A per-job lock (mkdir, with the owner's PID so a dead owner's lock is taken over) makes the second skip it. - Relative -s, -l, -p, ENV_FILE and AOB_CODE_TAR_DIR resolved against the repo root. They now resolve against the caller's directory; the defaults are still the repo's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Found reviewing the per-job task copies. A resume reuses the job's own manifests but mounts shared/ from the current -s, so a resume with another suite would pair one suite's manifests with another suite's data. Before the copies, Harbor's task lock refused that resume; now run.sh records the suite beside the job (<job>.suite) and refuses it, as it does for another runtime image. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
7 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes the
benchmarks/harbor/run.shproblems found while reviewing #558. They hold even when the runtime image is stable. Each commit covers one group:1. Job identity, task copies, exit status (
2494f1e)-m "X high" -m "X max"resumed thehighjob formax, and a different-pin the same results folder couldn't resume. Jobs are nowstirrup_agent__<profile>__<model>[__<effort>].benchmarks/harbor/datasets/jobs/, not into the results folder.run.shcould delete mid-run.harbor run … || truehid the only failuresharbor runreports: a job that can't start at all. A skipped router didn't count either. Both now make the script exit non-zero.2. Retries, code tar, key checks, env file, lock, paths (
66eb514)Resume retries. A resume reran only
NonZeroAgentExitCodeError. Harbor matches exact class names, so trials that failed on rate limits, network or auth errors, Ctrl-C, or environment and verifier failures stayed failed.run.shnow lists every error type that isn't the model's own work. Timeouts, context and output limits, and safety refusals stay as results.test_run_sh.pychecks the list against the installed Harbor, and fails when Harbor adds an agent error subclass the list doesn't place.Code tar. It was built once into
~/assetops-code.tarand reused forever. Now:AOB_CODE_TAR_DIR(default~/.cache/assetopsbench);So a job never switches code image, and a rebuild never rewrites a tar a job is using.
Router check. It only checked that the router's URL answered. Now:
/models) skips the model;FMSR_MODEL_ID's router is checked too;FMSR_MODEL_IDnames a router model. Otherwise the FMSR server rejects it and fmsr scenarios run withoutgenerate_failure_modes.Env file.
ENV_FILEwasn't the only credential source, because StirrupAgent also loaded the repo's.envand filled the gaps. WhenAOB_ENV_FILEis set, StirrupAgent now reads that file instead of searching, andrun.shsets it.Lock. A per-job lock stops a second
run.shfrom working on the same job. The lock holds its owner's PID, so a lock left by a crashed run is taken over.Paths. Relative
-s,-l,-p,ENV_FILEandAOB_CODE_TAR_DIRnow resolve against your current directory, not the repo root. The defaults are still the repo's.3. Suite guard (
d9eb32d), found while reviewing the task copies. A resume reuses the job's own manifests but mountsshared/from the current-s.run.shnow records the suite beside each job and refuses a resume from another one, the same way it handles another runtime image.Behaviour changes
Jobs started under the old
stirrup_agent__<model>names aren't resumed. A rerun starts new jobs beside them.The first run saves a new code tar (about 450 MB) under
~/.cache/assetopsbench.~/assetops-code.taris no longer used, and neither isAOB_CODE_TARfrom your shell.Documentation / Tutorial update
Testing
uv run pytest src/ -k "not integration"gives 718 passed, 19 failed and 3 skipped. The 19 failures are the same ones already on the branch, listed in feat(harbor): run AssetOpsBench scenarios as Harbor tasks #558. The 4 new tests are:test_aob_env_file_replaces_the_dotenv_search, which fails on the old code;test_run_sh.py, which catch both a misspelled name and a subclass left off the list.run.sh, with Harbor and Docker stubbed. The task generator runs for real, and so does the router check, against a local server that accepts one key and answers 404 on the TokenRouter path:-s -l -p ENV_FILEAOB_ENV_FILEandAOB_CODE_TARpassed to Harbor, exit 0FMSR_MODEL_ID/ with a routerFMSR_MODEL_ID/ with a non-router onerun.shon one job at the same time-s/modelswith 200 for the real key and 401 for a bogus one, which is what the check assumes.Checklist