This repository was archived by the owner on Oct 6, 2026. It is now read-only.
docs(harbor): correct stale docs, docstrings and comments - #590
Merged
Merged
Conversation
Text the later Harbor changes left behind, found reviewing #558. No behaviour changes except the wording of one error message. - stirrup.py module docstring: drop the first-draft run line (PYTHONPATH=agent, a watsonx model, --n-concurrent 16) and the claim that MCP servers get the environment through mcphub; the Stirrup runner passes it through mcp_server_env. The docker backend is no longer "redundant": the code-sandbox overlay provides a daemon and keeps code off CouchDB's hostname and the credentials, which local does not. - StirrupAgent's docker-backend error pointed at mounting a Docker socket; it now names the code-sandbox overlay. It keeps "no Docker daemon", which a test matches. - _load_dotenv: SETTING_ENV_VARS (FMSR_MODEL_ID) reach the container too. - stirrup_agent/runner.py: FMSR has no "standalone watsonx default" since #583. - README step 4 and CODE-SANDBOX.md's run.sh example used a watsonx model, which FMSR now rejects, so those runs had no generate_failure_modes. They use litellm_proxy/azure/gpt-5.6-sol, and step 4 says why. - README: the local backend no longer sits next to the open scenarios' ground truth (#587 removed it from the image); say what it does sit next to. The explicit-env section names mcp_server_env and plan-execute. - base-image/Dockerfile: check_models_list.sh does not exist, so say nothing checks models.txt against the catalog; the build line matches the script's default tags. - generate_tasks.py, publish-images.sh: "corpus" is now "suite"/"scenario data"; publish-images.sh's example namespace is quay.io/assetopsbench. The task template's comments are stale in the same way but are left alone: editing them changes every task's content hash, and Harbor then refuses to resume jobs started on the current tasks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
7 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes text in the Harbor port that later changes left out of date. These were found while reviewing #558 (item 7 of its known issues, plus the watsonx examples in item 5). Nothing changes behaviour except the wording of one error message.
Changes
src/assetops_harbor/stirrup.pyPYTHONPATH=agent, a watsonx model and--n-concurrent 16.mcp_server_env.localcode gets both.overlays/code-sandbox.yamlinstead. It still contains "no Docker daemon", whichtest_stirrup.pymatches._load_dotenvdocstring:SETTING_ENV_VARS(FMSR_MODEL_ID) also reach the container.src/agent/stirrup_agent/runner.py: the comment no longer mentions FMSR's "standalone watsonx default", which #583 removed.benchmarks/harbor/README.mdgenerate_failure_modes. It now useslitellm_proxy/azure/gpt-5.6-soland says why.--n-concurrentgoes from 16 to 4.mcp_server_envand says plan-execute uses it too.code_backendparagraph now describes the overlay path and the difference between the two backends.benchmarks/harbor/CODE-SANDBOX.md: therun.sh -mexample used the same watsonx model.benchmarks/harbor/base-image/Dockerfilecheck_models_list.shthat doesn't exist. It now says nothing checksmodels.txtagainst the catalog.generate_tasks.pyandpublish-images.sh: "corpus" becomes "suite" or "scenario data".publish-images.sh's example namespace is nowquay.io/assetopsbench.Left alone on purpose: the task template's comments (
template/task.toml,template/environment/Dockerfile) are stale in the same way. Editing them changes every generated task's content hash, and Harbor then refuses to resume jobs started on the current tasks. They're better changed together with the next template change that does something.Testing
uv run pytest src/assetops_harbor src/agent/tests/test_stirrup_mcp_env.pygives 24 passed.bash -npasses onpublish-images.sh.ruff checkfinds nothing in the changed lines.stirrup_agent/runner.pyalready had an import-order warning and aBLE001, which this PR doesn't touch.Checklist