Skip to content
This repository was archived by the owner on Oct 6, 2026. It is now read-only.

docs(harbor): correct stale docs, docstrings and comments - #590

Merged
ShuxinLin merged 1 commit into
feature/harbor-integrationfrom
fix/harbor-stale-docs
Sep 30, 2026
Merged

ShuxinLin merged 1 commit into
feature/harbor-integrationfrom
fix/harbor-stale-docs

Conversation

@ShuxinLin

Copy link
Copy Markdown
Collaborator

Description

Fixes text in the Harbor port that later changes left out of date. These were found while reviewing #558 (item 7 of its known issues, plus the watsonx examples in item 5). Nothing changes behaviour except the wording of one error message.

Changes

src/assetops_harbor/stirrup.py

  • Module docstring:
    • The example run no longer uses PYTHONPATH=agent, a watsonx model and --n-concurrent 16.
    • It no longer says the MCP servers get their environment through mcphub. The Stirrup runner passes it through mcp_server_env.
    • It no longer says the docker backend is "redundant" under Harbor. The code-sandbox overlay provides the daemon, and its code containers get neither CouchDB's hostname nor the forwarded credentials; local code gets both.
  • Docker-backend error: it told users to mount a Docker socket, and now names overlays/code-sandbox.yaml instead. It still contains "no Docker daemon", which test_stirrup.py matches.
  • _load_dotenv docstring: SETTING_ENV_VARS (FMSR_MODEL_ID) also reach the container.

src/agent/stirrup_agent/runner.py: the comment no longer mentions FMSR's "standalone watsonx default", which #583 removed.

benchmarks/harbor/README.md

  • Step 4 used a watsonx model, which FMSR now rejects, so that run had no generate_failure_modes. It now uses litellm_proxy/azure/gpt-5.6-sol and says why. --n-concurrent goes from 16 to 4.
  • Local backend: it said code runs next to the open scenarios' ground truth, which fix(harbor): build the runtime image from a clean clone of HEAD #587 took out of the image. It now says what that code does run next to: CouchDB, the forwarded credentials, and later the verifier.
  • Explicit env: the section now names mcp_server_env and says plan-execute uses it too.
  • Arms: the code_backend paragraph now describes the overlay path and the difference between the two backends.

benchmarks/harbor/CODE-SANDBOX.md: the run.sh -m example used the same watsonx model.

benchmarks/harbor/base-image/Dockerfile

  • It referred to a check_models_list.sh that doesn't exist. It now says nothing checks models.txt against the catalog.
  • The build line now matches the script's default tags.

generate_tasks.py and publish-images.sh: "corpus" becomes "suite" or "scenario data". publish-images.sh's example namespace is now quay.io/assetopsbench.

Left alone on purpose: the task template's comments (template/task.toml, template/environment/Dockerfile) are stale in the same way. Editing them changes every generated task's content hash, and Harbor then refuses to resume jobs started on the current tasks. They're better changed together with the next template change that does something.

  • Documentation / Tutorial update

Testing

  • uv run pytest src/assetops_harbor src/agent/tests/test_stirrup_mcp_env.py gives 24 passed.
  • bash -n passes on publish-images.sh.
  • ruff check finds nothing in the changed lines. stirrup_agent/runner.py already had an import-order warning and a BLE001, which this PR doesn't touch.

Checklist

  • I have signed off my commits (DCO).

Text the later Harbor changes left behind, found reviewing #558. No behaviour
changes except the wording of one error message.

- stirrup.py module docstring: drop the first-draft run line (PYTHONPATH=agent,
  a watsonx model, --n-concurrent 16) and the claim that MCP servers get the
  environment through mcphub; the Stirrup runner passes it through
  mcp_server_env. The docker backend is no longer "redundant": the
  code-sandbox overlay provides a daemon and keeps code off CouchDB's hostname
  and the credentials, which local does not.
- StirrupAgent's docker-backend error pointed at mounting a Docker socket; it
  now names the code-sandbox overlay. It keeps "no Docker daemon", which a test
  matches.
- _load_dotenv: SETTING_ENV_VARS (FMSR_MODEL_ID) reach the container too.
- stirrup_agent/runner.py: FMSR has no "standalone watsonx default" since #583.
- README step 4 and CODE-SANDBOX.md's run.sh example used a watsonx model,
  which FMSR now rejects, so those runs had no generate_failure_modes. They use
  litellm_proxy/azure/gpt-5.6-sol, and step 4 says why.
- README: the local backend no longer sits next to the open scenarios' ground
  truth (#587 removed it from the image); say what it does sit next to. The
  explicit-env section names mcp_server_env and plan-execute.
- base-image/Dockerfile: check_models_list.sh does not exist, so say nothing
  checks models.txt against the catalog; the build line matches the script's
  default tags.
- generate_tasks.py, publish-images.sh: "corpus" is now "suite"/"scenario
  data"; publish-images.sh's example namespace is quay.io/assetopsbench.

The task template's comments are stale in the same way but are left alone:
editing them changes every task's content hash, and Harbor then refuses to
resume jobs started on the current tasks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
@ShuxinLin
ShuxinLin merged commit 57dff2f into feature/harbor-integration Sep 30, 2026
5 checks passed
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant