Skip to content

feat(harbor): run AssetOpsBench scenarios as Harbor tasks - #558

Merged
ShuxinLin merged 72 commits into
aafeedback_changesfrom
feature/harbor-integration
Oct 1, 2026
Merged

ShuxinLin merged 72 commits into
aafeedback_changesfrom
feature/harbor-integration

Conversation

@DhavalRepo18

@DhavalRepo18 DhavalRepo18 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Description

This PR does two things:

  • It adds an opt-in way to run AssetOpsBench scenarios in parallel, as Harbor tasks with one task per scenario. Stirrup plugs in as a custom agent, and CouchDB runs as a per-trial Compose sidecar.
  • It brings along fixes to the MCP servers and agent runners, a new rule for choosing the FMSR model, and changes to the TSFM model catalog.

Harbor gives each trial its own Compose project and names every container, network and volume after it, so each scenario gets its own CouchDB. That makes parallel scenario runs safe by design, rather than safe only if they're carefully sequenced.

Why this matters. The current path can't be run in parallel safely:

  • scenario_suite_runner.py resets one shared CouchDB before each scenario, and loader._ensure_db drops each database before loading it.
  • Run that loop concurrently and it deletes a live scenario's iot, workorder and vibration databases partway through its run.
  • The agent then reads empty results, writes a plausible answer, and the trajectory scores as legitimate.
  • Nothing crashes, so parallel numbers would be quietly wrong rather than visibly broken.

Existing behaviour doesn't change. scenario_suite_runner.py is untouched, and Harbor is an opt-in extra (uv sync --extra harbor).

Size: 73 files, +8,132 / −249. Most of that is uv.lock, about 41 MB of TTM weights under artifacts/, and the catalog scripts under benchmarks/harbor/scripts/. The PR targets aafeedback_changes, not main; see the reviewer notes.

What the code changes do

1. Harbor integration (the main feature), benchmarks/harbor/

2. Clearer errors when a CouchDB database is missing (#554)

The iot, fmsr, wo and vibration servers now tell a missing database apart from a wrong key. Before, both produced errors like "unknown asset_id". Now a missing database returns "the data source does not exist in this environment; ... do not retry", without naming the database. Tests cover each server.

3. Agent runners pass the environment to MCP servers

Without an explicit env, the MCP SDK's stdio client passes on only HOME, PATH and a few other variables, so the servers never saw COUCHDB_URL and silently fell back to localhost:5984. The new mcp_server_env in src/agent/runner.py forwards the full environment, and the Stirrup, plan-execute, deep-agent and openai-agent runners now use it. This fixes a real bug and applies outside Harbor too. test_stirrup_mcp_env.py covers the Stirrup case.

4. FMSR model (#583)

  • generate_failure_modes takes its model only from FMSR_MODEL_ID; there is no built-in default.
  • Every agent runner sets it to the agent's own --model-id unless it is set explicitly (fmsr_env_overrides in src/agent/runner.py).
  • Only the litellm_proxy/ and tokenrouter/ prefixes are accepted; watsonx support is dropped.
  • The tool also works when the catalog hasn't been initialised.
  • INSTRUCTIONS.md and docs/running_benchmark.md document the new behaviour.

5. TSFM model catalog and weights

  • The repo's shared/tsfm/model_catalog.json drops ttm_96_28 and now lists six models: four local TTM checkpoints, the fine-tuned energy model ttm_energy_168_24, and ttm_1024_192_hub.
  • The local checkpoints are committed under artifacts/tsfm_models/ as plain git files, not Git LFS (about 41 MB). artifacts/output/, for checkpoints an agent fine-tunes during a trial, is now ignored by git and Docker.
  • Nine Python scripts under benchmarks/harbor/scripts/ generate, audit and fix the catalog, preload Hub models, fetch TTM checkpoints and smoke-test forecasts and tasks.
  • src/servers/tsfm/tests/test_catalog_checkpoints.py checks that every local checkpoint in the catalog resolves, fits and forecasts.

6. Dependencies (pyproject.toml)

  • tsfm moves from a dependency group to an optional extra and grows a lot: transformers[torch]>=5.3, gluonts, lightning, toto-models and others.
  • A new harbor optional extra is added, and src/assetops_harbor is added to the packages.
  • sktime goes up to >=1.2.0, and numba and pyod are added to the dev group.

Type of Change

  • Infrastructure / Tooling Improvement

Industry Relevance

Third-party evaluators need concurrency and per-trial isolation before they can publish AssetOpsBench numbers, and this PR builds both into the runner. Each result also records the runtime image's commit, and a run can be pinned to one build with runtime:<commit>.

Related Issues

Running the benchmark

Run every command from the repo root. You need Docker running (at least 4 GB of memory) and model credentials in .env (LITELLM_* and/or TOKENROUTER_*). Use a litellm_proxy/ or tokenrouter/ model, or set FMSR_MODEL_ID to one: FMSR's generate_failure_modes accepts only those two routers. benchmarks/harbor/QUICKSTART.md has the step-by-step version and common problems.

End to end (as of ba120aa)

# 0. Once: install the Harbor extra.
uv sync --dev --extra harbor

# 1. Build the runtime image from HEAD, and pin the run to this build's commit tag.
bash benchmarks/harbor/scripts/build-runtime-image.sh
RUNTIME=assetopsbench/runtime:$(git rev-parse HEAD | cut -c1-7)

# 2. Stirrup on the private mini suite, with the Docker code sandbox.
#    Rerun the same command to resume.
bash benchmarks/harbor/run.sh \
  -s <path-to>/AssetOpsBenchScenarioGeneration/scenarios_data \
  -l ~/AssetOpsBenchRuns/leaderboard \
  -p benchmarks/scenario_suite/mini.yaml \
  -r "$RUNTIME" -n 4 \
  -m "litellm_proxy/azure/gpt-5.6-sol max"

# 3. Results: Harbor's viewer, and per-category means for one job.
uv run harbor view ~/AssetOpsBenchRuns/leaderboard/harbor-jobs
uv run python benchmarks/harbor/metric.py --job-dir \
  ~/AssetOpsBenchRuns/leaderboard/harbor-jobs/stirrup_agent__mini__litellm_proxy-azure-gpt-5.6-sol__max
  • The -s folder: run.sh wires the private suite from it. It mounts the suite's shared/ read-only into each trial and generates each job's tasks from it.
  • Defaults: leave out -m to run the 8 default models, and -p to run the all profile. Each trial runs a privileged dind sidecar, so keep -n around 4 on a laptop-sized Docker VM.
  • Private catalog first: its ttm_energy_168_24 must point at artifacts/tsfm_models/ before you run on a new image:
    uv run python benchmarks/harbor/scripts/apply_catalog_fixes.py --catalog <path-to>/scenarios_data/shared/catalog/model_catalog.json --move-energy --write
  • Not the published image yet: quay.io/assetopsbench/runtime:dev was built before the model move and the --no-dev change, so build locally as in step 1.
  • Disk: the first run saves the code-sandbox image (about 450 MB) to ~/.cache/assetopsbench/.

By hand, without run.sh

The same suites with plain Harbor commands. Both run Stirrup with the Docker code sandbox.

  • Credentials come from .env, which StirrupAgent loads itself. It forwards them into the agent phase only, and fails at once if its model's router pair is missing. --ae KEY=VALUE overrides them for a single run.
  • AOB_RUNTIME_IMAGE is the runtime image the tasks build FROM. Set it to the $RUNTIME from step 1; unset, it falls back to the local assetopsbench/runtime:dev.
  • AOB_CODE_TAR is the code-sandbox image each trial loads. run.sh keeps its own copy under ~/.cache/assetopsbench/, so for a run by hand, save it once:
docker build -t assetops-code:dev \
  -f src/agent/stirrup_agent/Dockerfile.code src/agent/stirrup_agent
docker save assetops-code:dev -o ~/assetops-code.tar

Open suite. There is no --scenario-root, so the tasks use the repo's own scenario data (src/couchdb/scenarios_data):

export AOB_RUNTIME_IMAGE=$RUNTIME AOB_CODE_TAR=~/assetops-code.tar

uv run python benchmarks/harbor/adapter/generate_tasks.py \
  --profile benchmarks/scenario_suite/open.yaml \
  --output-dir benchmarks/harbor/datasets/assetopsbench-open \
  --dataset-name assetopsbench/open \
  --overwrite && \
uv run harbor run -y \
  -p benchmarks/harbor/datasets/assetopsbench-open \
  --extra-docker-compose benchmarks/harbor/overlays/code-sandbox.yaml \
  --agent assetops_harbor.stirrup:StirrupAgent \
  --model litellm_proxy/azure/gpt-5.6-sol \
  --ak code_enabled=true --ak code_backend=docker --ak allow_docker_backend=true \
  --ak workspace_dir=/workspace-share \
  --n-concurrent 4

Mini suite. --scenario-root points at the private suite, and overlays/private-data.yaml mounts its shared/ from AOB_PRIVATE_DIR. Give an absolute path: Compose would resolve a relative one against the task's folder, and the trial would fail at start.

export AOB_PRIVATE_DIR=<path-to>/AssetOpsBenchScenarioGeneration/scenarios_data
export AOB_RUNTIME_IMAGE=$RUNTIME AOB_CODE_TAR=~/assetops-code.tar

uv run python benchmarks/harbor/adapter/generate_tasks.py \
  --scenario-root "$AOB_PRIVATE_DIR" \
  --profile benchmarks/scenario_suite/mini.yaml \
  --output-dir benchmarks/harbor/datasets/assetopsbench-mini \
  --dataset-name assetopsbench/mini \
  --overwrite && \
uv run harbor run -y \
  -p benchmarks/harbor/datasets/assetopsbench-mini \
  --extra-docker-compose benchmarks/harbor/overlays/private-data.yaml \
  --extra-docker-compose benchmarks/harbor/overlays/code-sandbox.yaml \
  --agent assetops_harbor.stirrup:StirrupAgent \
  --model litellm_proxy/azure/gpt-5.6-sol \
  --ak code_enabled=true --ak code_backend=docker --ak allow_docker_backend=true \
  --ak workspace_dir=/workspace-share \
  --n-concurrent 4

Each private task's container sees only its own scenario_<id>/manifest.json and the suite's shared/, mounted read-only. It sees no other scenario's folder and no answer files.

  • Check before spending tokens: run the open suite's harbor run with --agent oracle in place of the --extra-docker-compose, --agent, --model and --ak flags. Expect 3 trials, 0 exceptions and reward 1.000; anything less is a setup problem, not a model problem.
  • Tools-only track: drop the code-sandbox overlay and the --ak flags, and pass --ak code_enabled=false. Avoid code_backend=local: agent code then runs next to CouchDB and the forwarded credentials, and could bypass the MCP tools the benchmark measures.
  • Subset: mini is 35 tasks. Add -i 'fmsr-*' (or another category) to the harbor run part for a quick subset.
  • Results land in jobs/<timestamp>/<task>__<id>/: result.json has the rewards and token totals, agent/trajectory.json the trajectory in ATIF form, and verifier/ the reward and evaluation logs. uv run harbor view jobs opens them in Harbor's viewer.
  • Resume: with the same exports, run uv run harbor jobs resume -p jobs/<timestamp>. Harbor refuses if the tasks changed since the job started, for example after regenerating from an edited template or suite.

Testing & Validation

  • Unit tests on this branch's head (ba120aa):

    uv run pytest src/ -k "not integration" gives 718 passed, 19 failed and 3 skipped. None of the 19 is caused by this PR:

    • 8 also fail on aafeedback_changes: evaluation/tests/test_static_json_scorer.py (6) and observability/tests/test_file_exporter.py (2).
    • 11 are the existing test_invalid_site tests in iot/tests/test_tools.py. This PR adds a test class to that file but doesn't change these tests. They fail with Connection refused on localhost:5984 when a local .env sets CouchDB variables and no CouchDB is running. Run them without .env, or with CouchDB up.

    test_catalog_checkpoints.py (12 passed) loads, fits and forecasts every local checkpoint at its catalog path, ttm_energy_168_24 included.

  • Scenario validation: Harbor's oracle agent (which writes the known answer) scored 1.000 with no exceptions:

    • on the 3 open tasks;
    • on 13 private tasks with overlays/private-data.yaml;
    • on 5 tasks (3 open, 2 private) with AOB_RUNTIME_IMAGE pointing at a different image;
    • on the 3 open tasks again with the current runtime image, runtime:102646e (ba120aa changes only the README).

    A 100% oracle pass is Harbor's own gate that a task can be scored at all.

  • Agent run: Stirrup with litellm_proxy/azure/gpt-5.6-sol, the local code backend and the open suite gave 3 trials, 0 exceptions and a mean reward of 0.952.

    • wosr-1 and wosr-2 scored 1.0.
    • wosr-3 scored 0.857. It matched 6 of 7 keys, and gave the end of the queried window rather than the last observation's timestamp.
    • No trajectory referenced groundtruth, /tests or /solution.
    • Mini with -i 'fmsr-*' and the Docker sandbox, set up as under "By hand", gave 5 trials, 0 exceptions, 3/5 passed and a mean reward of 0.760. The catalog, IoT and FMSR tools returned data in every trial.
    • fmsr-913 with feat(harbor): run mini/lite/all by mounting the private suite #577 (local code backend): the data loaded and the tools returned no errors.
  • Not yet run end to end: the agent runs above, and the private and alternate-image oracle runs, predate docs(harbor): correct stale docs, docstrings and comments #590, fix(harbor): harden run.sh jobs, retries, checks and inputs #591, the energy model's move and the runtime image changes in 2970b5d and 102646e. run.sh's changes in fix(harbor): harden run.sh jobs, retries, checks and inputs #591 were tested with Harbor and Docker stubbed (see fix(harbor): harden run.sh jobs, retries, checks and inputs #591). The open suite with the Docker sandbox, as written under "By hand", has not been run.

  • Per-trial isolation, observed during a concurrent run:

    uv run harbor run -y -p benchmarks/harbor/datasets/assetopsbench-open --agent oracle --n-concurrent 2 &
    sleep 20 && docker ps --format '{{.Names}}\t{{.Ports}}' | grep couchdb

    This shows two rows with different project prefixes and no host ports.

  • No answers in the images:

    • Searching the whole filesystem of the current runtime image (runtime:102646e) finds no groundtruth*, rubric.json, reference_answer*, scenario_meta.json or .env files and no .git. None of the local-only folders .dockerignore excludes (reports/, logs/, src/tmp/, .claude/, artifacts/output/, jobs/, generated tasks) is present.
    • All 231 all-profile tasks carry only manifest.json in their image.
  • Data integrity: generated tasks and run output are gitignored, and no scenario data is duplicated into the repo.

Reviewer notes

  • Base branch: aafeedback_changes is already fully merged into main. The only commits on main that it lacks are fix: distinct error message when a CouchDB database is missing #554's, and this branch already contains them. Retargeting to main would leave the same change, minus fix: distinct error message when a CouchDB database is missing #554's files.

  • Committed model weights: about 41 MB of safetensors, all under artifacts/tsfm_models/, as plain git files rather than Git LFS. The largest file is 20 MB.

  • ttm_energy_168_24 moved. It sat in artifacts/output/tuned_models/, which artifacts/README.md reserves for agent output that is never committed. It now ships from artifacts/tsfm_models/, and the repo's catalog points there.

    A catalog outside the repo that still names the old path fails for this model once the runtime image is rebuilt. The private suite's shared/catalog/model_catalog.json does. Update it with the apply_catalog_fixes.py command under "End to end" before the next private run on a new image.

  • Ground truth in the image: fixed.

  • Fixed since the first review:

    • resume reran only one error type (fix(harbor): harden run.sh jobs, retries, checks and inputs #591);
    • the MCP env test deleted a pre-set COUCHDB_URL;
    • stale docs, docstrings and template comments (docs(harbor): correct stale docs, docstrings and comments #590, 038d900 and 2daca75, which also cut the Harbor docs down to what a reader needs now);
    • duplicate dependencies in pyproject.toml and a duplicate .gitignore rule;
    • the energy model's location;
    • the code-sandbox loader's message when no code tar is given named no image (581a538);
    • the runtime image carried the dev group (pytest and the Jupyter stack), and a timed-out download from it failed a build (2970b5d);
    • every rebuild downloaded the 3.9 GB of Hub models again, and the image kept a 2.2 GB uv cache beside its venv (102646e). A source edit now rebuilds in seconds, and the image dropped from 8.75 GB to 6.5 GB.
  • Known issues, not fixed in this PR, most severe first:

    1. stirrup.py forwards every provider credential it finds (AWS, Anthropic, OpenAI, Gemini, watsonx) into the agent container, whatever router the model uses. Under code_backend=local, agent code can read them. run.sh uses the Docker sandbox, whose code containers don't get them.
    2. The verifier's environment in template/task.toml has no TOKENROUTER_* variables. A tokenrouter/ judge therefore can't authenticate, and every llm_judge scenario scores 0 without an error.
    3. Under code_backend=local, the agent is root in the container the verifier runs in. It could write a run record to /logs/agent or edit /opt/aob. The Docker sandbox, which run.sh uses, isn't affected.
    4. plan-execute and stirrup-agent still default to a watsonx model, which FMSR rejects, so a run with no --model-id has no FMSR generate tools. run.sh refuses such a model unless FMSR_MODEL_ID names a router model, and the docs no longer use one (docs(harbor): correct stale docs, docstrings and comments #590). The command-line tools are unchanged.
    5. run.sh edges: two processes taking over a dead job's lock in the same instant could both proceed; Harbor still loads a .env.local from the repo root, if there is one; old code tars and task copies are never cleaned up.

ShuxinLin and others added 21 commits September 25, 2026 16:12
A missing or unreachable database was reported the same way as a wrong key
(e.g. 'unknown asset_id', 'work order not found'). On not-found/failure paths
the iot, fmsr, wo and vibration servers now report that the database does not
exist in this environment and that retrying with other arguments won't help.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
…heck

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
fix: distinct error message when a CouchDB database is missing
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
ruff EXE001: the file has a shebang but shipped as mode 100644.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q7W9hWwuZU8Mkwo3ZCrR7m
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
Signed-off-by: Dhaval Patel <pateldha@us.ibm.com>
@ShuxinLin
ShuxinLin changed the base branch from main to aafeedback_changes September 28, 2026 14:15
@ShuxinLin

Copy link
Copy Markdown
Collaborator

What the code changes do

PR #558 does two things. It adds an opt-in way to run AssetOpsBench scenarios in parallel using Harbor, where each trial gets its own CouchDB. It also brings along a batch of fixes to the servers and changes to the TSFM model catalog. It's large: 57 files, +7,004 / −172, much of it uv.lock and model weights. It targets aafeedback_changes, not main.

1. Harbor integration (the main feature)

The reason for it: the current runner resets one shared CouchDB before every scenario, so running scenarios at once would wipe a live scenario's data. The agent would then get empty results and still produce a plausible-looking answer that gets scored. Harbor starts a separate Compose project per trial, so each scenario gets its own database.

  • Task generator (benchmarks/harbor/adapter/generate_tasks.py): turns each scenario in open.yaml into a Harbor task. That task contains the question, a thin Docker layer, the ground-truth files for the verifier, and an "oracle" solution that writes the correct answer.
  • Task template (benchmarks/harbor/template/):
    • task.toml sets the environment variables (COUCHDB_URL=http://couchdb:5984, …) and loads the scenario's data through a healthcheck running init_data.py <id>.
    • docker-compose.yaml adds a CouchDB sidecar that doesn't publish any host port.
    • test.sh runs the existing uv run evaluate, and to_reward.py converts its report into Harbor's {reward, passed}.
  • Base image (base-image/Dockerfile): Python 3.12 with the repo, uv sync --extra tsfm --group otel, and the Hugging Face models listed in models.txt pre-downloaded.
  • Stirrup as a Harbor agent (src/assetops_harbor/stirrup.py): runs uv run stirrup-agent inside the task container. It forwards only an allow-list of credential variables, uses Harbor's trial name as --run-id, and stops before anything is built if the router credentials are missing. After the run it converts the saved record into Harbor's trajectory format (ATIF) so token counts show up.
  • Code-track sandbox (overlays/code-sandbox.yaml): an optional Docker-in-Docker sidecar, so code the agent writes can't see CouchDB's hostname or credentials.
  • metric.py: per-category averages over a finished job.

2. Clearer errors when a CouchDB database is missing (from #554)

The iot, fmsr, wo and vibration servers now tell apart a missing database and a wrong key. Before, both produced errors like "unknown asset_id". Now a missing database returns "the data source does not exist … do not retry with other arguments", without naming the database. Tests cover this.

3. Stirrup runner passes the environment to MCP servers

src/agent/stirrup_agent/runner.py now sets "env": dict(os.environ). Without it, the MCP SDK passed on only HOME, PATH and a few other variables, so the servers silently fell back to localhost:5984. This fixes a real bug and applies outside Harbor too.

4. TSFM model catalog and weights

  • model_catalog.json is rewritten: the ttm_96_28 entry is removed and entries are added for several local TTM checkpoints, a fine-tuned energy model (ttm_energy_168_24), and many Hugging Face models (Chronos, Moirai, TimesFM, Toto, TSPulse, …).
  • About 41 MB of .safetensors files are committed under artifacts/ as plain git blobs, not Git LFS.
  • Nine helper scripts are added under benchmarks/harbor/scripts/ (catalog audit and fixes, preloading, smoke tests), plus src/servers/tsfm/tests/test_catalog_checkpoints.py.

5. Dependencies (pyproject.toml)

  • tsfm moves from a dependency group to an extra and grows a lot: transformers[torch]>=5.3, gluonts, lightning, toto-models, and others.
  • A new harbor extra is added, and sktime goes up to >=1.2.0.

- Layout lists the real paths (src/assetops_harbor, template/, overlays/)
  and marks datasets/ as generated.
- Run steps work from the repo root: correct Dockerfile path, adapter
  defaults instead of a --template pointing at a generated task, no
  PYTHONPATH=agent, no 'push'.
- Drop the pin to main@81265cb; the branch targets aafeedback_changes.
- Replace 'Still unverified' with the Docker-host results reported in #558.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
generate_tasks.py gains three options for corpora that do not ship in the
repo, such as AssetOpsBenchScenarioGeneration/scenarios_data:

- --runtime-image: the image each task builds FROM.
- --data-dir: where the corpus lives inside that image. The per-task
  layer copies the scenario there and the healthcheck runs init_data.py
  with SCENARIOS_DATA_DIR pointed at it; the agent keeps the repo copy,
  as scenario_suite_runner does.
- --skip-missing: warn and skip profile scenarios absent from the corpus
  (all.yaml lists wosr-62, which the corpus lacks).

corpus-image/Dockerfile bakes the corpus (2.1 GB, mostly shared/iot) into
one layer over the runtime image, so tasks share it instead of each
carrying it in their build context.

Also stop copying scenario 1's description and the wosr keyword into
every generated task.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Runs a scenario corpus through Harbor with the same profile, the same
stirrup-agent and the same Docker code sandbox as benchmarks/run.sh, but
with scenarios in parallel, each trial on its own CouchDB.

- Builds assetopsbench/runtime:corpus from -s, and the code sandbox tar
  for the per-trial Docker-in-Docker daemon.
- Regenerates the dataset from scratch (the generator never removes stale
  task folders), skipping profile entries the corpus lacks.
- One Harbor job per model under <leaderboard>/harbor-jobs; re-running
  resumes it, the equivalent of --skip-existing.
- Credentials reach the Harbor process through uv run --env-file only.

Uses [[ -z "${arr[*]+set}" ]] and ${arr[@]+...} so empty arrays work
under set -u in macOS's bash 3.2.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
corpus-image/ -> suite-image/, assetopsbench/runtime:corpus ->
assetopsbench/runtime:suite, /opt/corpus/scenarios_data ->
/opt/suite/scenarios_data, and the generated dataset
assetopsbench-corpus -> assetopsbench-suite, with matching wording in
run.sh, the generator's help and the README.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Completes 533bc5a, which only moved corpus-image/ to suite-image/:
assetopsbench/runtime:corpus -> assetopsbench/runtime:suite,
/opt/corpus/scenarios_data -> /opt/suite/scenarios_data, the dataset
assetopsbench-corpus -> assetopsbench-suite, and the wording in run.sh,
the generator's help and the README.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
base-image/Dockerfile does `COPY . .`, so building from the working
tree baked in whatever sat there. An untracked reports/results_table.csv
with a ground_truth column for 56 private scenarios (8 of mini, 41 of
lite) reached assetopsbench/runtime:dev that way, readable by any agent
running code in `main`, and publish-images.sh would have pushed it.

scripts/build-runtime-image.sh now builds from a one-commit clone of
HEAD, so only committed files enter the context. The clone keeps a real
.git, which StirrupAgent.get_version_command needs for `git rev-parse`;
a git worktree's .git pointer would not resolve inside the container.
publish-images.sh, run.sh and the docs use the script.

.dockerignore stays as the second guard, and the only one for a plain
`docker build .`: env files at any depth (a bare `.env` matched only the
context root), the open scenarios' answer files, which the verifier
never reads from the image, and the local folders reports/, logs/,
src/tmp/, .claude/ and the kdd_tutorial work.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
ShuxinLin and others added 5 commits September 30, 2026 13:50
Review of #587: the one-commit .git the clean clone kept still held
every committed blob, so `git show HEAD:.../groundtruth.txt` returned
the open scenarios' answers that .dockerignore had dropped from the
tree. Its reflog also recorded the builder's name and email, and its
index and reflog changed on every clone, so `COPY . .` never hit the
cache and each build re-downloaded the models.

build-runtime-image.sh now builds from `git archive HEAD` and passes
the commit as the AOB_COMMIT build arg. .dockerignore drops .git. A
`commit` stage in base-image/Dockerfile writes /opt/aob/.aob-commit,
which StirrupAgent.get_version_command now reads, and fails any build
without AOB_COMMIT, so a plain `docker build .` of the working tree
stops instead of baking in untracked files. The ARG lives only in that
stage, so a new commit invalidates the final layer and nothing else.

The script also keeps its default tag and --load when given other
flags (e.g. --no-cache), warns about untracked files it leaves out,
and checks for docker buildx up front.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Review of #586:
- generate_tasks.py copies into each task only manifest.json and the
  files it names inside the scenario folder. Skipping the verifier's
  input names let any other answer file through (a .bak, a new scorer
  input, GroundTruth.txt on a case-insensitive disk).
- overlays/private-data.yaml uses the long bind syntax with
  create_host_path: false. A mistyped or relative AOB_PRIVATE_DIR used
  to make Docker create an empty shared/ and load empty collections
  with the healthcheck passing; it now fails the trial at start with
  "bind source path does not exist". A path with ':' also works now.
- run.sh reports a failed resume instead of hiding it behind `|| true`.
  Harbor refuses to resume a job whose tasks or overlays no longer
  match its lock.json, which is every job from the old suite-image
  run.sh.
- The overlay no longer claims suite edits need no rebuild: a changed
  manifest.json needs the tasks regenerated. The README asks for an
  absolute AOB_PRIVATE_DIR.

Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
fix(harbor): run.sh mounts only the private suite's shared/
fix(harbor): build the runtime image from a clean clone of HEAD
Every task image builds FROM the runtime image, and the only way to choose it
was the Dockerfile's ARG default, so using a published image meant retagging it
as the local assetopsbench/runtime:dev. The template's docker-compose.yaml now
passes AOB_RUNTIME_IMAGE to that ARG as a build arg. Harbor runs `docker
compose build` with the shell's environment and every compose file, so
`AOB_RUNTIME_IMAGE=<image> harbor run ...` builds FROM that image, with no
task regeneration; unset, it keeps the local tag.

run.sh takes it as -r (or AOB_RUNTIME_IMAGE). A registry reference is pulled
first, because a local copy satisfies FROM and the build would otherwise run on
a stale one; a bare name must already exist locally, as before. It exports the
variable so `harbor jobs resume` sees it too, and prints the image id it uses.

The image changes only the base. Each task still adds its own manifest.json,
and shared/ still comes from overlays/private-data.yaml.

The template change alters every generated task, so Harbor will not resume a
job started before it; run.sh already says so and suggests moving it aside.

Verified with Harbor's oracle agent, AOB_RUNTIME_IMAGE set to a labelled copy
of runtime:dev: the 3 open tasks, and tsfm-1019 and fmsr-913 from the private
suite with the private-data overlay. All 5 trials scored 1.000 with no errors.
A nonexistent AOB_RUNTIME_IMAGE failed the build pulling exactly that name. The
tsfm-1019 task image built on the labelled base carries the label. With the
overlay's mount it sees scenario_1019/manifest.json, and a read-only shared/
where all five manifest paths resolve. run.sh's image step pulls a quay.io
reference and rejects a missing local tag. 22 assetops_harbor tests pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
ShuxinLin and others added 4 commits September 30, 2026 15:29
build-runtime-image.sh tagged only assetopsbench/runtime:dev, which every build
overwrites, so a run could not name the build it used after a rebuild. With no
-t given it now also tags assetopsbench/runtime:<short commit>. That tag does not
move, so `run.sh -r assetopsbench/runtime:<commit>` (or AOB_RUNTIME_IMAGE) pins a
run to one build while :dev stays the default the tasks build FROM.

publish-images.sh adds <namespace>/runtime:<commit> beside :TAG and :latest.
The commit is HEAD's, which is exactly what the image holds, because the build
comes from `git archive HEAD` whatever the working tree contains.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Review fixes for the runtime-image selection.

run.sh:
- Resolves the image as -r, then the shell's AOB_RUNTIME_IMAGE, then ENV_FILE's,
  then the local default. It used to take only the shell's value and export a
  default over it, so an AOB_RUNTIME_IMAGE in .env was ignored; `uv run
  --env-file` never overrides a variable the shell already has.
- Pulls whenever the image is not local, or its local copy came from that same
  repository. The old rule pulled only references naming a registry host, so a
  Docker Hub image such as publish-images.sh's own <ns>/runtime:latest was never
  refreshed. A local build (no RepoDigest for its name) is never pulled over.
- Pins the resolved image for the whole run under a tag private to the process
  (aob-runtime-pin:<id>-<pid>, removed on exit) and exports that. Every trial
  resolves FROM when it builds, so exporting the movable :dev let a rebuild or
  another run's pull switch later trials' base mid-run. FROM cannot name a bare
  image id, which is why this uses a tag.
- Records the image beside each job (<job>.runtime-image) and refuses to resume
  a job on a different one. Harbor's resume lock hashes the task files, not the
  base they build FROM, so a resume with another -r mixed two images in one job.
- Exits non-zero when a model's job could not start or resume.
- Runs the image step after the cheap input checks, and strips a sha256: prefix
  only if present when printing the id.

build-runtime-image.sh: treats -tNAME and -t=NAME as naming a tag, so they no
longer also get the default :dev and :<commit> tags. Its comment no longer
claims the commit tag never moves (rebuilding the same commit moves it), and it
says how to remove old commit tags, which now keep their images alive.

publish-images.sh: the NOTE points users at the commit tag it just pushed, with
:latest as the moving option.

Docs: QUICKSTART says tasks generated before the build arg existed ignore the
variable and need regenerating once, that the variable must be set per shell or
in .env, and its pull-error entry now covers an unset AOB_RUNTIME_IMAGE. README
describes the resolution order, pin, per-job record and exit status.

Verified end to end with run.sh, a one-scenario profile and a fake model id, so
trials fail fast without calling a provider:
- The first run printed the resolved id. The pin tag existed during the run and
  was gone after it, the job's .runtime-image record was written, and the
  trial built, ran and was scored (reward 0).
- A second run on the same image resumed the job, dropped the crashed trial and
  reran it, and exited 0.
- A third run with -r set to a different image refused to resume, named the
  job's original image, and exited 1.
- Helper checks: image_repo strips tags and digests; runtime:dev and :3bb4f12
  count as local builds; quay.io/... and couchdb:3.5 count as registry images.
  -r and the shell beat .env, which beats the default.
- `uv run pytest src/ -k "not integration"`: 714 passed, 19 failed. The base
  branch has the same 19 failures, in the scorer and trace-exporter tests and in
  the iot invalid-site tests, which need a CouchDB or no .env.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
feat(harbor): pick the runtime image per run; tag builds by commit
Text the later Harbor changes left behind, found reviewing #558. No behaviour
changes except the wording of one error message.

- stirrup.py module docstring: drop the first-draft run line (PYTHONPATH=agent,
  a watsonx model, --n-concurrent 16) and the claim that MCP servers get the
  environment through mcphub; the Stirrup runner passes it through
  mcp_server_env. The docker backend is no longer "redundant": the
  code-sandbox overlay provides a daemon and keeps code off CouchDB's hostname
  and the credentials, which local does not.
- StirrupAgent's docker-backend error pointed at mounting a Docker socket; it
  now names the code-sandbox overlay. It keeps "no Docker daemon", which a test
  matches.
- _load_dotenv: SETTING_ENV_VARS (FMSR_MODEL_ID) reach the container too.
- stirrup_agent/runner.py: FMSR has no "standalone watsonx default" since #583.
- README step 4 and CODE-SANDBOX.md's run.sh example used a watsonx model,
  which FMSR now rejects, so those runs had no generate_failure_modes. They use
  litellm_proxy/azure/gpt-5.6-sol, and step 4 says why.
- README: the local backend no longer sits next to the open scenarios' ground
  truth (#587 removed it from the image); say what it does sit next to. The
  explicit-env section names mcp_server_env and plan-execute.
- base-image/Dockerfile: check_models_list.sh does not exist, so say nothing
  checks models.txt against the catalog; the build line matches the script's
  default tags.
- generate_tasks.py, publish-images.sh: "corpus" is now "suite"/"scenario
  data"; publish-images.sh's example namespace is quay.io/assetopsbench.

The task template's comments are stale in the same way but are left alone:
editing them changes every task's content hash, and Harbor then refuses to
resume jobs started on the current tasks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
ShuxinLin and others added 2 commits September 30, 2026 19:15
docs(harbor): correct stale docs, docstrings and comments
Four run.sh problems that hold even with a stable runtime image.

- A job was named after the model alone, and a resume runs the job's saved
  config. So `-m "X high" -m "X max"` resumed the high job for max and ran no
  max trials, and another -p in the same leaderboard directory hit Harbor's
  lock and could not resume. Jobs are now
  stirrup_agent__<profile>__<model>[__<effort>].
- Tasks were regenerated on every run, and Harbor refuses to resume a job
  whose tasks differ from its lock. Any change to the template (comments
  included), the suite's scenario files or the generator made every job
  unresumable. Each new job now gets its own copy of the tasks, generated once
  when it starts and reused on resume. The copies hold the answers, so they go
  to the repo's gitignored benchmarks/harbor/datasets/jobs/, keyed by the job's
  path, not into the leaderboard directory.
- One shared task folder meant a second run.sh deleted tasks that a running
  job was still reading. Per-job copies remove the shared folder.
- `harbor run ... || true` hid the only failures harbor run reports: it exits 0
  when trials fail and non-zero when the job itself cannot run (a rejected
  config, or StirrupAgent's credential check aborting the job as its first
  trial starts). That, and a model skipped for an unreachable router, now set a
  non-zero exit. A resume whose task copy is gone says so.

Jobs started under the old names are not resumed; a rerun starts new jobs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
ShuxinLin and others added 12 commits September 30, 2026 19:34
The rest of the run.sh problems found reviewing #558.

- Resume reran only NonZeroAgentExitCodeError. Harbor matches exact class
  names, so its subclasses (rate limits, network, auth), Ctrl-C, and
  environment and verifier failures stayed failed. run.sh now lists every type
  that is not the model's own work; timeouts, context and output limits and
  safety refusals stay as results. test_run_sh.py checks the list against the
  installed Harbor and fails when a new agent error subclass is unplaced.
- The code sandbox tar was built once and reused forever. It is now built each
  run (cached), saved once per image id under AOB_CODE_TAR_DIR (default
  ~/.cache/assetopsbench), and recorded per job; a resume loads the job's own
  tar, so a job never switches code image and a rebuild never rewrites a tar
  in use.
- The router check only probed the base URL. It now rejects a key the router
  refuses (401/403 from /models), also checks FMSR_MODEL_ID's router, and skips
  a model with no router prefix unless FMSR_MODEL_ID names one, since FMSR
  would reject it and fmsr scenarios would run without generate_failure_modes.
- ENV_FILE was not the only credential source: StirrupAgent also loaded the
  repo's .env and filled its gaps. StirrupAgent now reads AOB_ENV_FILE instead
  of searching when it is set, and run.sh sets it.
- Two run.sh processes could work on one job. A per-job lock (mkdir, with the
  owner's PID so a dead owner's lock is taken over) makes the second skip it.
- Relative -s, -l, -p, ENV_FILE and AOB_CODE_TAR_DIR resolved against the repo
  root. They now resolve against the caller's directory; the defaults are still
  the repo's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Found reviewing the per-job task copies. A resume reuses the job's own
manifests but mounts shared/ from the current -s, so a resume with another
suite would pair one suite's manifests with another suite's data. Before the
copies, Harbor's task lock refused that resume; now run.sh records the suite
beside the job (<job>.suite) and refuses it, as it does for another runtime
image.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
fix(harbor): harden run.sh jobs, retries, checks and inputs
test_the_sdk_default_would_drop_couchdb_url set COUCHDB_URL in os.environ and
popped it afterwards, deleting any value the developer already had, so later
tests in the session ran without it. It now uses monkeypatch.setenv, which
restores the original.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
…ments

- pyproject.toml listed python-dotenv and granite-tsfm twice in dependencies,
  and numba's line in the dev group had trailing whitespace. uv.lock is
  unchanged (`uv lock --check` passes).
- .gitignore listed benchmarks/harbor/datasets/ twice, and its comment named
  harbor/adapter/generate_tasks.py without the benchmarks/ prefix.
- template/task.toml said mcphub carries COUCHDB_URL to the MCP servers "with
  no code change"; the runners pass it through mcp_server_env, because the MCP
  SDK's default environment drops it.
- template/environment/Dockerfile described the per-task COPY as how "the
  restricted corpus" layers its scenarios; it now says what it copies (the
  manifest and the files it names, never answers) and that an external suite
  lands in /opt/suite/scenarios_data.

Template edits change generated tasks' hashes. run.sh jobs are unaffected,
since each job keeps its own copy of its tasks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The fine-tuned energy model sat in artifacts/output/tuned_models/, the root
artifacts/README.md reserves for checkpoints an agent writes during a trial and
says is never committed and ignored by git and Docker. But the model catalog
serves it as a shipped model, so the ignore rules the README describes were
never added: adding them would have dropped the model and broken its card.

- Move the checkpoint (config.json, meta.json, model.safetensors) to
  artifacts/tsfm_models/ttm_energy_168_24 and repoint its card's three path
  fields (apply_catalog_fixes.py --move-energy --write).
- .gitignore and .dockerignore now exclude artifacts/output/, as the README
  says.
- generate_model_catalog.py no longer scans artifacts/output/tuned_models:
  agent output must not become catalog cards.
- apply_catalog_fixes.py: the docstring says the repo has moved and the flag is
  for catalogs that have not, such as a private suite's; the reminder prints
  only when a catalog still points at the old path.
- artifacts/README.md lists what tsfm_models/ holds, corrects the checkpoint
  sizes, and says save_to takes any path rather than implying agents write to
  output/ by themselves.

test_catalog_checkpoints.py (12 passed) loads, fits and forecasts the model at
the new path, and preload_models.py --check resolves it.

A catalog outside the repo that still names the old path fails for this model
once the runtime image is rebuilt. Update it with
  apply_catalog_fixes.py --catalog <path> --move-energy --write

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Cut the Harbor docs and code comments to what a reader needs now, and
remove content that no longer holds.

- README: drop the dated "verified" run logs, the smoke-test commands that
  repeated other sections, the CLI-history notes, the `git rm --cached jobs`
  step and the code-track section duplicated in CODE-SANDBOX.md; summarise
  run.sh as an options table and short steps.
- QUICKSTART: drop the code-track section (now a link), the troubleshooting
  entries for fixes already in the code (MCP env, otel group) and a
  registry image name that is not published.
- CODE-SANDBOX: replace the Harbor environment inventory with one line and
  drop the claim that main holds groundtruth.txt, which the image excludes.
- Overlays, template, Dockerfile, run.sh, image scripts: shorten comments.
  code-sandbox.yaml said the registry route sets "both variables" but
  showed one.
- stirrup.py, generate_tasks.py, metric.py, catalog scripts: shorten
  docstrings; drop "on main" / "yet" wording, the reference to a design doc
  that is not in the repo, and history about past export failures.
- Fix README paths in pyproject.toml, assetops_harbor/__init__.py and the
  generated dataset.toml header (harbor/ -> benchmarks/harbor/).
- stirrup_agent/runner.py: the MCP env comment no longer says mcphub covers
  the other runners; they all use mcp_server_env.

No behaviour change. Generated tasks still validate against Harbor's
TaskConfig; the unit suite shows the same 19 failures as before the change
(evaluation, observability, iot), all outside these files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
code-image-loader echoes $STIRRUP_CODE_IMAGE when no AOB_CODE_TAR is given,
but only `main` had that variable, so the message read "the daemon will pull
 instead". Give the loader the same STIRRUP_CODE_IMAGE as `main`.

Checked by rendering the overlay with docker compose config and running the
loader's no-tar branch: it now prints assetops-code:dev, or AOB_CODE_IMAGE
when set.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
`uv sync` installs the dev group by default, so the runtime image carried
pytest and the Jupyter/IPython stack, which nothing in a trial imports. They
were 63 of the image's 291 downloaded packages; a timed-out download of one of
them (debugpy, via ipykernel) failed a build.

Both syncs now pass --no-dev. numba, pyod, anyio and the OpenTelemetry
packages stay: other dependencies require them.

UV_NO_SYNC=1 is set once the environment is final. Every process in a trial
starts through `uv run`, which otherwise syncs the default groups first and
would reinstall the dev group in every trial; it would also do so in the
preload step of this build. Checked with uv 0.12.21, the version the image
installs: after `sync --no-dev`, a plain `uv run` installs the dev group and
`UV_NO_SYNC=1 uv run` does not.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
The Hub model download (3.9 GB, about 6 minutes) ran after `COPY . .`, so
every commit downloaded every model again, and each commit tag kept its
own 4 GB copy of the weights.

- Download the models in their own stage, which depends only on
  models.txt and preload_models.py (its --from-list path needs nothing
  but huggingface_hub, pinned to uv.lock's 1.33.0). The main stage copies
  /opt/hf below the dependency layer, since uv.lock changes more often
  than models.txt.
- Give both `uv sync` calls a BuildKit cache mount for the uv cache. The
  cache used to stay in the image (2.2 GB, beside a 2.0 GB venv); now a
  lock change reuses downloaded wheels.

Measured with the docker driver: a source edit rebuilds in 2 s (was a
full model download), a pyproject.toml edit in 24 s, and a cold build in
about 7 minutes. The image shrinks from 8.75 GB to 6.53 GB, and a new
commit adds only its ~60 MB of repo layers. COPY --link was tried for
the weights and did not avoid the copy on this driver, so it is not used.

Checked on the built image: no uv cache, no dev group, all 20 models
resolve offline, imports work, and Harbor's oracle scores 1.000 on the
3 open tasks with 0 exceptions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Replace "Running it", "Running a full scenario suite" and "Running a
private profile by hand" with three sections:

- Running a private profile end to end: build the image, pin the run to
  its commit tag, run.sh on mini, then harbor view and metric.py, followed
  by run.sh's options, steps, resume and locking (unchanged).
- Running a public profile by hand: the open profile with plain Harbor
  commands, including saving the code sandbox tar, the oracle check and
  the Docker sandbox run; tools-only, results and resume as notes.
- Running a private profile by hand: the same with --scenario-root, its
  own --output-dir and the private-data overlay, then the generator flags
  and "How the private data gets in" (unchanged).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
Harbor names a local dataset after its folder and builds
agent__model__dataset keys from it. Its CLI splits those keys on "__"
to print the results table, and Job splits them to pick the dataset's
metrics. run.sh named each job's tasks <job_name>-<checksum>, and job
names contain "__", so a run.sh job failed once every trial had finished:

  too many values to unpack (expected 2)
  Harbor could not run <job>; see the error above.

result.json and the trials were written, but with no dataset metrics,
and run.sh reported the job as failed.

Replace "__" with "--" in the folder name. A new test rebuilds the name
from run.sh's own line and checks it against Harbor's key format.
Reproduced with the oracle on the open tasks: a folder with "__" exits 1
with that error and empty metrics; with "--" it exits 0 and records
them.

A job started before this change keeps its old task folder, which
run.sh no longer finds, so it refuses to resume it; move that job aside
to rerun it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Shuxin Lin <linshuhsin@gmail.com>
@ShuxinLin
ShuxinLin merged commit 9750df7 into aafeedback_changes Oct 1, 2026
6 checks passed
@ShuxinLin
ShuxinLin deleted the feature/harbor-integration branch October 1, 2026 00:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants