Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
72 commits
Select commit Hold shift + click to select a range
6bbaba3
fix: distinct error message when a CouchDB database is missing
ShuxinLin Sep 25, 2026
f230812
fix: do not expose database names in missing-database errors
ShuxinLin Sep 25, 2026
4224cff
fix: cover find_assets_by_sensors and wo delete in missing-database c…
ShuxinLin Sep 25, 2026
81265cb
Merge pull request #554 from IBM/feat/couchdb-error-classification
DhavalRepo18 Sep 25, 2026
b91549f
feat(harbor): run AssetOpsBench scenarios as Harbor tasks
DhavalRepo18 Sep 26, 2026
96197eb
adding files
DhavalRepo18 Sep 26, 2026
6948500
revised code
DhavalRepo18 Sep 27, 2026
13d5ad3
Updated
DhavalRepo18 Sep 27, 2026
88f1bb8
revised code
DhavalRepo18 Sep 27, 2026
ae4d69c
adding more models and code for tsfm
DhavalRepo18 Sep 27, 2026
27ac496
revised models
DhavalRepo18 Sep 27, 2026
f53e383
chore(harbor): mark preload_models.py executable
DhavalRepo18 Sep 27, 2026
8987013
revised code
DhavalRepo18 Sep 27, 2026
abff786
revised git ignore and output
DhavalRepo18 Sep 27, 2026
672ec56
revised docker ignore
DhavalRepo18 Sep 27, 2026
69ca4af
revised code
DhavalRepo18 Sep 27, 2026
5421771
revised code
DhavalRepo18 Sep 27, 2026
e119b1f
revised catalog generator
DhavalRepo18 Sep 27, 2026
bc26dde
revised code
DhavalRepo18 Sep 28, 2026
8edc871
revised code
DhavalRepo18 Sep 28, 2026
e32abb7
revised code
DhavalRepo18 Sep 28, 2026
0c45322
docs(harbor): fix stale paths and run steps in README
ShuxinLin Sep 28, 2026
bba3fa6
feat(harbor): generate tasks against an external scenario corpus
ShuxinLin Sep 28, 2026
2dd44bc
style(harbor): ruff format generate_tasks.py
ShuxinLin Sep 28, 2026
7be1922
feat(harbor): run.sh, the Harbor counterpart of benchmarks/run.sh
ShuxinLin Sep 28, 2026
533bc5a
refactor(harbor): rename corpus to suite
ShuxinLin Sep 28, 2026
c453ad7
refactor(harbor): rename corpus to suite in contents
ShuxinLin Sep 28, 2026
836e9d3
fix(harbor): check the model router before a job; retry crashed trials
ShuxinLin Sep 28, 2026
958e3d8
docs(harbor): document generating mini/lite/all profiles
ShuxinLin Sep 29, 2026
1b9306f
feat(harbor): load .env for StirrupAgent credentials
ShuxinLin Sep 29, 2026
6eaa1d8
docs(harbor): correct the shared-data gap for non-open profiles
ShuxinLin Sep 29, 2026
1e1579f
Merge pull request #575 from IBM/feat/harbor-load-dotenv
ShuxinLin Sep 29, 2026
e45f338
Merge feature/harbor-integration into feat/harbor-suite-run
ShuxinLin Sep 29, 2026
cfecba9
feat(harbor): mount a private suite with AOB_PRIVATE_DIR
ShuxinLin Sep 29, 2026
1832899
refactor(harbor): infer the suite data dir from --scenario-root
ShuxinLin Sep 29, 2026
b4b01f3
feat(harbor): run mini/lite/all by mounting the private suite
ShuxinLin Sep 29, 2026
3dedb12
Merge pull request #577 from IBM/feat/harbor-private-data
ShuxinLin Sep 29, 2026
8459a7b
docs(harbor): align smoke tests with the PR comment; record agent runs
ShuxinLin Sep 29, 2026
460543f
revised code
DhavalRepo18 Sep 29, 2026
a30f036
fix(fmsr): take the FMSR model from FMSR_MODEL_ID, with no built-in d…
ShuxinLin Sep 30, 2026
ffa4c38
fix(fmsr): accept only the llm.routers prefixes, dropping watsonx
ShuxinLin Sep 30, 2026
2a1b421
fix(harbor): forward FMSR_MODEL_ID and check its router up front
ShuxinLin Sep 30, 2026
b77f60f
fix(fmsr): gate the integration test on the database as well as the LLM
ShuxinLin Sep 30, 2026
534b010
fix(fmsr): generate failure modes without an initialised catalog
ShuxinLin Sep 30, 2026
bca6aaa
Merge pull request #583 from IBM/fmsr-model-pin
ShuxinLin Sep 30, 2026
b74b621
Merge branch 'feature/harbor-integration' into harbor-private-mount-d…
ShuxinLin Sep 30, 2026
bc8886b
Merge pull request #581 from IBM/harbor-private-mount-docker-md
ShuxinLin Sep 30, 2026
4ee86af
fix(harbor): run.sh mounts only the private suite's shared/
ShuxinLin Sep 30, 2026
b7a7d84
fix(harbor): build the runtime image from a clean clone of HEAD
ShuxinLin Sep 30, 2026
47a806e
fix(harbor): ship no .git in the runtime image
ShuxinLin Sep 30, 2026
7f66b87
fix(harbor): allowlist task inputs and fail on a missing suite path
ShuxinLin Sep 30, 2026
ec62c34
Merge pull request #586 from IBM/fix/harbor-run-private-mount
ShuxinLin Sep 30, 2026
326598d
Merge pull request #587 from IBM/fix/harbor-clean-runtime-image
ShuxinLin Sep 30, 2026
e91aa07
feat(harbor): pick the runtime image at run time with AOB_RUNTIME_IMAGE
ShuxinLin Sep 30, 2026
3bb4f12
feat(harbor): tag each runtime build with its commit
ShuxinLin Sep 30, 2026
1ad4eea
fix(harbor): pin the runtime image for a run and guard resumes
ShuxinLin Sep 30, 2026
401197f
Merge pull request #588 from IBM/feat/harbor-runtime-image-param
ShuxinLin Sep 30, 2026
f2d1310
docs(harbor): correct stale docs, docstrings and comments
ShuxinLin Sep 30, 2026
57dff2f
Merge pull request #590 from IBM/fix/harbor-stale-docs
ShuxinLin Sep 30, 2026
2494f1e
fix(harbor): give each run.sh job its own name, tasks and exit status
ShuxinLin Sep 30, 2026
66eb514
fix(harbor): run.sh retries, code tar, key checks, env file, lock, paths
ShuxinLin Sep 30, 2026
d9eb32d
fix(harbor): refuse to resume a run.sh job on another suite
ShuxinLin Sep 30, 2026
cb058ab
Merge pull request #591 from IBM/fix/harbor-run-job-identity
ShuxinLin Sep 30, 2026
1f68511
fix(test): keep a pre-set COUCHDB_URL in the MCP env test
ShuxinLin Sep 30, 2026
038d900
chore(harbor): drop duplicate deps and ignore rules; fix template com…
ShuxinLin Sep 30, 2026
8676a3a
fix(tsfm): ship ttm_energy_168_24 from artifacts/tsfm_models
ShuxinLin Sep 30, 2026
2daca75
docs(harbor): trim docs and comments, drop stale content
ShuxinLin Sep 30, 2026
581a538
fix(harbor): name the code image in the loader's no-tar message
ShuxinLin Sep 30, 2026
2970b5d
fix(harbor): leave the dev group out of the runtime image
ShuxinLin Oct 1, 2026
102646e
perf(harbor): rebuild the runtime image in seconds after a source edit
ShuxinLin Oct 1, 2026
ba120aa
docs(harbor): organise the README by how a run is driven
ShuxinLin Oct 1, 2026
8f95ad2
fix(harbor): keep "__" out of run.sh's per-job task folder
ShuxinLin Oct 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 46 additions & 0 deletions .dockerignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# Filters the runtime image's build context, which build-runtime-image.sh makes
# from `git archive HEAD`: committed files not listed here are baked into the
# image by `COPY . .` in benchmarks/harbor/base-image/Dockerfile.

# Credentials. A bare name matches only at the context root, hence **/.
**/.env
**/.env.*
!**/.env.public

# Host virtualenv: the image builds its own with `uv sync --frozen`.
.venv/

# Agent-written checkpoints. artifacts/tsfm_models/ ships in the image.
artifacts/output/

# Harbor output and generated tasks.
jobs/
benchmarks/harbor/datasets/

# Caches and local noise.
**/__pycache__/
*.py[cod]
.pytest_cache/
.ruff_cache/
.mypy_cache/
traces/
.DS_Store

# Answers: the image is the agent's container. The verifier reads the copies
# Harbor uploads to /tests after the agent phase.
src/couchdb/scenarios_data/**/groundtruth*
src/couchdb/scenarios_data/**/rubric.json
src/couchdb/scenarios_data/**/reference_answer.json
src/couchdb/scenarios_data/**/scenario_meta.json

# Local work that is not part of the benchmark.
reports/
logs/
src/tmp/
.claude/
notebook/kdd_tutorial/
artifacts/kdd_tutorial/

# Git data holds every committed blob, answers included. The commit reaches the
# image as /opt/aob/.aob-commit instead.
.git
7 changes: 7 additions & 0 deletions .env.public
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,13 @@ LITELLM_BASE_URL=
TOKENROUTER_API_KEY=
TOKENROUTER_BASE_URL=https://api.tokenrouter.com/v1

# ── FMSR MCP server (generate_failure_modes) ─────────────────────────────────
# Model used inside the FMSR server. Agent runners always pin it: this value if
# set, otherwise the agent's --model-id. Set it to fix one model across agents.
# There is no built-in default, so a standalone server with this unset reports
# the generate_* tools as unavailable rather than picking a provider for you.
FMSR_MODEL_ID=

# ── Benchmark code execution
SCENARIO_DIR=
LEADERBOARD_DIR=
Expand Down
10 changes: 9 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -199,8 +199,16 @@ benchmark/cods_track2/.env.local
CLAUDE.md
mcp/couchdb/sample_data/bulk_docs.json
.env
mcp/servers/tsfm/artifacts/tsfm_models/
# Agent-written checkpoints. artifacts/tsfm_models/ is tracked.
artifacts/output/
src/tmp/
trainer_output

# Observability artifacts (OTLP-JSON traces + per-run trajectory JSON).
traces/

# Generated Harbor tasks; they hold the scenarios' answers.
benchmarks/harbor/datasets/

# Harbor run output.
jobs/
1 change: 1 addition & 0 deletions INSTRUCTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -97,6 +97,7 @@ See [MCP Servers](#mcp-servers) for available tools and [docs/mcp-servers.md](do
| `WO_DBNAME` | `workorder` | Work order database name |
| `FAILURE_CODE_DBNAME` | `failure_code` | FCC failure-code database name |
| `FAILURE_MODE_DBNAME` | `failure_mode` | FMSR failure-mode database name |
| `FMSR_MODEL_ID` | agent `--model-id` | LLM for FMSR `generate_failure_modes`; runners always pin it (explicit value, else the agent model). No built-in default: unset and standalone, the `generate_*` tools report `LLM unavailable` |
| `VIBRATION_DBNAME` | `vibration` | Vibration sensor database name |
| `CATALOG_DBNAME` | `catalog` | Shared sensor/asset/failure-mode catalog database name |
| `MODEL_CATALOG_DBNAME` | `model_catalog` | TSFM model catalog database name |
Expand Down
66 changes: 66 additions & 0 deletions artifacts/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
# artifacts/

Model weights that do not come from the HuggingFace Hub.

```
artifacts/
tsfm_models/ read-only. Committed here, baked into the image.
output/
tuned_models/ writable. For checkpoints an agent fine-tunes during a trial. Never committed.
```

## tsfm_models/ - shipped checkpoints

Four TinyTimeMixer checkpoints (`ttm_512_96`, `ttm_512_720`, `ttm_1536_96`,
`ttm_1536_720`) and `ttm_energy_168_24`, a TTM fine-tuned for energy load
forecasting. A fine-tuned checkpoint that ships with the repo belongs here, not
in `output/`. An optional `meta.json` beside the weights records what they
cannot (domain, lineage, training data) for
`benchmarks/harbor/scripts/generate_model_catalog.py`.

Each entry is a `save_pretrained` directory (`config.json` plus weights) that a
catalog card points at:

```json
"source": "local_artifact",
"hf_repo": null,
"artifact_path": "artifacts/tsfm_models/ttm_512_96",
"model_checkpoint": "artifacts/tsfm_models/ttm_512_96",
"params": { "model_path": "artifacts/tsfm_models/ttm_512_96" }
```

Keep the location fields in step. Only `params.model_path` is read at load
time, but the agent copies this shape when it registers its own cards.

The path resolves against the working directory: the repo root locally,
`/opt/aob` in the container. A missing directory is treated as a Hub repo id,
and the resulting error does not mention the directory, so check instead:

```bash
uv run python benchmarks/harbor/scripts/preload_models.py --check
```

That resolves every active card, Hub and local alike, and exits non-zero when
anything would fail at fit time.

### What belongs here

Public, redistributable weights only: check both the base model's licence and
what the checkpoint was fine-tuned on. A model tuned on internal or customer
data belongs in the private set, whatever the base licence says.

Plain git is fine at these sizes (0.15 MB to 20 MB); use Git LFS if a file
approaches 50 MB.

## output/tuned_models/ - agent output

For checkpoints an agent fine-tunes during a trial: `run_recipe`'s `save_to`
names the directory (any path works), and `register_finetuned` points a new
card at it. It lives in the trial container and is discarded with it;
`.gitignore` and `.dockerignore` both exclude it.

Cards for these models carry `created_by: "agent.tsfm.finetune"`, which is what
`preload_models.py --check` uses to skip them.

If you mount this root from the host, give each trial its own subdirectory, or
parallel trials will overwrite each other.
64 changes: 64 additions & 0 deletions artifacts/tsfm_models/ttm_1536_720/config.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
{
"adaptive_patching_levels": 3,
"architectures": [
"TinyTimeMixerForPrediction"
],
"categorical_vocab_size_list": null,
"context_length": 1536,
"d_model": 384,
"d_model_scale": 3,
"decoder_adaptive_patching_levels": 0,
"decoder_d_model": 256,
"decoder_d_model_scale": 2,
"decoder_mode": "common_channel",
"decoder_num_layers": 2,
"decoder_raw_residual": false,
"distribution_output": "student_t",
"dropout": 0.4,
"dtype": "float32",
"enable_forecast_channel_mixing": false,
"exogenous_channel_indices": null,
"expansion_factor": 2,
"fcm_context_length": 1,
"fcm_gated_attn": true,
"fcm_mix_layers": 3,
"fcm_prepend_past": true,
"fcm_prepend_past_offset": null,
"fcm_use_mixer": true,
"frequency_token_vocab_size": 8,
"gated_attn": true,
"head_dropout": 0.4,
"huber_delta": 1,
"init_embed": "pytorch",
"init_linear": "pytorch",
"init_processing": true,
"init_std": 0.02,
"loss": "mse",
"mask_value": 0,
"masked_context_length": null,
"mode": "common_channel",
"model_type": "tinytimemixer",
"norm_eps": 1e-05,
"norm_mlp": "LayerNorm",
"num_input_channels": 1,
"num_layers": 2,
"num_parallel_samples": 100,
"num_patches": 12,
"patch_last": true,
"patch_length": 128,
"patch_stride": 128,
"positional_encoding_type": "sincos",
"post_init": false,
"prediction_channel_indices": null,
"prediction_filter_length": null,
"prediction_length": 720,
"quantile": 0.5,
"resolution_prefix_tuning": false,
"scaling": "std",
"self_attn": false,
"self_attn_heads": 1,
"stride_ratio": 1,
"transformers_version": "4.57.6",
"use_decoder": true,
"use_positional_encoding": false
}
Binary file not shown.
64 changes: 64 additions & 0 deletions artifacts/tsfm_models/ttm_1536_96/config.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
{
"adaptive_patching_levels": 3,
"architectures": [
"TinyTimeMixerForPrediction"
],
"categorical_vocab_size_list": null,
"context_length": 1536,
"d_model": 384,
"d_model_scale": 3,
"decoder_adaptive_patching_levels": 0,
"decoder_d_model": 256,
"decoder_d_model_scale": 2,
"decoder_mode": "common_channel",
"decoder_num_layers": 2,
"decoder_raw_residual": false,
"distribution_output": "student_t",
"dropout": 0.4,
"dtype": "float32",
"enable_forecast_channel_mixing": false,
"exogenous_channel_indices": null,
"expansion_factor": 2,
"fcm_context_length": 1,
"fcm_gated_attn": true,
"fcm_mix_layers": 3,
"fcm_prepend_past": true,
"fcm_prepend_past_offset": null,
"fcm_use_mixer": true,
"frequency_token_vocab_size": 8,
"gated_attn": true,
"head_dropout": 0.4,
"huber_delta": 1,
"init_embed": "pytorch",
"init_linear": "pytorch",
"init_processing": true,
"init_std": 0.02,
"loss": "mse",
"mask_value": 0,
"masked_context_length": null,
"mode": "common_channel",
"model_type": "tinytimemixer",
"norm_eps": 1e-05,
"norm_mlp": "LayerNorm",
"num_input_channels": 1,
"num_layers": 2,
"num_parallel_samples": 100,
"num_patches": 12,
"patch_last": true,
"patch_length": 128,
"patch_stride": 128,
"positional_encoding_type": "sincos",
"post_init": false,
"prediction_channel_indices": null,
"prediction_filter_length": null,
"prediction_length": 96,
"quantile": 0.5,
"resolution_prefix_tuning": false,
"scaling": "std",
"self_attn": false,
"self_attn_heads": 1,
"stride_ratio": 1,
"transformers_version": "4.57.6",
"use_decoder": true,
"use_positional_encoding": false
}
Binary file not shown.
64 changes: 64 additions & 0 deletions artifacts/tsfm_models/ttm_512_720/config.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
{
"adaptive_patching_levels": 3,
"architectures": [
"TinyTimeMixerForPrediction"
],
"categorical_vocab_size_list": null,
"context_length": 512,
"d_model": 192,
"d_model_scale": 3,
"decoder_adaptive_patching_levels": 0,
"decoder_d_model": 128,
"decoder_d_model_scale": 2,
"decoder_mode": "common_channel",
"decoder_num_layers": 2,
"decoder_raw_residual": false,
"distribution_output": "student_t",
"dropout": 0.4,
"dtype": "float32",
"enable_forecast_channel_mixing": false,
"exogenous_channel_indices": null,
"expansion_factor": 2,
"fcm_context_length": 1,
"fcm_gated_attn": true,
"fcm_mix_layers": 3,
"fcm_prepend_past": true,
"fcm_prepend_past_offset": null,
"fcm_use_mixer": true,
"frequency_token_vocab_size": 5,
"gated_attn": true,
"head_dropout": 0.4,
"huber_delta": 1,
"init_embed": "pytorch",
"init_linear": "pytorch",
"init_processing": true,
"init_std": 0.02,
"loss": "mse",
"mask_value": 0,
"masked_context_length": null,
"mode": "common_channel",
"model_type": "tinytimemixer",
"norm_eps": 1e-05,
"norm_mlp": "LayerNorm",
"num_input_channels": 1,
"num_layers": 2,
"num_parallel_samples": 100,
"num_patches": 8,
"patch_last": true,
"patch_length": 64,
"patch_stride": 64,
"positional_encoding_type": "sincos",
"post_init": false,
"prediction_channel_indices": null,
"prediction_filter_length": null,
"prediction_length": 720,
"quantile": 0.5,
"resolution_prefix_tuning": false,
"scaling": "std",
"self_attn": false,
"self_attn_heads": 1,
"stride_ratio": 1,
"transformers_version": "4.57.6",
"use_decoder": true,
"use_positional_encoding": false
}
Binary file not shown.
Loading
Loading