Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
26 commits
Select commit Hold shift + click to select a range
4c9514b
transformer asset integration initial commit
Rohith-Kanathur Mar 27, 2026
c8c7fd0
added failure modes for transformer
Rohith-Kanathur Mar 29, 2026
2060777
add transformer health index prediction tool
Rohith-Kanathur Apr 1, 2026
1d0e12b
added transformer asset docs
Rohith-Kanathur Apr 1, 2026
6ee4925
added mocks and docstrings
Rohith-Kanathur Apr 1, 2026
f9d25c2
fix(fmsr): allow local health prediction without LLM credentials
Sagar-CK Sep 30, 2026
e87f4e6
feat(llm): support generation token budgets and verify provider requests
Sagar-CK Sep 30, 2026
5659599
feat(llm): add Claude Code generation adapter
Sagar-CK Sep 30, 2026
d03bb31
feat(llm): add Codex generation adapter
Sagar-CK Sep 30, 2026
a163948
feat(llm): add Z.ai GLM generation adapter
Sagar-CK Sep 30, 2026
aa6f7fe
feat(scenarios): validate tool references, grounding and duplicate re…
Sagar-CK Sep 30, 2026
d51c997
feat(scenarios): add literature retrieval and research synthesis
Sagar-CK Sep 30, 2026
e308d65
feat(servers): expose asset and vibration coverage for grounding
Sagar-CK Sep 30, 2026
4aa06c1
feat(scenarios): add configurable generation and repair pipeline
Sagar-CK Sep 30, 2026
7dcc686
docs(scenarios): explain generator workflow, CLI and environment setup
Sagar-CK Sep 30, 2026
bc77f1a
feat(scenarios): support z.ai and cap generation at model output limits
Sagar-CK Sep 30, 2026
6375f58
feat(benchmark): measure native agent runs and grade independently in…
Sagar-CK Sep 30, 2026
44bfb56
docs(benchmarks): publish transformer runs, full traces, graphs and o…
Sagar-CK Sep 30, 2026
5c4c7b0
feat(benchmark): repeat comparisons with matched snapshots and report…
Sagar-CK Sep 30, 2026
fc345d4
docs: explain agent execution and independent evaluation
Sagar-CK Sep 30, 2026
45fe5a7
fix(benchmark): count exhausted executions without inventing judge re…
Sagar-CK Sep 30, 2026
e0e1707
fix(benchmark): resume finished repetitions without replaying scenarios
Sagar-CK Sep 30, 2026
a457523
feat(benchmark): display median repetition pass rate in comparison cards
Sagar-CK Sep 30, 2026
37a4d06
fix(benchmark): sanitize interrupted orchestration metadata in publis…
Sagar-CK Sep 30, 2026
f3bccc0
docs(benchmarks): publish three transformer repetitions with averages…
Sagar-CK Sep 30, 2026
8253359
docs(benchmarks): publish average scores for three paper criteria
Sagar-CK Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
29 changes: 29 additions & 0 deletions .env.public
Original file line number Diff line number Diff line change
Expand Up @@ -30,3 +30,32 @@ STIRRUP_CODE_IMAGE=assetops-code
# Rancher Desktop: unix:///Users/you/.rd/docker.sock
# DOCKER_HOST=
# ASSETOPS_SHARED_DIR=/tmp/assetops_shared

# ── Additional database names (optional overrides; defaults shown) ─────────
# Asset registry used to match --mode open requests to installed assets.
ASSET_DBNAME=asset
# Vibration measurements and coverage for live asset grounding.
VIBRATION_DBNAME=vibration
# Curated failure-mode catalog used by FMSR and scenario grounding.
FAILURE_MODE_DBNAME=failure_mode

# ── Scenario research: Semantic Scholar ─────────────────────────────────────
# Optional for --retriever semantic_scholar; recommended for authenticated rate limits.
# Not needed for arXiv or when --reuse-research FILE skips retrieval.
# Put real credentials in the Git-ignored .env; keep this public template blank.
SEMANTIC_SCHOLAR_API_KEY=

# ── Scenario generation: Z.ai GLM ───────────────────────────────────────────
# Required only for --backend glm. Put the real key in the Git-ignored .env.
ZAI_API_KEY=
# Optional API endpoint override; the adapter uses this URL by default.
# ZAI_BASE_URL=https://api.z.ai/api/paas/v4/
# Optional: enabled or disabled. GLM-5.3 (the default model) requires thinking
# and enables it automatically. Other models default to disabled.
# ZAI_THINKING=enabled

# ── Scenario generation: local CLI authentication ──────────────────────────
# --backend claude-code and --backend codex use their own local login; no API key here.
# Optional Claude Code response limit; explicit per-stage limits take precedence.
# CLAUDE_CODE_MAX_OUTPUT_TOKENS=8192
# Choose the backend/model with CLI flags or GeneratorConfig, not environment keys.
6 changes: 6 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -201,6 +201,12 @@ mcp/couchdb/sample_data/bulk_docs.json
.env
mcp/servers/tsfm/artifacts/tsfm_models/
src/tmp/
src/scenarios/tmp/

# Observability artifacts (OTLP-JSON traces + per-run trajectory JSON).
traces/
# ignore generated scenarios
generated_scenarios.json
logs/
# ignore generated scenario outputs and manual scenario analyses
generated/
10 changes: 6 additions & 4 deletions INSTRUCTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -375,10 +375,12 @@ See [docs/observability.md](docs/observability.md) for span attribute reference,

## Evaluation

**[Run agents and evaluate their results](docs/running-evaluations.md)** — step-by-step commands for the Transformer suite, independent Fable grading, five-model execution, live progress and three repetitions.

Offline scoring of saved trajectories against ground-truth scenarios. Three-stage flow:

```
agent run → trajectory (run_id) → uv run evaluate → reports/<run_id>.json
agent run → trajectory (run_id) → uv run evaluate → reports/_aggregate.json
```

End-to-end against a ground-truth file:
Expand All @@ -396,11 +398,11 @@ uv run evaluate \
--judge-model litellm_proxy/azure/gpt-5.4
```

Output lands under `reports/` — one `<run_id>.json` per trajectory plus `_aggregate.json` for the rollup.
Output lands in `reports/_aggregate.json`; its `results` array contains each trajectory's score and rubric details.

> [!NOTE]
> If `llm_judge` is used, `--judge-model` must not match the trajectory's `model`
> for any evaluated run. The evaluator now rejects self-judging rows with a clear error.
> Same-model judging is rejected by default. `--allow-self-judge` explicitly permits it;
> the native Claude Code judge still uses a fresh, tool-free session separate from execution.

Scorer families follow MLflow's evaluator/scorer split: `llm_judge` is wired up; `exact_string_match`, `numeric_match`, and `semantic_similarity` ship as skeletons (raise `NotImplementedError`).

Expand Down
37 changes: 37 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,7 @@ Or jump in instantly:
- 🚀 **[Run on Colab](https://colab.research.google.com/github/IBM/AssetOpsBench/blob/main-0.x/notebook/LLM_Agent.ipynb)** — no install required (illustration of LLM Agent)
- 🎮 **[Try the HF Playground](https://huggingface.co/spaces/ibm-research/AssetOps-Bench)** — interactive demo
- 📖 **[Read INSTRUCTIONS.md](./INSTRUCTIONS.md)** — full setup, MCP servers, plan-execute runner
- **[Run agents and evaluate results](./docs/running-evaluations.md)** — copyable execution, grading, live progress and three-repetition commands

> [!NOTE]
> Active development is on `main`. The codebase used for various publication venues continues to be maintained on separate branches, for example, ACL 2026 [`IndustryAssetEQA`](https://github.com/IBM/AssetOpsBench/tree/IndustryAssetEQA) and prior experimental work is maintained on [`main-0.x`](https://github.com/IBM/AssetOpsBench/tree/main-0.x).
Expand Down Expand Up @@ -124,6 +125,42 @@ Some tasks focus on a single domain, others are multi-step end-to-end workflows.

## Leaderboards

### Generated transformer comparison · k = 3 · September 30, 2026

Five models each ran the same 52 open-form scenarios three times from the same initial database snapshot: **780 assigned trials, 779 independent Fable 5.1 judgments, one terminal GLM execution failure**. The [repeated-run report](benchmarks/runs/2026-09-30-transformer-k3/README.md) includes every repetition, median pass rates, means and sample standard deviations, scenario repeatability, full traces and [offline HTML](benchmarks/runs/2026-09-30-transformer-k3/comparison.html).

The [latest paper](https://arxiv.org/html/2506.03828v4#S5) reports task completion, data retrieval accuracy and result verification separately. Below are the corresponding averages from our saved judgments, giving each execution repetition equal weight. Values are **mean ± sample SD in percentage points**; each criterion is averaged independently of the strict overall pass gate.

| Model | Task completion (%) | Data retrieval accuracy (%) | Result verification (%) | Judged / assigned |
|---|---:|---:|---:|---:|
| Opus 5.5 | 57.7 ± 5.1 | 96.2 ± 1.9 | 68.6 ± 6.8 | 156/156 |
| GPT-6 Astra | 54.5 ± 2.9 | 90.4 ± 0.0 | 62.8 ± 4.4 | 156/156 |
| GLM 5.3 (low) | 58.1 ± 5.8 | 93.6 ± 4.4 | 55.5 ± 9.6 | 155/156 |
| GPT-6.1 Sol | 51.9 ± 1.9 | 96.2 ± 1.9 | 59.6 ± 1.9 | 156/156 |
| Fable 5.1 | 64.7 ± 7.8 | 99.4 ± 1.1 | 76.3 ± 2.9 | 156/156 |

![Transformer average criterion scores](benchmarks/runs/2026-09-30-transformer-k3/graphs/criterion-averages.png)

[Download criterion averages](benchmarks/runs/2026-09-30-transformer-k3/criterion-averages.csv). These runs use one successful Fable 5.1 judgment per execution across three execution repetitions; the paper averages five Llama-4-Maverick judgments per trajectory. The criterion names match, while the judge protocol, scenarios and environment differ. GLM's missing judgment is excluded from criterion averages and remains a nonpassing assigned trial in the overall pass rate.

![Transformer mean pass rates and variation](benchmarks/runs/2026-09-30-transformer-k3/graphs/pass-rate.png)

![Transformer execution time across three repetitions](benchmarks/runs/2026-09-30-transformer-k3/graphs/execution-time.png)

HTML cards show the median of the three repetition pass rates; tables and graphs retain means ± sample SD. Execution timing covers the entire agent invocation, grading is separate, and all failed/retried attempts are retained. FMSR's unconfigured Watsonx backend remains an environment limitation; see the report for the method and evidence.

### Generated transformer comparison · September 30, 2026

Five models completed 52 open-form scenarios, each graded in an independent Fable 5.1 session. The [run report](benchmarks/runs/2026-09-30-transformer/README.md) includes per-scenario results, full observed traces, runtime settings, token/tool metrics and the [offline HTML comparison](benchmarks/runs/2026-09-30-transformer/comparison.html).

![Transformer pass rates](benchmarks/runs/2026-09-30-transformer/graphs/pass-rate.png)

![Transformer execution times](benchmarks/runs/2026-09-30-transformer/graphs/execution-time.png)

Timing covers the entire agent invocation; grading is measured separately. Results include three retained GLM retries. FMSR's unconfigured Watsonx backend affected tool availability across models; see the run report for the environment, rubric and interpretation limits.

### Earlier benchmark results

- To be revised (WIP with latest models)
- Evaluated with **7 Large Language Models**
- Trajectories scored using **LLM Judge (Llama-4-Maverick-17B)**
Expand Down
14 changes: 14 additions & 0 deletions benchmarks/generated-comparison.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"suite": "benchmarks/runs/2026-09-30-transformer/suite",
"judge": "claude-code/claude-fable-5-1",
"judge_runtime": "separate tool-free Claude Code CLI session per scenario",
"allow_same_model_judge": true,
"database_policy": null,
"targets": [
{"name":"opus-5-5","agent":"claude","model_id":"claude-opus-5-5"},
{"name":"gpt-6-astra","agent":"codex","model_id":"gpt-6-astra"},
{"name":"glm-5-3-low","agent":"openai","model_id":"zai/glm-5.3","reasoning_effort":"low"},
{"name":"gpt-6-1-sol","agent":"codex","model_id":"gpt-6.1-sol"},
{"name":"fable-5-1","agent":"claude","model_id":"claude-fable-5-1"}
]
}
176 changes: 176 additions & 0 deletions benchmarks/generated-scenarios.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,176 @@
# Comparing agents on generated scenarios

For setup and copyable commands, see [Run agents and evaluate their results](../docs/running-evaluations.md).

Use `python -m benchmark.generated_suite_runner` from the repository with
`PYTHONPATH=src`. The runner requires a completed generation manifest and the
exact requested positive/negative counts. It reads both scenario JSON files,
runs tool-enabled agents, and saves separate trajectories and logs per target.
It never resets or reloads the live databases.

Start with `--dry-run`, then `--limit 1` to verify model access and tool calls.
Remove those flags for the full suite. Repeating a command resumes matching
saved scenarios. Changing the model, harness or generation run under an existing
target name is rejected. Failed calls stop the target and retain its logs.

```bash
# Replace RUN_DIR with the completed generated run directory.
PYTHONPATH=src .venv/bin/python -m benchmark.generated_suite_runner RUN_DIR \
--output-dir generated/comparisons/transformer --name opus-5-5 \
--agent claude --model-id claude-opus-5-5 --limit 1

# Uses the locally authenticated Codex subscription.
PYTHONPATH=src .venv/bin/python -m benchmark.generated_suite_runner RUN_DIR \
--output-dir generated/comparisons/transformer --name gpt-6-astra \
--agent codex --model-id gpt-6-astra --limit 1

PYTHONPATH=src .venv/bin/python -m benchmark.generated_suite_runner RUN_DIR \
--output-dir generated/comparisons/transformer --name glm-5-3 \
--agent openai --model-id zai/glm-5.3 --limit 1
```

Opus uses Claude Agent SDK with Claude Code authentication. Bare OpenAI model
IDs with `--agent openai` use `OPENAI_API_KEY`. The `--agent codex` runner uses
your existing `codex login` subscription, exposes the benchmark MCP servers,
and saves standard trajectories for the existing evaluator.
Proxy model IDs may instead use `litellm_proxy/` or `tokenrouter/` with the
corresponding credentials. GLM uses `ZAI_API_KEY` and optional `ZAI_BASE_URL`.
Set `GLM_REASONING_EFFORT=low` (or `high` / `max`) for explicit GLM reasoning.
Record this setting separately for each comparison target.
The CLI loads local `.env`; credentials must never go in this document.

The examples compare different SDK harnesses. For a model comparison under a
single harness, use `--agent openai` for every target and an OpenAI-compatible
proxy route for Opus. Verify each exact model route with the one-scenario smoke
run before launching the full comparison.

To score, use the same judge across all targets. Add
`--judge-model MODEL_ID` to each full-run command, using an independent judge session. Same-model judging is rejected by default;
`--allow-self-judge` explicitly permits it with a separate session. Alternatively invoke
`python -m evaluation.cli` on its trajectory directory with both scenario files.
Native Claude Code and `zai/` model IDs are supported as judges. A judge score
uses the generated characteristic behavior as its rubric; it is not independently
verified ground truth.

Open-form work-order tools can modify database state. Compare against equivalent
starting data for every target. This runner does not restore snapshots; establish
the benchmark database reset/snapshot policy before a full comparison. Review
logs for tool failures, not only final scores.

## Fresh-run measurements

All targets now measure the entire child-agent invocation, including CLI/SDK
startup, MCP connection, model and tool work, persistence, cleanup and exit.
The record is written before launch and finalized for completed, failed,
timed-out or cancelled invocations. Retries retain separate attempt files.
No model runs are started by reporting or configuration commands.

Each target writes:

- `settings.json`: exact model/provider/harness, installed CLI/SDK versions,
suite and rubric hashes, reasoning setting, limits, timeout/retry policy,
database policy, concurrency and order. Unknown provider defaults are null.
- `measurements/<run-id>.json`: start/end/duration, run status/error,
observed tokens and tool statistics, and separately timed grading outcomes.
- `traces/<run-id>.jsonl`: incrementally saved full observed messages,
request inputs/responses, tool arguments/results/errors and UTC timestamps.
Traces survive interrupted invocations. Judge traces have a `.judge` suffix.
- `trajectories/<run-id>.json`: the established evaluator input format.

The five target definitions are in `benchmarks/generated-comparison.json`.
Fable 5.1 judges each model in a fresh, tool-free Claude Code session. Pass
`--allow-self-judge` for the Fable execution target, as explicitly requested;
this permits a separate judge session using the same model. `--judge-model`
uses measured per-scenario grading rather than a second timing aggregate.
Set GLM explicitly with `--reasoning-effort low`. Record total simultaneous
execution targets using `--concurrency-level N`; each target remains serial.

First-response timing means the first observed assistant/protocol response,
not first-token latency. OpenAI-compatible SDK request timing includes its
internal retry handling. Claude/Codex do not expose all underlying request
boundaries, retry counts, reasoning-token subdivisions or compactions; those
fields remain null when unavailable. CLI token costs are provider-reported
estimates, not actual subscription charges. Actual billed API cost is null
unless a billing source supplies it; no guessed prices are used in this report.

For authoritative database-write auditing, use an **existing isolated**
`eval_...` CouchDB namespace with `python -m benchmark.database_audit
--prefix PREFIX --port PORT --audit-dir DIR`. It reads upstream connection
credentials from `.env`. Set `BENCHMARK_DB_PROXY_URL` to its localhost URL and
`BENCHMARK_DB_AUDIT_DIR` to the same directory before executing a target. The
standalone path mode adds the run ID to the proxy path; writes and record changes
are joined by that ID. For IoT, use the root-URL mode described below. Without auditing, database action metrics remain null. Store the
snapshot hash, source, reset policy and proxy namespace in the target's
`environment.json` before launch. This helper does not reset the live database.

Generate comparisons from measured runs only:

```bash
PYTHONPATH=src .venv/bin/python -m benchmark.comparison_report \
generated/comparisons/transformer --output generated/comparisons/transformer/comparison.html
```

Pass rates and rubric success rates use graded cases, with observed denominators
shown. Median/p95 execution time uses completed invocations with measured time;
failed and cancelled attempts remain in reliability counts. Missing token and
tool values are excluded from means and totals, and availability counts are
included. Historical timing is never reconstructed.

The configured five-model comparison can be launched with:

```bash
PYTHONPATH=src .venv/bin/python tools/run_generated_comparison.py
```

This captures one common source snapshot, clones a namespace per target, checks
connectivity with the actual IoT CouchDB client, runs the targets concurrently,
and grades each target with independent Fable sessions. For the local Codex
subscription, it uses the app's bundled CLI; the older CLI on PATH rejects
GPT-6.1 Sol. The measured records preserve the exact CLI version used.

The comparison launcher uses one root-URL proxy per target and an active-run
file to attribute writes, because the IoT `couchdb3` client discards URL paths.
`BENCHMARK_DB_RUN_ID_FILE` tells the suite runner to update that attribution
before each invocation. Authentication requests do not count as database writes.
The standalone path-based proxy mode is intended for clients that retain paths.

Live view:

```bash
PYTHONPATH=src .venv/bin/python tools/live_evaluation/server.py
```

Open `http://127.0.0.1:8765`. The final offline HTML shares the live view's layout.

## Independent repetitions

Keep the published transformer comparison as repetition 1 and execute two more:

```bash
PYTHONPATH=src .venv/bin/python tools/run_repeated_comparison.py \
--output-dir generated/comparisons/transformer-k3

PYTHONPATH=src .venv/bin/python tools/live_evaluation/server.py --port 8766 \
--experiment generated/comparisons/transformer-k3/experiment.json
```

Repetitions run serially; the five model targets and independent grading workers
run concurrently within each repetition. The launcher verifies the reference
suite and initial database hash, preserves one snapshot, and clones fresh isolated
namespaces for each repetition/model. Scenario order, harnesses, reasoning,
timeouts and judge remain the same. Each measurement records its repetition index.

When all three repetitions have finished:

```bash
uv run tools/publish_repeated_comparison.py \
--experiment generated/comparisons/transformer-k3/experiment.json
```

The report shows equal-weight means and sample standard deviations across full
repetitions, individual repetition results, pooled metrics and per-scenario pass
frequencies. Failed/retried invocation attempts remain in reliability metrics;
they are not additional repetitions. Missing metrics remain missing. The live
view uses only fully completed repetitions for its averages. The publisher checks
suite/settings consistency and distinct judge sessions before emitting a final
comparison, graphs, CSV/JSON data and portable compressed traces.
Loading