Smart Grid Transformer Integration & Automated Scenario Generation Pipeline - #288
Rohith-Kanathur wants to merge 25 commits into
Conversation
|
Hi @Sagar-CK @Rohith-Kanathur I would like to know why only 5 data points are in the bulk_docs_transformer.json |
|
Hi @ShuxinLin The current bulk_docs_transformer.json only contains 5 data points because it was initially added as lightweight mock/test data to verify if the CouchDB ingestion was happening correctly. The intent wasn’t to model a full production level smart grid dataset but rather to validate the loader + CouchDB integration flow first. We can definitely expand the dataset with additional transformer documents. I was thinking about using this dataset since it contains around 400 transformer data points: https://data.mendeley.com/datasets/rz75w3fkxy/1 Please let me know what you think. Thank you. |
|
@Rohith-Kanathur You needed these sample data for what? These PR is for generating scenarios? or testing Scenario? |
|
@DhavalRepo18 I was talking about this github issue: #303 |
711d9f1 to
1acf560
Compare
1acf560 to
6ee4925
Compare
… live comparisons
…ffline comparison
Overview
This PR delivers two major contributions to AssetOpsBench:
ScenarioGeneratorAgent) that scales benchmark creation to new asset classes without manual authoring1. Smart Grid Transformer Integration
CouchDB Data
Four New FMSR Tools
predict_health_indexinterpret_dgaassess_winding_temperatureassess_load_profileTests
2. Scenario Generation Pipeline (
ScenarioGeneratorAgent)A three-phase automated pipeline that generates physically plausible, causally consistent, and tool-reachable benchmark scenarios for any onboarded asset class.
Phase 1: Asset Profiling
Discovers live asset instances, sensors, and failure mappings from CouchDB, retrieves and synthesizes domain literature from ArXiv or Semantic Scholar, and merges everything with MCP tool schemas into a single structured
AssetProfilethat grounds all downstream generation.Phase 2: Budget Allocation
Distributes the total scenario count across focusses (
iot,fmsr,tsfm,wo,vibration,multiagent) proportionally to the asset's available data modalities and tool coverage, withmultiagentcapped at 75% of total budget to preserve lane diversity.Phase 3: Scenario Generation & Validation
Generates per-focus scenarios conditioned on the asset profile, runs each candidate through an LLM-based repair step.
Output
Each run produces a timestamped directory with
scenarios.json. Each scenario object contains anid,type,text,category, andcharacteristic_form.CLI