Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
34 commits
Select commit Hold shift + click to select a range
d8982f4
First ipman agentic forecasting commit with data
vbeohar Sep 2, 2026
19f86a9
Added agent predictor
vbeohar Sep 22, 2026
3c909b9
Added 5 more macro expl vars
vbeohar Sep 22, 2026
532170e
Merge pull request #1 from VectorInstitute/main
minaattari Sep 22, 2026
255b592
Add XGBoost manufacturing stress predictor
nasas Sep 22, 2026
6722520
Merge pull request #2 from minaattari/add-xgboost-manufacturing-stress
minaattari Sep 22, 2026
9abcd93
feat(manufacturing): add notebook workbench and cached LLMP backtest
minaattari Sep 23, 2026
6877b06
fix(manufacturing): harden agent backtest evaluation
vbeohar Sep 23, 2026
f9e41e8
Merge pull request #3 from minaattari/branch_mina
pratyushpatodia-dev Sep 23, 2026
ea2bd3b
safety-bug-fix
vbeohar Sep 23, 2026
be1834c
Expanded data fetch. Updated downstream components to ingest expanded…
pratyushpatodia-dev Sep 23, 2026
3f13a08
controlled param sweep
vbeohar Sep 23, 2026
ad6c993
Merge pull request #4 from minaattari/day2-data-fetch-pp
vbeohar Sep 23, 2026
9759cb6
Added NY Fed GSCPI data
vbeohar Sep 23, 2026
6c5baa3
Experimented with GSCPI NYFed data
vbeohar Sep 23, 2026
f6af597
Added stress param hyperparams
vbeohar Sep 24, 2026
47df9d3
Merge origin/main into fix/manufacturing-agent-json-output
vbeohar Sep 24, 2026
d2ca24e
Merge pull request #5 from minaattari/fix/manufacturing-agent-json-ou…
pratyushpatodia-dev Sep 24, 2026
2c70f66
Fix manufacturing-stress credit spread feature and drop GSCPI
vbeohar Sep 24, 2026
bce63ed
Add --stress-threshold-pct flag to the parameter sweep
vbeohar Sep 24, 2026
9ced7b1
Merge pull request #6 from minaattari/fix/credit-spread-full-history
minaattari Sep 25, 2026
0dfa932
Improve manufacturing parameter tuning
minaattari Oct 2, 2026
b885a2c
Update FRED adapter and getting-started notebooks
minaattari Oct 2, 2026
c94c527
Merge remote-tracking branch 'bmo-c2/main' into copilot/push-to-bmo-c2
minaattari Oct 2, 2026
0f32996
Merge branch 'more_regulation_on_classic_methods' into HEAD
minaattari Oct 2, 2026
afd0ea8
Document manufacturing stress expansion plan
minaattari Oct 2, 2026
a43b4d9
Add manufacturing stress hybrid workbench
minaattari Oct 2, 2026
475f2c9
Use bounded hybrid adjustments as forecast probabilities
minaattari Oct 2, 2026
2112239
Add manufacturing adaptive evaluation workbenches
minaattari Oct 5, 2026
0492213
Publish manufacturing reports and planning artifacts
minaattari Oct 5, 2026
062259f
Add manufacturing adaptive comparison report
minaattari Oct 5, 2026
775ca7a
Add manufacturing workbench notebooks and reports
minaattari Oct 5, 2026
ea5aaa8
Add learned manufacturing-strategy skill state, SKILL.md and adaptati…
minaattari Oct 6, 2026
8412850
Remove optional backtest from adaptive agent workbench
minaattari Oct 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -42,6 +42,10 @@ implementations/**/data/
!implementations/energy_oil_forecasting/data/predictions/
!implementations/energy_oil_forecasting/data/predictions/**

# Manufacturing workbench reports are reproducible project artifacts.
!implementations/manufacturing_stress_forecasting/reports/
!implementations/manufacturing_stress_forecasting/reports/**

# workspace-specific agents
.github/agents/
.github/prompts/
Expand Down
3 changes: 2 additions & 1 deletion aieng-forecasting/aieng/forecasting/data/adapters/fred.py
Original file line number Diff line number Diff line change
Expand Up @@ -86,7 +86,8 @@ def __init__(
refresh: bool = False,
) -> None:
self._series_id = series_id
self._api_key = api_key or os.environ.get("FRED_API_KEY")
raw_key = (api_key if api_key is not None else os.environ.get("FRED_API_KEY", "")).strip()
self._api_key = raw_key or None
self._cache_dir = Path(cache_dir) if cache_dir is not None else None
self._refresh = refresh

Expand Down
4 changes: 2 additions & 2 deletions aieng-forecasting/aieng/forecasting/methods/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,7 +102,7 @@ from aieng.forecasting.methods.agentic import (

| Module | Class / Function | Description |
|---|---|---|
| `agentic/adk_runner.py` | `AdkTextRunner` | Async text-in / text-out wrapper around ADK `InMemoryRunner`. Manages ADK sessions (fresh-per-message or sticky) and optionally traces each turn to Langfuse via `propagate_attributes`. |
| `agentic/adk_runner.py` | `AdkTextRunner` | Async text-in / text-out wrapper around ADK `InMemoryRunner`. Drains each event stream, joins non-thought text parts from the first usable final response, prefers structured JSON submitted through `set_model_response`, manages ADK sessions (fresh-per-message or sticky), and optionally traces each turn to Langfuse via `propagate_attributes`. |
| `agentic/adk_runner.py` | `AdkTextRunnerConfig` | Pydantic configuration for `AdkTextRunner` (session mode, Langfuse fields). |
| `agentic/agent_factory.py` | `build_adk_agent` | Generic ADK `LlmAgent` factory with optional code execution, context retrieval, skills, generation controls, and structured output schema. `search_web` runs a cutoff-aware googleSearch sub-call plus an independent leakage verifier on historical origins (skipped when `as_of` is today or later); both inner LLM calls emit nested Langfuse generations (`search_web.google_search`, `search_web.leakage_verifier`) under the agent trace. |
| `agentic/agent_factory.py` | `AgentConfig` | Pydantic configuration for reusable ADK agents. `output_schema=None` supports interactive/free-form agents; a structured `AgentForecastOutput` schema supports Track 1 predictors. The `function_tools` field attaches conventional ADK tools (e.g. `ForecastTool`). Use-case-specific prompts and presets should live in `implementations/<use-case>/`. |
Expand All @@ -111,5 +111,5 @@ from aieng.forecasting.methods.agentic import (
| `agentic/outputs.py` | `ContinuousAgentForecastOutput` | Canonical continuous forecasting output schema. Declares `modality = "continuous"`, requires one forecast per task horizon and the standard quantile grid, then converts to `ContinuousForecast` payloads. |
| `agentic/outputs.py` | `DiscreteAgentForecastOutput` | Binary event output schema (`modality = "discrete"`): one probability plus `reasoning` / `key_signals` metadata, converted to a `BinaryForecast` payload. |
| `agentic/outputs.py` | `CategoricalAgentForecastOutput` | Ordered-categorical output schema (`modality = "categorical"`): one `{label, probability}` row per task category, validated against `task.categories` and converted to a `CategoricalForecast` payload. |
| `agentic/predictor.py` | `AgentPredictor` | Track 1 `Predictor` that builds prompts, runs an ADK agent through `AdkTextRunner`, validates structured JSON, and converts it to `Prediction` objects. Accepts an optional injected runner for tests or custom observability. |
| `agentic/predictor.py` | `AgentPredictor` | Track 1 `Predictor` that builds prompts, runs an ADK agent through `AdkTextRunner`, validates structured JSON, and converts it to `Prediction` objects. Accepts an optional injected runner for tests or custom observability; schema-validation failure logs include the Langfuse trace ID when available. |
| `agentic/predictor.py` | `ForecastPromptBuilder` | Protocol for task-specific prompt builders that turn `(task, context)` into the text passed to the agent. |
15 changes: 12 additions & 3 deletions aieng-forecasting/aieng/forecasting/methods/agentic/adk_runner.py
Original file line number Diff line number Diff line change
Expand Up @@ -295,15 +295,24 @@ async def run_text_async(
content = genai_types.Content(role="user", parts=[genai_types.Part(text=prompt)])

async def drain_run() -> str:
final_text: str | None = None
async for event in self._runner.run_async(
user_id=user_id,
session_id=session_id,
new_message=content,
run_config=run_config,
):
if event.is_final_response() and event.content and event.content.parts:
return event.content.parts[0].text or ""
return ""
if final_text is None and event.is_final_response() and event.content and event.content.parts:
text_parts = [
part.text
for part in event.content.parts
if isinstance(part.text, str)
and part.text.strip()
and getattr(part, "thought", False) is not True
]
if text_parts:
final_text = "".join(text_parts)
return final_text or ""

async def run_and_resolve() -> str:
"""Run the agent and return the best available output string.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -296,7 +296,11 @@ def predict(self, task: ForecastingTask, context: ForecastContext) -> list[Predi
try:
output = self.output_schema.model_validate(json.loads(output_str))
except Exception:
logger.warning("Raw agent response (schema validation failed):\n%s", output_str)
logger.warning(
"Raw agent response (schema validation failed; langfuse_trace_id=%s):\n%s",
self._runner.last_trace_id,
output_str,
)
raise

# Convert output to list of predictions
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -67,3 +67,12 @@ def test_missing_api_key_without_cache_raises(tmp_path: Path) -> None:
adapter = FREDAdapter("EXCAUS", api_key=None, cache_dir=cache_dir)
with pytest.raises(ValueError, match="FRED API key not provided"):
adapter.fetch()


def test_api_key_whitespace_is_trimmed() -> None:
"""Whitespace around a pasted key must not be included in the request URL."""
fake = _fred_cls_returning(_raw_fred_series())
with patch("fredapi.Fred", fake):
FREDAdapter("EXCAUS", api_key=" fake-key\n", cache_dir=None, refresh=True).fetch()

assert fake.call_args.kwargs["api_key"] == "fake-key"
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ def _final_event(text: str) -> MagicMock:
event.is_final_response.return_value = True
part = MagicMock()
part.text = text
part.thought = False
event.content.parts = [part]
return event

Expand Down Expand Up @@ -373,6 +374,41 @@ async def test_returns_text_from_first_final_event(self, patch_runner_cls, mock_
runner = AdkTextRunner(mock_agent, config=AdkTextRunnerConfig(app_name="app"))
assert await runner.run_text_async("hi") == "hello world"

async def test_joins_final_text_parts_and_ignores_thought_parts(self, patch_runner_cls, mock_agent) -> None:
"""Use final answer text, not internal thought or function-call parts."""
event = _final_event("")
thought = MagicMock()
thought.text = "internal thought"
thought.thought = True
json_start = MagicMock()
json_start.text = '{"probability":'
json_start.thought = False
json_end = MagicMock()
json_end.text = "0.4}"
json_end.thought = False
event.content.parts = [thought, json_start, json_end]
patch_runner_cls.run_async.return_value = _stream(event)

runner = AdkTextRunner(mock_agent, config=AdkTextRunnerConfig(app_name="app"))

assert await runner.run_text_async("forecast") == '{"probability":0.4}'

async def test_drains_stream_after_first_final_event(self, patch_runner_cls, mock_agent) -> None:
"""Consume cleanup events after capturing the final response."""
consumed = False

async def stream_with_trailing_event():
nonlocal consumed
yield _final_event("hello world")
yield _intermediate_event()
consumed = True

patch_runner_cls.run_async.return_value = stream_with_trailing_event()
runner = AdkTextRunner(mock_agent, config=AdkTextRunnerConfig(app_name="app"))

assert await runner.run_text_async("hi") == "hello world"
assert consumed

async def test_returns_empty_string_when_stream_has_no_final_event(self, patch_runner_cls, mock_agent) -> None:
"""Stream with only non-final events yields an empty string."""
patch_runner_cls.run_async.return_value = _stream(_intermediate_event())
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -306,6 +306,16 @@ def test_schema_validation_errors_are_not_swallowed(self) -> None:
with pytest.raises(ValidationError):
predictor.predict(_task([1]), _context())

def test_schema_error_log_includes_langfuse_trace_id(self, caplog: pytest.LogCaptureFixture) -> None:
"""Include trace IDs in parse failures for easier generation inspection."""
predictor, _ = _make_predictor(response="")
predictor._runner.last_trace_id = "trace-123"

with caplog.at_level(logging.WARNING), pytest.raises(json.JSONDecodeError):
predictor.predict(_task([1]), _context())

assert "langfuse_trace_id=trace-123" in caplog.text


# ---------------------------------------------------------------------------
# Async/sync bridge
Expand Down
Binary file added data/fred/CPIAUCSL.parquet
Binary file not shown.
Binary file added data/fred/DFF.parquet
Binary file not shown.
Binary file added data/fred/DGS10.parquet
Binary file not shown.
Binary file added data/fred/DGS2.parquet
Binary file not shown.
Binary file added data/fred/FEDFUNDS.parquet
Binary file not shown.
Binary file added data/fred/ICSA.parquet
Binary file not shown.
Binary file added data/fred/IPMAN.parquet
Binary file not shown.
Binary file added data/fred/UNRATE.parquet
Binary file not shown.
Binary file added data/fred/VIXCLS.parquet
Binary file not shown.
Binary file added data/yfinance/spy_adj_close_1d.parquet
Binary file not shown.
Binary file added data/yfinance/xli_adj_close_1d.parquet
Binary file not shown.
Binary file added implementations/__pycache__/__init__.cpython-312.pyc
Binary file not shown.
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
Metadata-Version: 2.4
Name: agentic-forecasting-implementations
Version: 0.1.0
Summary: Reference method implementations for the Agentic Forecasting Bootcamp
Author-email: Vector AI Engineering <ai_engineering@vectorinstitute.ai>
Requires-Python: >=3.12
Description-Content-Type: text/markdown
Requires-Dist: aieng-forecasting[agentic,documents,llm,numerical]
Requires-Dist: beautifulsoup4<5,>=4.12
Requires-Dist: xgboost<4,>=2.1

# implementations

Self-contained reference implementations and their helper code.

This is a local uv workspace package. It is installed automatically when you run `uv sync` from the repository root, but it is not a separately published public API.

Some use cases are notebook-only. Others expose a small importable helper package so shared analysis, plotting, or data-registration code can live in Python modules instead of large notebook cells.

---

## Directory layout

Numbered in the recommended order (mirrors the bootcamp progression: conventional numerical methods → LLM Processes → agents → agentic evaluation). The directories are not renamed — the numbers are an ordering convention used across the docs, and each directory stays an importable package (`from sp500_forecasting.data import ...`).

```text
implementations/
|-- getting_started/ # 0 · CPI gasoline hello-world (start here)
| `-- specs/ # backtest and eval YAML
|-- sp500_forecasting/ # 1 · S&P 500 multivariate numerical comparison (financial markets)
| `-- specs/ # backtest YAML (smoke + full)
|-- food_price_forecasting/ # 2 · CFPR-style food CPI experiment
| `-- specs/ # backtest YAML
|-- energy_oil_forecasting/ # 3 · Daily WTI oil price forecasting experiment
| `-- specs/ # backtest and eval YAML
|-- boc_rate_decisions/ # 4 · Discrete-event reference: BoC cut/hold/hike direction
| `-- specs/ # direction + binary backtest / eval / smoke YAML
|-- tests/ # tests for implementation-specific helper modules
`-- pyproject.toml # local workspace packaging
```

YAML backtest and eval specs live under each use case in `specs/`. Each directory is independent; see its `README.md` for the walkthrough. For the build-phase moves — onboarding data, standing up an experiment, customizing an agent, auditing a result — see [`guides/`](../guides/). To chat with the concierge or a domain starter in the ADK browser UI, see [`guides/05-access-adk-web-via-ssh-tunnel.md`](../guides/05-access-adk-web-via-ssh-tunnel.md) (includes the Coder SSH tunnel).

Every domain use case (all except `getting_started`) also ships a `starter_agent/` module and a `99_starter_agent.ipynb` — a fresh, hackable **starter agent** that is the consistent "build your own" entry point for that use case (toggleable news search + code execution, two lightweight tool-usage skills, an interactive cell, and one scored forecast).

`getting_started/` additionally ships a **`concierge_agent/`** module and **`99_repo_concierge.ipynb`** — a repo onboarding helper (not a forecaster) that answers questions about how the codebase works using a committed public-`main` knowledge digest. From the repository root: `uv run adk run implementations/getting_started/concierge_agent` (or `uv run adk web implementations/getting_started/concierge_agent` for the browser UI — [guide 5](../guides/05-access-adk-web-via-ssh-tunnel.md)). See [`getting_started/README.md`](getting_started/README.md) and the notebook for full usage.

---

## Relationship to `aieng-forecasting`

- `aieng-forecasting` (`aieng.forecasting`) owns reusable infrastructure and reusable reference predictors under `aieng.forecasting.methods`.
- `implementations/` owns use-case material: walkthrough notebooks, experiment-specific helper modules, plotting/analysis code, and task-specific framing.

If code becomes broadly reusable across use cases, promote it into `aieng-forecasting`.

---

## Adding a new use case

1. Create `implementations/<use-case>/`.
2. Add a `README.md` describing the task, the data, and what the notebooks cover.
3. Add YAML specs under `implementations/<use-case>/specs/`.
4. Start with notebooks as the primary user surface.
5. If notebook code becomes bulky or repeated, extract small helper modules into that use-case directory.
6. Add tests under `implementations/tests/<use-case>/` for non-trivial helper logic.
7. Promote code into `aieng-forecasting` once it is clearly reusable across more than one use case.

For architecture principles and cross-cutting extension ideas, see `planning-docs/roadmap.md`.
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
README.md
pyproject.toml
agentic_forecasting_implementations.egg-info/PKG-INFO
agentic_forecasting_implementations.egg-info/SOURCES.txt
agentic_forecasting_implementations.egg-info/dependency_links.txt
agentic_forecasting_implementations.egg-info/requires.txt
agentic_forecasting_implementations.egg-info/top_level.txt
boc_rate_decisions/__init__.py
boc_rate_decisions/analysis.py
boc_rate_decisions/data.py
boc_rate_decisions/plots.py
boc_rate_decisions/press_releases.py
boc_rate_decisions/rationale_eval.py
boc_rate_decisions/analyst_agent/__init__.py
boc_rate_decisions/analyst_agent/agent.py
boc_rate_decisions/predictors/__init__.py
boc_rate_decisions/predictors/llmp_binary.py
boc_rate_decisions/predictors/llmp_direction.py
boc_rate_decisions/predictors/logistic_baseline.py
boc_rate_decisions/starter_agent/__init__.py
boc_rate_decisions/starter_agent/agent.py
energy_oil_forecasting/__init__.py
energy_oil_forecasting/analysis.py
energy_oil_forecasting/data.py
energy_oil_forecasting/paths.py
energy_oil_forecasting/prophet_baseline.py
energy_oil_forecasting/tasks.py
energy_oil_forecasting/viz.py
energy_oil_forecasting/adaptive_agent/__init__.py
energy_oil_forecasting/adaptive_agent/agent.py
energy_oil_forecasting/adaptive_agent/skill_state.py
energy_oil_forecasting/adaptive_agent/skill_tools.py
energy_oil_forecasting/adaptive_agent/curriculum/snapshot_utils.py
energy_oil_forecasting/analyst_agent/__init__.py
energy_oil_forecasting/analyst_agent/agent.py
energy_oil_forecasting/starter_agent/__init__.py
energy_oil_forecasting/starter_agent/agent.py
energy_oil_forecasting/starter_agent/tools.py
food_price_forecasting/__init__.py
food_price_forecasting/analysis.py
food_price_forecasting/data.py
food_price_forecasting/plots.py
food_price_forecasting/reports.py
food_price_forecasting/smoke_report.py
food_price_forecasting/predictors/__init__.py
food_price_forecasting/predictors/llmp_quantile_grid.py
food_price_forecasting/predictors/llmp_sampled_trajectory.py
food_price_forecasting/starter_agent/__init__.py
food_price_forecasting/starter_agent/agent.py
getting_started/__init__.py
getting_started/concierge_agent/__init__.py
getting_started/concierge_agent/agent.py
getting_started/concierge_agent/catalog.py
getting_started/concierge_agent/catalog_build.py
getting_started/concierge_agent/knowledge.py
manufacturing_stress_forecasting/__init__.py
manufacturing_stress_forecasting/data.py
manufacturing_stress_forecasting/features.py
manufacturing_stress_forecasting/run_agent_backtest.py
manufacturing_stress_forecasting/run_agent_prediction.py
manufacturing_stress_forecasting/run_smoke.py
manufacturing_stress_forecasting/targets.py
manufacturing_stress_forecasting/analyst_agent/__init__.py
manufacturing_stress_forecasting/analyst_agent/agent.py
manufacturing_stress_forecasting/predictors/__init__.py
manufacturing_stress_forecasting/predictors/logistic.py
manufacturing_stress_forecasting/predictors/xgboost.py
sp500_forecasting/__init__.py
sp500_forecasting/analysis.py
sp500_forecasting/data.py
sp500_forecasting/leaderboard.py
sp500_forecasting/plots.py
sp500_forecasting/predictors/__init__.py
sp500_forecasting/predictors/llmp_sampled_trajectory.py
sp500_forecasting/starter_agent/__init__.py
sp500_forecasting/starter_agent/agent.py
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@

Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
aieng-forecasting[agentic,documents,llm,numerical]
beautifulsoup4<5,>=4.12
xgboost<4,>=2.1
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
boc_rate_decisions
energy_oil_forecasting
food_price_forecasting
getting_started
manufacturing_stress_forecasting
sp500_forecasting
Loading