Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Binary file added data/fred/BAMLH0A0HYM2.parquet
Binary file not shown.
Binary file added data/fred/CPIAUCSL.parquet
Binary file not shown.
Binary file added data/fred/DFF.parquet
Binary file not shown.
Binary file added data/fred/DGS10.parquet
Binary file not shown.
Binary file added data/fred/DGS2.parquet
Binary file not shown.
Binary file added data/fred/FEDFUNDS.parquet
Binary file not shown.
Binary file added data/fred/ICSA.parquet
Binary file not shown.
Binary file added data/fred/IPMAN.parquet
Binary file not shown.
Binary file added data/fred/UNRATE.parquet
Binary file not shown.
Binary file added data/fred/VIXCLS.parquet
Binary file not shown.
Binary file added data/yfinance/spy_adj_close_1d.parquet
Binary file not shown.
Binary file added data/yfinance/xli_adj_close_1d.parquet
Binary file not shown.
Binary file added implementations/__pycache__/__init__.cpython-312.pyc
Binary file not shown.
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
Metadata-Version: 2.4
Name: agentic-forecasting-implementations
Version: 0.1.0
Summary: Reference method implementations for the Agentic Forecasting Bootcamp
Author-email: Vector AI Engineering <ai_engineering@vectorinstitute.ai>
Requires-Python: >=3.12
Description-Content-Type: text/markdown
Requires-Dist: aieng-forecasting[agentic,documents,llm,numerical]
Requires-Dist: beautifulsoup4<5,>=4.12
Requires-Dist: xgboost<4,>=2.1

# implementations

Self-contained reference implementations and their helper code.

This is a local uv workspace package. It is installed automatically when you run `uv sync` from the repository root, but it is not a separately published public API.

Some use cases are notebook-only. Others expose a small importable helper package so shared analysis, plotting, or data-registration code can live in Python modules instead of large notebook cells.

---

## Directory layout

Numbered in the recommended order (mirrors the bootcamp progression: conventional numerical methods → LLM Processes → agents → agentic evaluation). The directories are not renamed — the numbers are an ordering convention used across the docs, and each directory stays an importable package (`from sp500_forecasting.data import ...`).

```text
implementations/
|-- getting_started/ # 0 · CPI gasoline hello-world (start here)
| `-- specs/ # backtest and eval YAML
|-- sp500_forecasting/ # 1 · S&P 500 multivariate numerical comparison (financial markets)
| `-- specs/ # backtest YAML (smoke + full)
|-- food_price_forecasting/ # 2 · CFPR-style food CPI experiment
| `-- specs/ # backtest YAML
|-- energy_oil_forecasting/ # 3 · Daily WTI oil price forecasting experiment
| `-- specs/ # backtest and eval YAML
|-- boc_rate_decisions/ # 4 · Discrete-event reference: BoC cut/hold/hike direction
| `-- specs/ # direction + binary backtest / eval / smoke YAML
|-- tests/ # tests for implementation-specific helper modules
`-- pyproject.toml # local workspace packaging
```

YAML backtest and eval specs live under each use case in `specs/`. Each directory is independent; see its `README.md` for the walkthrough. For the build-phase moves — onboarding data, standing up an experiment, customizing an agent, auditing a result — see [`guides/`](../guides/). To chat with the concierge or a domain starter in the ADK browser UI, see [`guides/05-access-adk-web-via-ssh-tunnel.md`](../guides/05-access-adk-web-via-ssh-tunnel.md) (includes the Coder SSH tunnel).

Every domain use case (all except `getting_started`) also ships a `starter_agent/` module and a `99_starter_agent.ipynb` — a fresh, hackable **starter agent** that is the consistent "build your own" entry point for that use case (toggleable news search + code execution, two lightweight tool-usage skills, an interactive cell, and one scored forecast).

`getting_started/` additionally ships a **`concierge_agent/`** module and **`99_repo_concierge.ipynb`** — a repo onboarding helper (not a forecaster) that answers questions about how the codebase works using a committed public-`main` knowledge digest. From the repository root: `uv run adk run implementations/getting_started/concierge_agent` (or `uv run adk web implementations/getting_started/concierge_agent` for the browser UI — [guide 5](../guides/05-access-adk-web-via-ssh-tunnel.md)). See [`getting_started/README.md`](getting_started/README.md) and the notebook for full usage.

---

## Relationship to `aieng-forecasting`

- `aieng-forecasting` (`aieng.forecasting`) owns reusable infrastructure and reusable reference predictors under `aieng.forecasting.methods`.
- `implementations/` owns use-case material: walkthrough notebooks, experiment-specific helper modules, plotting/analysis code, and task-specific framing.

If code becomes broadly reusable across use cases, promote it into `aieng-forecasting`.

---

## Adding a new use case

1. Create `implementations/<use-case>/`.
2. Add a `README.md` describing the task, the data, and what the notebooks cover.
3. Add YAML specs under `implementations/<use-case>/specs/`.
4. Start with notebooks as the primary user surface.
5. If notebook code becomes bulky or repeated, extract small helper modules into that use-case directory.
6. Add tests under `implementations/tests/<use-case>/` for non-trivial helper logic.
7. Promote code into `aieng-forecasting` once it is clearly reusable across more than one use case.

For architecture principles and cross-cutting extension ideas, see `planning-docs/roadmap.md`.
Original file line number Diff line number Diff line change
@@ -0,0 +1,76 @@
README.md
pyproject.toml
agentic_forecasting_implementations.egg-info/PKG-INFO
agentic_forecasting_implementations.egg-info/SOURCES.txt
agentic_forecasting_implementations.egg-info/dependency_links.txt
agentic_forecasting_implementations.egg-info/requires.txt
agentic_forecasting_implementations.egg-info/top_level.txt
boc_rate_decisions/__init__.py
boc_rate_decisions/analysis.py
boc_rate_decisions/data.py
boc_rate_decisions/plots.py
boc_rate_decisions/press_releases.py
boc_rate_decisions/rationale_eval.py
boc_rate_decisions/analyst_agent/__init__.py
boc_rate_decisions/analyst_agent/agent.py
boc_rate_decisions/predictors/__init__.py
boc_rate_decisions/predictors/llmp_binary.py
boc_rate_decisions/predictors/llmp_direction.py
boc_rate_decisions/predictors/logistic_baseline.py
boc_rate_decisions/starter_agent/__init__.py
boc_rate_decisions/starter_agent/agent.py
energy_oil_forecasting/__init__.py
energy_oil_forecasting/analysis.py
energy_oil_forecasting/data.py
energy_oil_forecasting/paths.py
energy_oil_forecasting/prophet_baseline.py
energy_oil_forecasting/tasks.py
energy_oil_forecasting/viz.py
energy_oil_forecasting/adaptive_agent/__init__.py
energy_oil_forecasting/adaptive_agent/agent.py
energy_oil_forecasting/adaptive_agent/skill_state.py
energy_oil_forecasting/adaptive_agent/skill_tools.py
energy_oil_forecasting/adaptive_agent/curriculum/snapshot_utils.py
energy_oil_forecasting/analyst_agent/__init__.py
energy_oil_forecasting/analyst_agent/agent.py
energy_oil_forecasting/starter_agent/__init__.py
energy_oil_forecasting/starter_agent/agent.py
energy_oil_forecasting/starter_agent/tools.py
food_price_forecasting/__init__.py
food_price_forecasting/analysis.py
food_price_forecasting/data.py
food_price_forecasting/plots.py
food_price_forecasting/reports.py
food_price_forecasting/smoke_report.py
food_price_forecasting/predictors/__init__.py
food_price_forecasting/predictors/llmp_quantile_grid.py
food_price_forecasting/predictors/llmp_sampled_trajectory.py
food_price_forecasting/starter_agent/__init__.py
food_price_forecasting/starter_agent/agent.py
getting_started/__init__.py
getting_started/concierge_agent/__init__.py
getting_started/concierge_agent/agent.py
getting_started/concierge_agent/catalog.py
getting_started/concierge_agent/catalog_build.py
getting_started/concierge_agent/knowledge.py
manufacturing_stress_forecasting/__init__.py
manufacturing_stress_forecasting/data.py
manufacturing_stress_forecasting/features.py
manufacturing_stress_forecasting/run_agent_backtest.py
manufacturing_stress_forecasting/run_agent_prediction.py
manufacturing_stress_forecasting/run_smoke.py
manufacturing_stress_forecasting/targets.py
manufacturing_stress_forecasting/analyst_agent/__init__.py
manufacturing_stress_forecasting/analyst_agent/agent.py
manufacturing_stress_forecasting/predictors/__init__.py
manufacturing_stress_forecasting/predictors/logistic.py
manufacturing_stress_forecasting/predictors/xgboost.py
sp500_forecasting/__init__.py
sp500_forecasting/analysis.py
sp500_forecasting/data.py
sp500_forecasting/leaderboard.py
sp500_forecasting/plots.py
sp500_forecasting/predictors/__init__.py
sp500_forecasting/predictors/llmp_sampled_trajectory.py
sp500_forecasting/starter_agent/__init__.py
sp500_forecasting/starter_agent/agent.py
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@

Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
aieng-forecasting[agentic,documents,llm,numerical]
beautifulsoup4<5,>=4.12
xgboost<4,>=2.1
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
boc_rate_decisions
energy_oil_forecasting
food_price_forecasting
getting_started
manufacturing_stress_forecasting
sp500_forecasting
118 changes: 118 additions & 0 deletions implementations/manufacturing_stress_forecasting/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,118 @@
# Manufacturing stress forecasting

This implementation asks one Track 1 question:

> Given information available at a monthly forecast origin, what is the
> probability that U.S. manufacturing will be under stress three months later?

The feature service provides trailing 1-, 3-, and 6-month IPMAN changes plus
the requested FRED and Yahoo Finance fields: `FEDFUNDS`, `YC_SPREAD`,
`CPIAUCSL`, `CPI_YOY`, `UNRATE`, `ICSA`, `VIXCLS`, `HY_SPREAD`, and 3- and
12-month returns for both `SPY` and `XLI`. FRED levels are collapsed to
monthly observations, derived fields use the documented source series, and
Yahoo returns use monthly adjusted-close prices.

## Target

A month is labelled `1` (stress) when IPMAN has declined by at least 2% over
its preceding three months; otherwise it is `0`. The threshold is a provisional
version-1 definition and should be reviewed visually before expanding the
project.

The forecast made at month `t` predicts the stress label at `t + 3 months`.
That distinction makes this forecasting rather than current-state detection.

## Predictors

- `HistoricalFrequencyPredictor`: the visible historical stress rate.
- `ManufacturingStressLogisticPredictor`: fit-at-origin logistic regression on
the IPMAN and macro variables.
- `ManufacturingStressXGBoostPredictor`: a small fit-at-origin gradient-boosted
tree classifier using the same IPMAN and macro variables and cutoff-safe
training rows.
- `manufacturing_stress_analyst`: a structured LLM predictor receiving the same
cutoff-safe IPMAN and macro signals plus recent IPMAN history and historical
base rates.

All predictors return `BinaryForecast` probabilities; backtested predictors are scored with Brier score.

## Data and cutoff assumptions
Compare XGBoost with logistic regression and historical frequency rather than judging it
in isolation, because this small monthly dataset can overfit flexible models.
`FREDAdapter` caches the required FRED series under `data/fred/`, and
`YFinanceDailyAdapter` caches `SPY` and `XLI` under `data/yfinance/`.
Yahoo refreshes request history from 1998 onward explicitly so the provider's
default recent-history window cannot replace the long-term cache.
IPMAN is conservatively treated as available one month after its reference
month. Daily rate observations are treated as available on the next business
day and collapsed to their final monthly observation. The standard FRED API
does not provide full point-in-time vintages, so historical observations may
still contain later revisions; a production study should use ALFRED vintages.

## Run

The primary interactive entry point is
[`manufacturing_stress_workbench.ipynb`](manufacturing_stress_workbench.ipynb).
Open it in VS Code or Jupyter and use its configuration cell to refresh FRED
data, run the deterministic smoke test, and explicitly opt in to the cached
LLMP backtest without using the terminal.

From the repository root, put a personal FRED key in `.env` or export it:

```bash
export FRED_API_KEY="..."
```

Populate the cache and inspect the registered series:

```bash
uv run python scripts/fetch_manufacturing_stress.py
```

Run the deterministic small backtest:

```bash
uv run --directory implementations python -m manufacturing_stress_forecasting.run_smoke
```

The output prints one mean Brier score per predictor; lower is better. The
logistic model should be compared against historical frequency, not judged in
isolation.

Run the token-limited LLMP backtest explicitly:

```bash
uv run --directory implementations python -m manufacturing_stress_forecasting.run_agent_backtest
```

This evaluates historical frequency, logistic regression, XGBoost, and the
`manufacturing_stress_analyst` agent through the same binary backtest and
Brier-score calculation. The agent run uses the default lite model, a
12-month IPMAN history, compact JSON prompts, a 384-token response cap, and
one retry per failed origin. Calendar dates are replaced by relative month
offsets in retrospective agent prompts to reduce historical-event recall.

Complete results are cached under a specification-fingerprinted directory in
`data/predictions/`, so changing the stride, horizon, dates, or warmup cannot
silently reuse an incompatible result. Incomplete runs with skipped origins
are not cached. Use `--force-refresh` to intentionally re-run every predictor.
The command also verifies that every reported model was scored on the exact
same origin and forecast-date pairs.

This remains a retrospective LLM pseudo-backtest: anonymizing dates reduces,
but cannot eliminate, the possibility that a modern model recognizes a
historical episode from its training knowledge. Use prospectively recorded
forecasts for a clean out-of-sample LLM evaluation.

Run one current forecast, including the structured agent:

```bash
uv run --directory implementations python -m manufacturing_stress_forecasting.run_agent_prediction
```

## Next steps

1. Plot IPMAN and the derived stress months; confirm or revise the 2% threshold.
2. Compare the expanded macro-panel score with the earlier IPMAN-only result.
3. Compare the cached agent backtest against the deterministic baselines only
after checking scored and skipped origin counts.
16 changes: 16 additions & 0 deletions implementations/manufacturing_stress_forecasting/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
"""Minimal IPMAN-based manufacturing-stress forecasting use case."""

from manufacturing_stress_forecasting.data import (
IPMAN_SERIES_ID,
STRESS_SERIES_ID,
build_manufacturing_stress_service,
)
from manufacturing_stress_forecasting.predictors import ManufacturingStressLogisticPredictor


__all__ = [
"IPMAN_SERIES_ID",
"STRESS_SERIES_ID",
"ManufacturingStressLogisticPredictor",
"build_manufacturing_stress_service",
]
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
"""Quantitative-only manufacturing-stress analyst agent."""

from manufacturing_stress_forecasting.analyst_agent.agent import (
ManufacturingStressPromptBuilder,
build_manufacturing_stress_agent_config,
build_manufacturing_stress_agent_predictor,
)


__all__ = [
"ManufacturingStressPromptBuilder",
"build_manufacturing_stress_agent_config",
"build_manufacturing_stress_agent_predictor",
]
Binary file not shown.
Binary file not shown.
Loading