Forecastbench is a dynamic, contamination-free benchmark of LLM forecasting accuracy with human comparison groups, serving as a valuable proxy for general intelligence.
-
Updated
Sep 10, 2026 - Python
Forecastbench is a dynamic, contamination-free benchmark of LLM forecasting accuracy with human comparison groups, serving as a valuable proxy for general intelligence.
Local Qwen 3.5 4B yes/no forecaster. Brier 0.186 on 1,662 held-out ForecastBench questions. No API key, no cloud.
Reproducible LLM evaluation of ForecastBench rationales, with dual blinded scoring, human validation, accuracy analysis, and an interactive dashboard
Co-evolving Synthetic Intelligence Systems (SIS) with dignity, rhythm, and sacred potential.
A verified map of the AI-forecasting research area: ForecastBench, OpenForecaster, and FutureX — what each measures, what their artifacts really contain, which numbers are current, and where the open problems are.
To associate your repository with the forecastbench topic, visit your repo's landing page and select "manage topics."