Stress-Testing the 4D Occupancy Forecasting Chain
NeurIPS 2026
Y.Zheng1, J.Hu1, J.Xiong2, R.Liu3, J.Zheng3,4, K.Yang1, J.Zhang1,†
1 Hunan University · 2 University of Oxford · 3 Karlsruhe Institute of Technology · 4 ETH Zurich
† Corresponding author.
News · Quick Start · Datasets · Models · Results · Citation
- 2026-10-06: Released Datasets.
- 2026-10-06: Released Code.
- 2026-10-06: Released Leaderboard.
- 2026-10-05: Paper available on arXiv.
- 2026-09-26: Accepted to NeurIPS 2026.
How do perception errors affect future occupancy forecasts? OccStress is a benchmark for evaluating robustness across the occupancy forecasting chain, from camera/LiDAR inputs to 3D occupancy states and multi-horizon 4D predictions.
3 datasets · 21 corruption families · 61 severity configurations · 6 future horizons
- Upstream track: sensor corruptions pass through a 3D occupancy estimator, exposing how perception errors propagate into future forecasts.
- Manual track: controlled occupancy-state corruptions isolate the sensitivity of the forecaster, independently of a particular upstream model.
- Temporal diagnostics: Current-only, Recent-burst and History-only test different failure regimes. A single-state position sweep separately isolates the effect of corruption timing with a fixed corruption budget.
Evaluation uses six future frames at +0.5 to +3.0 seconds, with paper averages over 1, 2 and 3 seconds. Each model retains its native observation window; traffic mirroring consistently transforms inputs, targets and motion metadata. See the evaluation specification for controls, class mappings and the distinction between historical and unified metrics.
Clone the repository, then use a Python 3.10+ environment:
git clone https://github.com/InSAI-Lab/OccStress.git
cd OccStress
python -m pip install -e .
python examples/minimal_eval.py --output-dir /tmp/occstress-demo
python tools/validate_protocol.py examples/fixtures/protocol.json --expected-anchors 1The demo writes clean.json, synthetic_error.json and summary.json under
/tmp/occstress-demo. These are synthetic scores, not model or paper results;
no dataset, checkpoint or CUDA installation is needed.
- Download the selected data and prepare its external GT and metadata.
- Choose a model below, install its isolated environment, and obtain the specified checkpoint files.
- Follow its setup/evaluation guide: start with a small clean test, then run the required protocols. Use the formal settings and resume and summary workflow.
The lightweight core does not install the forecasting models. See the full setup guide for paths and preflight checks.
| Benchmark | Source | Domain | Anchors per protocol |
|---|---|---|---|
| OccStress-nuScenes | nuScenes / Occ3D | Real-world driving | 4,519 |
| OccStress-Waymo | Waymo / Occ3D | Real-world driving | 5,978 |
| OccStress-CARLA | UniOcc CARLA subset | Simulated driving | 330 |
All three use a shared layout with 2 Hz observations and six future targets. The canonical record contains four historical states plus the current state; individual forecasters select their native input window.
Downloads: insailab/OccStress contains 75 archives, 40.95 GB compressed across the three datasets. You can download only the track or upstream source you need; the package index records archive sizes and checksums.
Example: download the Waymo manual track only
python -m pip install -e '.[download]'
python tools/list_data_packages.py --fetch --dataset waymo --track manual > selection.json
python tools/download_data.py --selection selection.json \
--output-dir /datasets/OccStress-downloadThe selector pins a dataset revision; the downloader resumes interrupted
downloads and verifies the archives. Follow the
extraction and setup instructions before evaluation.
Use --dataset nuscenes|waymo|carla to choose a dataset.
The archives contain protocols and derived occupancy assets, not original
images, point clouds, clean GT or model weights. Prepare those separately as
needed using the external data guide.
See dataset layout for the shared OccStress/ root.
Each method links to its setup and evaluation guide.
| Method | Model input |
|---|---|
| OccWorld | 5 occupancy states |
| I$^2$-World | 5 occupancy states |
| COME | 4 occupancy states |
| GenieDrive | 4 occupancy states |
| DOME | 4 occupancy states |
| SparseWorld-TC | 5 camera states |
Occupancy-input methods consume manual or upstream-exported states. SparseWorld-TC forecasts directly from cameras and has no manual-state or point-upstream track. Input windows and control policies remain model-specific; see method interfaces.
Export integrations are provided for ALOcc, STCOcc, FusionOcc, SDGOcc, CVT-Occ, EFFOcc and FlashOcc. Use the source-specific export guide for the correct dataset, configuration and checkpoint pairing.
No upstream model is required to consume the released occupancy exports. Use the prepared states directly with your chosen occupancy-input forecaster. OccFusion source and existing helpers are also retained, but are not part of the seven-source runtime checks.
Explore source-specific scores, temporal diagnostics and visualizations on the interactive Leaderboard. The paper reports the benchmark experiments; the evaluation specification documents metrics and controls. Historical paper scores and newly computed present-class scores must not be treated as interchangeable.
Code validation: small-sample A100 checks have passed for the six forecasters and seven upstream sources. These checks validate integration, not full-paper numerical reproduction or completion of every model/dataset combination. See the tested scope and limitations.
Do I need to download the original camera and LiDAR data?
Not for occupancy-input forecasting from the released states. You still need the corresponding clean GT, model metadata and forecaster checkpoints. Original sensor inputs are needed when re-exporting upstream occupancy or running a camera-direct method such as SparseWorld-TC. See external prerequisites.
Are pretrained checkpoints included?
No. Use the checkpoint guide for official links and exact file selections. Some weight identities and project-owned mirrors remain under review; source availability does not imply checkpoint availability.
Can all methods share one Python environment?
No. Several methods use incompatible versions of packages with the same import
name, including mmdet3d. Keep the core and model-specific dependencies separate
and follow the environment recipes.
- Add a forecasting method.
- Generate manual corruptions or export an upstream source.
- Inspect the protocol, metric and result workflow.
Repository structure
occstress/ Protocols, paths, corruptions, metrics and results
configs/ Evaluation contracts, suites and resource registry
EXIST/ Model-native integrations and original attribution
environments/ Isolated model installation profiles
scripts/ Dataset construction, upstream export and diagnostics
tools/ Evaluation, summaries and validation
examples/ Minimal synthetic workflow
tests/ CPU regression tests
docs/ Setup, method interfaces and attribution
See code layout for implementation details. Maintainers should follow the public packaging procedure instead of publishing the complete internal working tree or its Git history.
If you use OccStress, please cite our paper:
@inproceedings{zheng2026occstress,
title = {{OccStress}: Stress-Testing the {4D} Occupancy Forecasting Chain},
author = {Zheng, Yu and Hu, Jie and Xiong, Jiaqi and Liu, Ruiping and Zheng, Junwei and Yang, Kailun and Zhang, Jiaming},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026},
url = {https://arxiv.org/abs/2512.15621}
}Machine-readable citations: BibTeX and CFF. Please also cite the original methods and datasets used in your experiments.
OccStress builds on the original occupancy methods, Occ3D, nuScenes, Waymo, UniOcc/CARLA, RoboBEV and Robo3D. We thank their authors for making their work available to the research community.
Original OccStress code is released under MIT. Third-party code, adapted corruption operators, data and checkpoints retain their own terms; see third-party notices and the license inventory.
Join our Feishu community for discussions; we chose Feishu so new members can access the chat history.


