A human-verified benchmark for scientific code generation
Paper · Dataset · GitHub data release · Run the benchmark · Evaluation protocol · Audit trail · Upstream SciCode
- August 2026: The SciCode-Verified paper is now on arXiv,
and the versioned
v2benchmark data is available on Hugging Face.
SciCode-Verified is an independent, human-in-the-loop correction of the SciCode test benchmark. It preserves the scientific reasoning challenge while repairing contradictions, missing conventions, incorrect frozen targets, non-deterministic tests, and other defects that can reject valid solutions.
| 262 verified corrections |
63 / 64 problems changed |
192 / 262 defects that reject correct code |
155 / 287 affected subproblems |
SciCode subproblems are cumulative, and a whole problem passes only when every scored subproblem passes. One defective specification, frozen answer, or test can therefore propagate downstream and erase an otherwise correct solution. The audit found this score-suppressing failure mode in 58 of 64 main problems.
In the manuscript's matched twelve-model-snapshot evaluation, changing only the benchmark data moves the observed frontier from 45.3–60.3% to 83.7–98.3% on subproblems, and from 9.4–26.6% to 68.8–92.2% on whole problems. The model output protocol, evaluation harness, with-background condition, pass@1 setting, and multi-environment-OR grading are held fixed.
SciCode-Verified is not an easier rewrite. Corrections state what is required to make each task well posed and its grading faithful. Weak tests are tightened; the scientific derivations and algorithms remain the model's responsibility.
Every accepted correction is recorded in a ledger, propagated from a single source of truth,
and checked against the released JSONL, HDF5, and manifest. See
CLEANING_LOG.md for the complete audit and
analysis/README.md for release statistics.
Clone the repository and install the small runtime:
git clone https://github.com/flyingwagner/scicode-verified.git
cd scicode-verified
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -U pip openai h5py numpy scipy matplotlib sympyDownload the released grading targets from Hugging Face and verify their checksum:
hf download shhu2001/SciCode-Verified test_data_cleaned.h5 \
--repo-type dataset \
--local-dir scicode_verified
md5sum scicode_verified/test_data_cleaned.h5
# 2b41a7df40ddc23ce651ec05b8ecb6f8The same file is mirrored in GitHub Releases:
gh release download data --repo flyingwagner/scicode-verified --dir scicode_verified
md5sum scicode_verified/test_data_cleaned.h5
# 2b41a7df40ddc23ce651ec05b8ecb6f8Run a one-problem smoke test:
export DEEPSEEK_API_KEY="..."
python eval_clean/run_deepseek_eval.py \
--model pro \
--dataset cleaned \
--background on \
--run smoke \
--only 58 \
--workers 1Runs are resumable. Generated code, raw responses, per-step scores, and aggregate results are
saved under eval_clean/ds_runs/. OpenRouter models work with the same entry point by setting
OPENROUTER_API_KEY and passing the raw model slug to --model.
For generation-only runs, cached re-grading, provider configuration, and original-versus-
verified comparisons, see eval_clean/README.md.
| Canonical setting | |
|---|---|
| Evaluation set | 64 main problems / 287 scored subproblems |
| Sampling | pass@1 |
| Whole-problem score | Pass only if every scored subproblem passes |
| Context | With and without background reported separately |
| Step construction | Cumulative: step k receives code from steps 1..k-1 |
| Timeout | 1,800 seconds per step and environment |
| Environment drift | Pass if correct in either pinned 2024-era or 2025-era scientific Python |
| Integrity | Dataset hashes must match scicode_verified/manifest.json |
The release excludes original test problem 2 because its specification does not determine a unique, verifiable answer. The remaining 64 problems are the exact matched set used in the before/after evaluation above.
| Path | What it contains |
|---|---|
scicode_verified/ |
Released benchmark, source problems, target patches, and manifest |
eval_clean/ |
Portable generation, scoring, and re-grading harness |
analysis/ |
Machine-readable defect and evaluation analyses |
ledger/ |
Per-change provenance and focused defect investigations |
tools/ |
Dataset assembly and release verification gate |
The large test_data_cleaned.h5 is distributed through the
Hugging Face dataset and mirrored in
GitHub Releases, not Git.
The evaluator verifies its hash before running.
SciCode-Verified is derived from SciCode and redistributed under the Apache License 2.0. The
vendored upstream notice and license are in eval_clean/vendor/. If you
use this release, please cite both the original SciCode benchmark and SciCode-Verified.
@article{hu2026scicodeverified,
title = {{SciCode-Verified}: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models},
author = {Hu, Sihan and Huang, Lyuhan and Deng, Youjin and Chen, Kun},
year = {2026},
eprint = {2608.04975},
archivePrefix = {arXiv},
primaryClass = {cs.SE},
url = {https://arxiv.org/abs/2608.04975}
}
