imread-benchmark is a reproducible framework for measuring JPEG decoder and
PyTorch DataLoader supply throughput. It is built around immutable dataset
packages, deterministic experiment plans, fresh-process runs, and validated
schema-2 evidence bundles.
The repository no longer ships historical result JSON or a second execution path. New paper tables and figures must be generated from committed schema-2 bundles through the canonical publication layer.
Preprint: Choosing a JPEG Decoder for PyTorch DataLoaders: Workload-Specific Throughput on Four CPUs.
The paper fixes num_workers=8 and compares each decoder with Pillow on the
same CPU and JPEG workload. It reports camera originals and social-media
processed JPEGs separately. The full worker sweep and absolute images-per-second
tables are in the appendix, so the main result stays a speed comparison rather
than a training-time claim.
Two protocols are currently supported:
decode-memory: decode already-resident JPEG bytes into a fully materialized, C-contiguous(H, W, 3)RGBuint8NumPy array;loader-supply: traverse a real PyTorchDataLoaderover resident JPEG bytes, including worker scheduling, batching, queues, and delivery to the consumer process.
Neither protocol measures storage download, archive verification, model training, GPU transfer, or augmentation. Claims about epoch time require a separate end-to-end experiment.
Every timed configuration runs in a fresh subprocess. Validation and warmup are outside the timer. Pillow explicitly materializes the image, converts to RGB, and copies it into an owned NumPy array before returning.
A benchmark campaign pins:
- exact JPEG bytes through
package_id,manifest_id, and ordered item IDs; - a pre-timing operational or common support set;
- decoder threads, DataLoader workers, batch size, prefetch, persistence, and multiprocessing start method;
- lock-backed environment and stable platform descriptors;
- source revision, randomized repetition block, and run position.
Each completed run is an immutable directory containing raw samples, phase
events, runtime worker probes, full provenance, derived statistics, payload
hashes, and a final COMMITTED.json. A result is invisible until the marker
and every checksum validate.
The article uses selected scenes from the Forchheim Image Database (FODB):
fodb-native: original camera JPEGs, providing the large-resolution regime;fodb-mixed: the same matched scenes and devices after Facebook, Instagram, Telegram, Twitter, and WhatsApp processing, providing a realistic mixture of resolutions, quantization tables, compression ratios, and metadata.
These are workload comparisons, not a causal estimate of resolution or JPEG
quality. Encoder quality is generally unavailable; quality_estimate is only
an estimator derived from quantization tables. See
Experiment design for the core matrix and the
controlled resolution × quality ablation needed for a causal mechanism claim.
The ablation builder and exact interpretation rules are documented in
Controlled resolution and JPEG-quality ablation.
Install uv and sync the locked development
environment:
uv sync --frozen --group dev --extra mainstream
uv run pytest -qThese commands reproduce the committed dependency freeze. Before a new paper
campaign, update to the latest stable compatible set once with uv lock --upgrade, run the full gate, and commit the new lock before any smoke. Do not
upgrade between machines. See Experiment design.
List decoder capability contracts:
uv run imread-benchmark list-decodersAfter downloading the FODB ZIP parts, build the selected native and mixed workloads. The builder extracts only selected complete scenes, verifies ZIP CRC values, records JPEG descriptors, hard-links the two local views, and creates one deduplicated uncompressed tar package.
uv run imread-benchmark dataset fodb-package \
--archive ~/data/fodb-part01.zip \
--archive ~/data/fodb-part02.zip \
--archive ~/data/fodb-part03.zip \
--output-root ~/data/fodb-benchmark \
--scene-count 12 \
--seed 20260729Upload the returned descriptor to a private GCS prefix:
uv run imread-benchmark dataset publish \
~/data/fodb-benchmark/packages/<package-id>/package.json \
--store gs://YOUR_BUCKET/imread \
--prefix datasetsThe command returns the remote descriptor object key used by local and cloud materializers. Dataset redistribution rights are not implied; keep the bucket private and follow FODB's terms.
For the separate causal ablation, start from a pinned lossless PNG source set and generate every matched factor cell with one command:
uv run imread-benchmark dataset controlled-package \
--source-dir /data/pinned-lossless-png \
--output-root /data/controlled-jpeg \
--source-name SOURCE_DATASET_NAME \
--source-release SOURCE_DATASET_RELEASE \
--source-license SOURCE_DATASET_LICENSE \
--long-edge 512 --long-edge 1024 --long-edge 2048 \
--quality 50 --quality 75 --quality 90 --quality 95 \
--include-native \
--subsampling 4:2:0 \
--compressed-byte-limit 2147483648This produces 16 workloads with identical source membership and order. Generate
their plans from
examples/controlled-ablation.template.yaml
with plan instantiate; do not interpret FODB's naturally confounded
size/quality mixture as this controlled effect.
An experiment plan is schema 2 YAML and must pin the IDs printed in
package.json. Example protocol profiles:
matrix:
decoders:
pillow: {threads: [default]}
opencv: {threads: [default, 1]}
simplejpeg: {threads: [default]}
torchvision: {threads: [default, 1]}
protocols:
decode-memory: {}
loader-supply:
worker_profiles:
- workers: [0]
batch_size: 1
- workers: [2, 4, 8]
batch_size: 1
prefetch_factor: 1
persistent_workers: true
multiprocessing_start_method: spawn
execution:
per_run_subprocess: true
run_timeout_seconds: 3600
checkpoint_each_run: true
maximum_memory_fraction: 0.6Generate one pinned, validated plan for each selected workload. Omit
--workload to generate plans for every workload in the package:
uv run imread-benchmark plan instantiate \
examples/fodb-experiment.template.yaml \
--package-descriptor /data/packages/<package-id>/package.json \
--output-dir plans/fodb \
--workload fodb-native \
--workload fodb-mixedThe command prints each plan ID and its run count per platform. Inspect or materialize the deterministic randomized matrix when needed:
uv run imread-benchmark plan validate plans/fodb/fodb-native.yaml \
--package-descriptor /data/packages/<package-id>/package.json
uv run imread-benchmark plan expand plans/fodb/fodb-native.yaml \
--package-descriptor /data/packages/<package-id>/package.json \
--output expanded-plan.jsonThe complete five-block FODB matrix is available as
examples/fodb-experiment.template.yaml.
The generated workload plans are reused unchanged across CPU platforms;
captured platform identity remains part of every run key.
Before launching that matrix, instantiate
examples/fodb-smoke.template.yaml. It
contains the preregistered nine-run Pillow/OpenCV gate: both protocols, workers
0 and 2, one repetition block.
Capture platform provenance and provision the frozen worker environment:
REVISION=$(git rev-parse HEAD)
uv run imread-benchmark platform capture \
--output artifacts/platform.json \
--machine-type local \
--location local
uv run imread-benchmark environment provision \
--group mainstream \
--runner-revision "$REVISION" \
--project-root . \
--cache-root .cache/environmentsThen pass the emitted environment descriptor and Python path to the campaign:
<environment-python> -m imread_benchmark.cli campaign run experiment.yaml \
--package-descriptor /data/packages/<package-id>/package.json \
--environment-descriptor <environment.json> \
--platform-descriptor artifacts/platform.json \
--artifact-root artifacts \
--attempts-root attempts \
--runner-revision "$REVISION" \
--worker-python <environment-python>The launcher uploads a content-addressed source snapshot and plan, creates an
ephemeral VM, materializes the dataset from one sequential GCS object, restores
or builds a content-addressed frozen environment, and checkpoints every completed bundle. DONE
or FAILED is written last; the VM then deletes itself unless failure retention
was explicitly requested.
./gcp/run.sh \
--plan experiment.yaml \
--dataset-store gs://YOUR_BUCKET/imread \
--dataset-descriptor datasets/<package-id>/package.json \
--results-store gs://YOUR_BUCKET/imread-results \
--environment-store gs://YOUR_BUCKET/imread-cache \
--machine-type c4-standard-16 \
--groups mainstream \
--no-waitStarting another VM with the same source, plan, package, environment, and platform pulls valid committed bundles first and launches only missing runs. See GCP campaigns.
uv run imread-benchmark artifacts validate artifacts
uv run imread-benchmark publish publication.yaml \
--artifact-root artifacts \
--output-dir generated
uv run imread-benchmark publish publication.yaml \
--artifact-root artifacts \
--output-dir generated \
--checkPublication output includes raw sample values, repetition-level configuration groups, and a provenance sidecar with all bundle IDs, filters, claim scope, generator revision, and output hash. Training claims are rejected for decoder-only or loader-only evidence.
imread_benchmark/
datasets/ package building, FODB selection, GCS materialization
environments/ descriptor, frozen provisioner, remote tar.zst cache
plans/ schema-2 plan loading and deterministic expansion
support/ pre-timing support audits and pinned intersections
execution/ campaign coordinator, one-run workers, attempts
artifacts/ atomic bundles and remote commit protocol
analysis/ canonical loader, statistics, claim gate, publication
decoders/ entry-point adapters and capability contracts
Run the complete local gate before submitting changes:
uv run pytest -q
uv run pre-commit run --all-filesSee CONTRIBUTING.md for decoder adapter requirements.
@misc{iglovikov2026choosingjpegdecoder,
title={Choosing a JPEG Decoder for PyTorch DataLoaders: Workload-Specific Throughput on Four CPUs},
author={Vladimir Iglovikov and Dmitry Kosarevsky},
year={2026},
eprint={2605.08731},
archivePrefix={arXiv},
primaryClass={cs.PF},
url={https://arxiv.org/abs/2605.08731}
}