This source-folder release is a clean reimplementation of the final disclosed
method for a two-view model that predicts eight cephalometric measurements and
two three-class skeletal-pattern endpoints. It was written for publication after
the study analyses were frozen. It is not the historical internal pipeline that
trained the archived study-model weights or produced the frozen aggregate
snapshots under reference/results/.
The reimplementation provides preprocessing, partition reconstruction, model training, ensemble inference, conformal calibration, referral, evaluation, and a defined numerical verification procedure. Its validation covers unit-tested statistical formulas, model and configuration contracts, deterministic partition behavior, strict cohort and output-path checks, the reference checksum manifest, and comparison of 46 quantities from controlled frozen inputs at a fixed tolerance. This validation does not establish that every archived aggregate was regenerated by this source tree.
The public repository contains source code, final configurations, thresholds, frozen calibration and referral values, aggregate reference results, and their checksums. It contains no clinical record, photograph, identifier mapping, case-level prediction, reader-level record, or trained study-model weight.
The following inputs remain controlled under the applicable data-use agreement:
- the pseudonymized clinical cohort and frozen partition;
- case-level calibration and internal-test predictions;
- the coded operator-experience mapping;
- trained study-model weights.
Facial photographs are not distributed. The end-to-end pipeline assumes that an authorized data holder has obtained the controlled data and has lawful local access to the photographs.
The model-arm vocabulary is defined in ARM_REGISTRY.md.
Access conditions and a request template are in
DATA_ACCESS.md. No public file is a substitute for an
executed data-use agreement. Aggregate and experimental analysis inputs are
specified in ANALYSIS_INPUTS.md.
The minimum controlled bundle for numerical verification and core aggregate analysis does not automatically include trained weights, historical training histories, or prediction archives for the learning arms. Those materials are request-specific and may be supplied only under separate approval.
generated/ is the only permitted output tree. It may contain normalized face
images, pseudonymized case-level files, fitted calibration state, partitions,
and weights. Treat it as controlled, access-restrict it, and never publish or
share it. The ignore rule prevents accidental commits but is not a privacy
control. Every output destination must be new; existing files are never
overwritten.
Install the small reproduction environment and verify the controlled frozen predictions:
python -m pip install -r requirements-reproduce.txt
python reproduce.py --data-dir /path/to/authorized_bundle --operator-map /path/to/operator_experience.json
If operator_experience.json is already inside the authorized bundle, omit
--operator-map. The command recomputes and checks 46 quantities:
- 8 inter-operator ICC values and 8 single-tracing errors;
- 2 operator-experience stratum offsets;
- 2 reliability-ceiling approximations;
- 8 mean absolute errors and 8 coefficients of determination;
- 2 balanced accuracies;
- 8 split-conformal quantiles.
Inputs are read-only. An optional JSON verification report is created only when
an unused path inside generated/ is passed with --output. The default
absolute tolerance is 5e-6.
These 46 quantities are the numerical verification contract for this release.
They are a deliberately bounded subset of the reported analyses. To print every
computed value beside its frozen reference and numerical difference, add
--show-values:
python reproduce.py --data-dir /path/to/authorized_bundle --operator-map /path/to/operator_experience.json --show-values
With authorized inputs, the public source can execute the final-method
preprocessing, partition reconstruction, training, prediction, calibration,
referral, evaluation, and aggregate-analysis paths. The published experiment
nevertheless uses the controlled frozen partition, predictions, weights, and
training histories. A new run creates new artifacts under generated/; it does
not replace or claim identity with the historical artifacts.
The inference loader accepts checkpoints written by this release and the
validated historical five-member c4b checkpoint layout. Historical
compatibility is intentionally limited to that main-arm layout; this release
does not claim support for checkpoints from other historical arms or arbitrary
legacy formats. An ensemble directory must use one complete native or historical
convention; mixed filenames or checkpoint envelopes are rejected. Compatibility
does not imply that weights are public or that access to them has been approved.
Native checkpoints honor their declared fold count; the historical c4b
layout always requires five members. Inference-time changes to batch size,
worker count, mixed precision, or channels-last execution do not alter
checkpoint identity. Model arm, subset fraction, image size, fold count, seed,
scaler, state structure, and the remaining declared training-identity fields
remain validated. Native members must also retain one identical complete stored
training configuration.
The analyze command generates nine aggregate reports from the controlled
measurement table, frozen partition, and aligned calibration and internal-test
prediction archives:
bland_altman.json;age_strata_<arm>.json;shrinkage_<arm>.json;conformal_adaptivity_<arm>.json;sigma_patient_level_<arm>.json;threshold_sensitivity_<arm>.json;posthoc_<arm>.json;cost_sensitive_<arm>.json;boundary_analysis.json.
It also writes analysis_status.json, which distinguishes generated reports,
implemented utilities that require additional controlled intermediates, and
frozen artifacts whose historical generator is not included. The latter group
includes the complete reliability report, the historical c0a geometry
pipeline, and specified learning-curve outputs. See
ANALYSIS_INPUTS.md for the exact boundary.
The controlled bundle is sufficient for the point estimands above. Exact
historical Monte Carlo interval endpoints additionally require the original
pre-pseudonymization case order. The frozen boundary artifact's full-source SD
also requires pre-eligibility measurements absent from the controlled bundle.
The analyzer reports the eligible-cohort SD under an explicit name and records
both unavailable historical inputs in analysis_status.json; it never
synthesizes either input.
The clean implementation is not asserted to have produced any frozen file under
reference/results/. Those immutable files remain provenance and comparison
artifacts protected by reference/SHA256SUMS. The current release regenerates
aggregate analytical values, not the article's complete typeset tables or figure
panels.
The supported full workflow is an editable installation from this source folder using Python 3.12. The reported environment was Windows 11, Python 3.12.13, CUDA 12.8, and one NVIDIA GeForce RTX 5070 Ti:
python -m pip install -r requirements.txt
python -m pip install -e .
face2ceph --help
uv may be used as the environment manager:
uv venv --python 3.12
uv pip install -r requirements.txt
uv pip install -e .
Run these commands from the release root. The source tree supplies configs/,
schemas/, and reference/ and anchors every persistent output to
generated/. A standalone wheel or sdist installation is not part of this
release contract. Do not infer wheel portability from the editable install
metadata. The small reproduction command above is intentionally runnable from
the release root without installing the full package.
The explicit assets command is the only pipeline step that downloads files.
All assets are verified by SHA-256. Training and inference never perform an
implicit download or use an undeclared cache.
face2ceph assets face_landmarker profile_segmenter image_backbone
The stronger-backbone arm additionally requires:
face2ceph assets stronger_backbone
Downloaded third-party assets remain subject to their upstream licenses. Exact
asset provenance and notices are recorded in
THIRD_PARTY_NOTICES.md.
Preprocessing accepts one combined CSV. Required columns are case_id, age,
sex, the eight targets (ANB, Wits, SN_MP, FMA, PP_MP, Jarabak,
Y_axis, LAFH_TAFH), frontal_path, and profile_path. Classification labels
are optional; when supplied, they must match the frozen threshold scheme. Photo
paths may be absolute or relative to the input CSV. Case codes must be unique,
opaque ASCII values and must not contain direct identifiers.
The output cohort.csv retains every eligible row in its original order, removes
the raw photo paths, and adds normalized relative paths plus two canonical QC
flags:
usable: both normalized RGB views are available;analyzed:usableis true and the profile signed-distance field is available.
RGB-only arms select usable; all SDF arms select analyzed. The final cohort
contains 17,481 eligible cases, 17,392 RGB-usable cases, and 17,368 analyzed
cases. The main analyzed split contains 12,510 train/validation, 1,388
calibration, and 3,470 internal-test cases. preprocessing_status.csv is a
controlled QC result containing only case code, the two flags, and rejection
reasons; it is not an execution log.
The 24 RGB-usable cases not included in analyzed failed the fixed profile
silhouette/SDF quality-control rules; outcomes are not used in this decision.
Referral reference statistics are fitted only on observed training sex-by-age-band strata. An unseen stratum is rejected rather than silently imputed. This fail-closed rule does not arise in the declared eligible cohort, whose training split contains every declared stratum, and prevents unsupported deployment inputs from receiving an apparently valid discordance score.
The machine-readable contracts are in schemas/.
Normalize authorized photographs into a new controlled directory:
face2ceph preprocess --cohort /path/to/authorized/cohort.csv --output-dir preprocessing
For the published experiment, use the controlled frozen partition supplied with the data agreement. It contains all 17,481 eligible cases and is never modified. The partition command is available only when reconstructing assignments from the same eligible cohort in its original row order:
face2ceph partition --cohort generated/preprocessing/cohort.csv --output partitions/frozen.csv
Train the final five-fold ensemble. Each output checkpoint is created below the new run directory:
face2ceph train --cohort generated/preprocessing/cohort.csv --partition /path/to/authorized/frozen_partition.csv --image-root generated/preprocessing --output-dir runs/main --arm main
In addition to the five checkpoints, train writes
validation_history.json. It contains only fold training sizes and aggregate
validation MAE and balanced accuracies for each epoch, in the structure defined
by schemas/analysis_arm_histories.schema.json.
Generate calibration predictions, fit conformal quantiles, and fit referral state:
face2ceph predict --cohort generated/preprocessing/cohort.csv --partition /path/to/authorized/frozen_partition.csv --image-root generated/preprocessing --checkpoints generated/runs/main/checkpoints --split calibration --output predictions/main_calibration.npz
face2ceph calibrate --cohort generated/preprocessing/cohort.csv --partition /path/to/authorized/frozen_partition.csv --predictions generated/predictions/main_calibration.npz --output calibration/main.json
face2ceph fit-referral --cohort generated/preprocessing/cohort.csv --partition /path/to/authorized/frozen_partition.csv --predictions generated/predictions/main_calibration.npz --output referral/main.npz
Generate and evaluate internal-test predictions:
face2ceph predict --cohort generated/preprocessing/cohort.csv --partition /path/to/authorized/frozen_partition.csv --image-root generated/preprocessing --checkpoints generated/runs/main/checkpoints --split internal_test --output predictions/main_internal_test.npz
face2ceph evaluate --cohort generated/preprocessing/cohort.csv --partition /path/to/authorized/frozen_partition.csv --predictions generated/predictions/main_internal_test.npz --conformal generated/calibration/main.json --referral generated/referral/main.npz --output evaluation/main_internal_test.json
Prediction archives are controlled case-level derivatives. --include-features
adds an aligned feature vector for every case, and --perturbation <tag> runs a
registered deterministic input perturbation and creates another case-level
prediction archive. Both options require authorized images and checkpoints.
Their NPZ outputs must remain access-restricted under generated/ and must not
be published as aggregate results. The registered grid and downstream aggregate
scoring API are documented in
ANALYSIS_INPUTS.md.
face2ceph predict is the cohort-aligned batch path for the declared study
splits. It requires a cohort table, its matching partition, and a declared
split; it is not a standalone two-photograph interface, a clinical device, or a
deployment API. This release makes no claim of clinical utility or deployment
readiness.
Generate the nine core aggregate analyses into a new directory:
face2ceph analyze --cohort /path/to/authorized/measurements.csv --partition /path/to/authorized/frozen_partition.csv --calibration-predictions /path/to/authorized/c4b_calibration.npz --test-predictions /path/to/authorized/c4b_internal_test.npz --output-dir analysis/main --arm main
When measurements.csv already contains the validated split and fold
columns supplied in the controlled bundle, omit --partition. Otherwise pass
the matching controlled partition as shown above.
The analysis cohort must include the paired-measurement and distinct tracer
fields described in schemas/measurements.schema.json. The command reads all
controlled inputs without modifying them and writes aggregate JSON only to the
new generated/analysis/main directory.
The training, prediction, and evaluation entry points support all declared arms
in configs/arms/.
The exact mapping between semantic command names, frozen identifiers such as
c1 through c4b, QC availability populations, checkpoint criteria, and study
status is in ARM_REGISTRY.md. Pass the matching --arm
when training, predicting, or evaluating a non-main archive. Classification-only evaluation omits
conformal and referral inputs; homoscedastic evaluation omits conformal input.
The configured value 42 is a base seed, not a run identifier. Partitioning and
training use this base; model initialization and optimization for zero-based
fold k use 42 + 1000 * k, giving 42, 1042, 2042, 3042, and 4042. Learning-arm
subset sampling uses 42 + k, giving 42, 43, 44, 45, and 46.
The evaluation command uses the separate bootstrap base seed 0. Within each
classification report, overall accuracy uses that base, per-class sensitivity
uses base + 1, and balanced accuracy uses base + 2. Regression and
conformal bootstrap routines receive the evaluation base directly. The core
analyze dispatcher uses seed 0 for its age-strata and threshold-sensitivity
bootstraps and the declared seed 42 for its post hoc route. Experimental
confound, learning, and perturbation utilities expose their own seed parameters;
their function signatures, rather than the pipeline seed, define those
defaults. Each member of the five-fold main ensemble has 56,642,966 trainable
parameters.
Retraining is seeded but is not promised to be bit-wise identical across CUDA versions or hardware. Frozen-prediction quantities are checked numerically. Seeded bootstrap confidence limits additionally depend on the frozen case order; exact archived CI endpoints therefore require the original prediction order.
By default, commands emit only a final count or a concise error;
reproduce --show-values is the explicit opt-in exception. Commands create no
execution log, timestamp, telemetry, or hardware inventory. The declared
validation_history.json artifact contains aggregate validation metrics only;
it contains no case-level value or execution metadata.
Before distributing a source-folder snapshot, run the publication guard from the release root in the full environment:
python verify_release.py
The check verifies the frozen reference manifest, parses JSON/YAML/TOML files,
and rejects generated outputs, caches, model weights, clinical-image formats,
identifier-mapping names, local absolute paths, non-English text, and
logging/telemetry imports. It is intended for the clean distribution tree; a
working tree that still contains caches or controlled outputs should fail.
The reference/** -text rule in .gitattributes preserves the exact bytes
covered by reference/SHA256SUMS across Git platforms; do not replace it with
repository-wide line-ending rules or normalize the frozen reference files.
The reported reliability ceiling is a modeled approximation, not a universal
performance bound or a per-case credibility score. It models a same-radiograph
re-tracing setting; repeated image acquisition adds positioning and projection
variation and may yield a lower ceiling. The calculation assumes Gaussian
single-tracing error with constant variance and uses Phi(d / sigma), where
d is the distance from an observed label to the nearest decision boundary.
The observed label is only a proxy for the unobserved noise-free value.
For middle-band labels, using only the nearer boundary ignores the farther boundary and can only increase the approximation. The direction of error caused by the observed-label proxy is not known. Archived Gaussian and constant-error checks do not remove that proxy assumption. The 46-quantity command recomputes only the two point estimates; it does not regenerate confidence intervals or the archived assumption analyses. The estimate should not be generalized to external acquisition protocols.
CITATION.cff records the intended public repository URL and identifies the
accompanying manuscript as unpublished. A
journal, DOI, software release date, and author
identifiers are intentionally omitted until they are assigned or authoritatively
available; no placeholder value should be interpreted as bibliographic
metadata.
Source code is licensed under the MIT License. Aggregate reference data is
licensed under CC BY-NC 4.0. Controlled inputs, derived case-level artifacts,
and trained study-model weights remain governed by their data-use agreement.
See DATA_LICENSE.md, DATA_ACCESS.md,
and THIRD_PARTY_NOTICES.md.