Clinical prediction · CDC NVSS Natality · TRIPOD+AI 2024

Antenatal prediction of preterm birth, done to reporting standard

A reproducible, TRIPOD+AI-reported risk model built from information available early in pregnancy — with an enforced no-leakage boundary, temporal external validation, calibration, decision-curve analysis, SHAP explainability, and race/ethnicity + maternal-age fairness auditing.

Leakage guard enforced Temporal external validation Calibration + DCA SHAP explainability Subgroup fairness 22 tests passing
Purpose-built methodological sample by Moocean Studio. Numbers below come from a schema-accurate synthetic version of the CDC natality file so the whole pipeline runs offline; they are illustrative. The real, open CDC data path is a documented one-liner that changes no downstream code. View the code on GitHub →

Headline result — external temporal validation

Endpoint: preterm birth (<37 weeks). Develop on 2015–2017 (n=120,000), validate on the later, non-overlapping 2019–2020 cohort (n=80,000). Model: isotonic-calibrated gradient boosting.

0.658
External AUC (95% CI 0.651–0.665)
1.06
Calibration slope (ideal 1.0)
0.086
Brier score
32
Antenatal-only predictors

Internal 5-fold CV AUC (0.649) and external AUC (0.658) agree closely — the internal estimate generalised. AUC ≈0.66 is an honest number: population preterm prediction from routine antenatal factors is genuinely hard.

Evidence

ROC curves, development vs external
Discrimination. ROC for development vs the external temporal cohort.
Calibration reliability diagram
Calibration. Reliability diagram; slight under-prediction externally from a secular rise in the outcome rate — surfaced, not hidden.
Decision-curve analysis
Clinical usefulness. Net benefit beats treat-all, treat-none, and the "treat if prior preterm birth" clinical rule across relevant thresholds.
SHAP feature importance
Explainability. SHAP: prior preterm birth, prenatal-care timing, WIC, plurality, BMI, age, chronic hypertension — consistent with known epidemiology.
AUC by race/ethnicity
Fairness — race/ethnicity. AUC 0.64–0.68 across groups.
AUC by maternal age band
Fairness — maternal age. AUC 0.65–0.67 across age bands.

The fairness finding

Discrimination is similar across groups, but calibration is not — under-prediction is largest in the smallest subgroups (mothers ≥40, "other" race/ethnicity). Equal discrimination with unequal calibration is a recognised fairness failure mode: a single global threshold would mis-target treatment. The pipeline reports this per subgroup instead of averaging it away.

Race/ethnicitynPreterm %AUCCalibration-in-the-large
Non-Hispanic White41,2409.40.653+0.15
Non-Hispanic Black11,53912.70.651+0.21
Hispanic18,82610.20.653+0.15
NH Asian / PI5,6099.40.652+0.13
NH AIAN85510.90.681+0.06
Other1,93111.30.642+0.32

Why the method is sound

No data leakage — enforced

A blocklist enumerates every delivery-time and gestational-age-derived field (birth weight, Apgar, NICU, induction, delivery route, …). An assert_no_leakage() guard runs before every model fit; leakage raises an error instead of quietly inflating AUC. Only variables knowable before labour and delivery are eligible.

Real external validity

Development and validation cohorts are non-overlapping birth years — a temporal holdout under genuine distribution shift, not a random split of a single year.

2003 birth-certificate revision handled

Field availability changed with the revised U.S. certificate (adopted through 2014). The schema drops revision-specific fields for pre-2014 analyses so no model trains on a feature that doesn't exist for part of its cohort.