Clinical prediction · CDC NVSS Natality · TRIPOD+AI 2024
A reproducible, TRIPOD+AI-reported risk model built from information available early in pregnancy — with an enforced no-leakage boundary, temporal external validation, calibration, decision-curve analysis, SHAP explainability, and race/ethnicity + maternal-age fairness auditing.
Endpoint: preterm birth (<37 weeks). Develop on 2015–2017 (n=120,000), validate on the later, non-overlapping 2019–2020 cohort (n=80,000). Model: isotonic-calibrated gradient boosting.
Internal 5-fold CV AUC (0.649) and external AUC (0.658) agree closely — the internal estimate generalised. AUC ≈0.66 is an honest number: population preterm prediction from routine antenatal factors is genuinely hard.
Discrimination is similar across groups, but calibration is not — under-prediction is largest in the smallest subgroups (mothers ≥40, "other" race/ethnicity). Equal discrimination with unequal calibration is a recognised fairness failure mode: a single global threshold would mis-target treatment. The pipeline reports this per subgroup instead of averaging it away.
| Race/ethnicity | n | Preterm % | AUC | Calibration-in-the-large |
|---|---|---|---|---|
| Non-Hispanic White | 41,240 | 9.4 | 0.653 | +0.15 |
| Non-Hispanic Black | 11,539 | 12.7 | 0.651 | +0.21 |
| Hispanic | 18,826 | 10.2 | 0.653 | +0.15 |
| NH Asian / PI | 5,609 | 9.4 | 0.652 | +0.13 |
| NH AIAN | 855 | 10.9 | 0.681 | +0.06 |
| Other | 1,931 | 11.3 | 0.642 | +0.32 |
A blocklist enumerates every delivery-time and gestational-age-derived field
(birth weight, Apgar, NICU, induction, delivery route, …). An
assert_no_leakage() guard runs before every model fit; leakage
raises an error instead of quietly inflating AUC. Only variables knowable
before labour and delivery are eligible.
Development and validation cohorts are non-overlapping birth years — a temporal holdout under genuine distribution shift, not a random split of a single year.
Field availability changed with the revised U.S. certificate (adopted through 2014). The schema drops revision-specific fields for pre-2014 analyses so no model trains on a feature that doesn't exist for part of its cohort.
make setup && make test && make run regenerates every figure and table.