PhysioNet / CinC Challenge 2019 · results

Predicting sepsis six hours before the clinician does

Hourly risk scoring across 40,336 ICU admissions from two independent health systems, scored on the challenge's own time-dependent clinical utility rather than on accuracy.

1,552,210 ICU hours 345 causal features 1.8% positive hours ~90% of lab cells missing external site held back until final scoring
40,336ICU admissionstwo health systems
0.819AUROC, internal testhospital A, scored once
0.802AUROC, externalhospital B, never seen
33 hmedian warningbefore clinical onset
84%caught before onsetat 4.8 false alarms each

Generated by sepsis evaluate. Every number below comes from data the model in question never saw during fitting or tuning: hyperparameters and blend weights were selected on the validation split, the hospital A test split was scored once, and hospital B was untouched until this table.

Cohort

splithoursadmissionsseptic_admissionsseptic_admission_ratepositive_hour_rate
train551,90814,2341,2530.08800.0217
val119,2223,0512680.08780.0215
test119,0853,0512690.08820.0217
external761,99520,0001,1420.05710.0141

Headline results

Hospital A — held-out test

modelaurocauprcutilitysensitivityprecisionalert_rate
clinical_rule0.64780.03710.13780.44500.04000.2416
logistic0.79210.08090.35940.54990.07380.1616
xgboost0.81350.13910.39390.68610.06100.2441
gru0.77020.06990.31220.53290.06330.1827
ensemble0.81930.10620.40960.67410.06540.2235

Hospital B — external validation, never seen in any form

modelaurocauprcutilitysensitivityprecisionalert_rate
clinical_rule0.65960.02740.01100.36420.02920.1763
logistic0.73110.05700.25660.44650.06050.1044
xgboost0.80130.06930.32750.51970.06960.1057
gru0.71580.02800.05300.55120.03100.2519
ensemble0.80150.06570.25730.63430.04600.1950

Best model by clinical utility: ensemble — test AUROC 0.819 (95% CI 0.797–0.840, patient-level cluster bootstrap), normalised utility 0.410. On hospital B the same frozen model and threshold give AUROC 0.802 and utility 0.257.

Are the differences real?

DeLong test for correlated ROC curves — the models score identical rows, so an unpaired comparison would be needlessly conservative.

championchallengerauc_aauc_bdiffzp_value
ensemblexgboost0.8190.8130.0058755.77e-07
ensemblelogistic0.8190.7920.0272123.74e-33
ensemblegru0.8190.770.049116.12.29e-58
ensembleclinical_rule0.8190.6480.17229.94.18e-196

What the search actually bought

The Optuna search ran 60 trials (20 pruned) against 3-fold grouped cross-validation, reaching a best out-of-fold utility of 0.4483. The same configuration scores 0.3939 on the held-out test split.

That +0.0544 gap is selection optimism, and it is expected: taking the maximum over 60 trials of a noisy, piecewise-constant objective selects partly for luck. It is reported rather than hidden because the CV number is the one that would be quoted if nobody had held a test set back — which is exactly why one was held back.

What this looks like at the bedside

modeldetection_ratemedian_lead_time_hfalse_alarm_rate_per_admissionfalse_alarms_per_true_detection
clinical_rule0.82936.0000.7629.502
logistic0.77336.0000.3484.659
xgboost0.81038.0000.4255.417
gru0.79935.0000.4265.512
ensemble0.83633.0000.3854.760

detection_rate counts a septic admission as caught only if the first alert lands before clinical onset; an alert raised after the team already suspected sepsis is not an early warning.

One admission at a time

A median lead time of 36 hours is a fact about 268 admissions, and it does not answer the question a clinician asks first, which is what this thing would have done in front of one patient. Below, four admissions replayed hour by hour at the frozen operating point: the risk the model produced, when it crossed the alert threshold, and when the care team actually started acting.

The cases come from the validation split, never from test. Validation was already spent on tuning, calibration and threshold selection, so nothing is lost by looking at it again. Choosing illustrative cases from the test split would quietly convert a held-out set into a presentation set. It also means the probabilities on screen are in-sample for the isotonic map, which was fitted on this split: the replay is about timing, not about calibration quality.

The headline case is the median, not the best. p014990 alerts 36 hours before onset because 36 hours is where the median of the 225 caught admissions sits. Picking the largest lead time in the cohort would have been easy and would have described nothing. Two of the four cases are failures for the same reason: p006176 alerts 3 hours after the team was already acting, which is not an early warning, and p018494 is a 25-year-old who never became septic and sat above the threshold for 56 of 58 hours. p002551 is the third kind of case worth seeing: caught, with 1 hour of warning, which is a catch by the metric and close to useless at the bedside.

Validation split, xgboost at the frozen threshold of 0.0308. Ticks under the axis mark hours when a sparse lab was drawn.

admissionwhat it showsstay (h)sepsisonset hourfirst alertwarning (h)alert hours
p014990caught, at the cohort's median lead time87yes8549+3639
p002551caught, but only just8yes76+13
p006176not caught in time50yes4851-31
p018494no sepsis, alarm anyway58no—1—56

Lab draws are marked on the same timeline as the risk. The feature-block ablation found that ordering behaviour alone reaches 97% of the full matrix's AUROC, so what the team chose to measure belongs next to the risk it produced.

Experiments

Each experiment changes exactly one thing, refits, and reports what that change was worth. Every one of them asserts its own invariants before returning, so a variant that failed to commit the mistake it was studying raises rather than publishing a plausible wrong number.

What leakage is worth, measured

Splitting ICU hours at random instead of splitting admissions inflates AUROC by +0.0275 (0.821 to 0.848). It is not an exotic bug: it is what train_test_split does to a dataframe of rows, and it puts 17,236 of 17,285 admissions on both sides of the split at once. Filling gaps from the future with a single bfill inflates AUROC by +0.0056. Everything else is held fixed across the three rows: same hyperparameters, same boosting rounds, same seed, and for variant B the same admissions in train and test as the baseline.

The more useful number is the one underneath. Clinical utility inflates by +0.0796 and +0.0478 for the two mistakes, roughly 3x and 9x the corresponding AUROC movement. Leakage does not merely make the ranking look better; it moves the whole risk distribution, so the threshold sweep finds an operating point that does not exist on honest data. A project reporting AUROC alone would see a fraction of the damage it is actually doing.

variantaurocauroc_inflationutilityutility_inflationn_train_rowsn_test_rows
honest (admission split, causal features)0.82060.00000.41480.0000503767167363
A: split ICU hours at random0.84810.02750.49440.0796503347167783
B: fill gaps from the future (bfill)0.82620.00560.46260.0478503767167363

What each feature block buys

With all 345 features the model scores AUROC 0.8206. No single block costs more than 0.0080 AUROC to remove — the most expensive is recency (hours since each channel was last measured). Yet every block scores between 0.720 and 0.782 on its own. That combination has one explanation: the matrix is enormously redundant, and the same physiology is reachable through several different encodings of it.

The sharpest number here is the ordering-only row. Using 109 features that contain no measured value whatsoever — only which channel was sampled, how recently, and how often — the model reaches AUROC 0.7943, or 97% of what the full matrix achieves. Nothing about the patient's physiology is in that subset. It is a record of what the care team chose to look at, and it is nearly as predictive as the measurements themselves.

Two readings follow. The optimistic one: missingness is signal, and the recency and intensity blocks earn their place rather than padding the matrix. The uncomfortable one: a model this dependent on ordering behaviour is partly learning clinical suspicion rather than physiology, so it would degrade wherever ordering habits differ — which is exactly what the drop from hospital A to hospital B looks like.

blockwhat it isn_featuresauroc_withoutloo_costauroc_soloutility_withoututility_loo_cost
recencyhours since each channel was last measured340.81260.00800.75990.41170.0031
clinicalSIRS, qSOFA, partial SOFA, shock index400.81720.00340.78220.4179-0.0031
rolling6h and 24h level, spread and trend1280.81750.00310.72740.4194-0.0047
intensityhow often each channel has been sampled680.81780.00280.77640.41380.0010
locfcarried-forward channel values340.81860.00200.74990.4204-0.0057
missingpanel-level ordering activity70.81970.00090.71960.4158-0.0011
deviationeach channel against the patient's own running mean340.82000.00060.73650.4181-0.0034

Case mix or degradation: decomposing the hospital B drop

At the operating point frozen on hospital A's validation split (threshold 0.519, applied unchanged everywhere), hospital A scores 0.4243 normalised utility and hospital B scores 0.2946: a gap of 0.1297 (95% CI 0.0655 to 0.1817). Hospital B's admissions are reweighted to hospital A's baseline case mix -- demographics, unit, early vitals and first-six-hour ordering volume -- by a propensity density ratio trimmed at the 1st and 99th percentiles. That weighting cuts mean covariate imbalance from 0.220 to 0.075 |SMD| and retains an effective 53% of 20,000 admissions.

Reweighted, hospital B scores 0.3120. So of the 0.1297 gap, +0.0174 is case mix (95% CI -0.0041 to +0.0419) and +0.1123 is what remains after adjustment (95% CI +0.0450 to +0.1724) — most of the gap survives the adjustment, and the case-mix component is not separable from zero. Both intervals come from a 200-replicate patient-level cluster bootstrap that refits the propensity model on every replicate, so the uncertainty of the reweighting is inside the interval rather than assumed away.

That component is not robust to how the adjustment is specified, which is itself the result. Matching to hospital A's test split rather than its training split gives +0.0206; using only covariates fixed at admission — dropping the six hours of early vitals and ordering volume, which are care process and plausibly mediators of the site difference being adjusted for — gives +0.0056. The pre-registered specification is the one reported above; these are reported beside it because a number that moves when a defensible choice moves belongs in the open as a range.

Two cautions on reading the second number as degradation. The adjustment can only correct for the 17 covariates it was given, so anything that differs between two health systems and is not in that list — treatment practice, labelling behaviour, unmeasured acuity — stays in the residual. And reweighting on baseline covariates does not force the septic rate to match: it moves from 5.7% to 6.7% against hospital A's 8.8%. The residual is an upper bound on degradation, not a measurement of it.

One candidate explanation can be ruled out without any modelling. The threshold is frozen at hospital A's, by design, because that is what deploying a model means, and a frozen threshold is exactly the kind of thing that stops working at a new site. Here it does not: re-picking the threshold on hospital B alone moves its utility from 0.2946 only to 0.3006, 5% of the gap. Hospital A's threshold is already very nearly the right threshold for hospital B, so the loss is in the scores themselves, not in where the line is drawn. The MICU-to-SICU experiment in the next section runs the same check across a different boundary and gets the opposite answer, which is the useful pairing: a model can cross a boundary intact while its operating point does not, and it can carry its operating point across intact and still lose most of its value.

cohortn_admissionsseptic_admissions_pctaurocutilityutility_gap_vs_A
hospital A (test)30518.81680.82290.42430.0000
hospital B (as observed)200005.71000.78680.2946-0.1297
hospital B (reweighted to A's case mix)200006.72760.78560.3120-0.1123

MICU to SICU: a shift this cohort can actually support

Trained on 2,727 medical ICU admissions and evaluated at a threshold frozen on held-out MICU patients (0.445), the model scores AUROC 0.8046 (95% CI 0.7794 to 0.8348) on MICU admissions it has not seen and 0.7810 (0.7477 to 0.8183) on surgical ICU admissions. The difference is +0.0236 with a 95% interval of -0.0133 to +0.0635, from an independent two-sample cluster bootstrap — DeLong does not apply here, because the two cohorts are different patients rather than two scores on the same rows. So the interval on that difference spans zero, so this experiment does not establish a ranking loss across the unit boundary. That is not the same as showing there is none: with 4,650 SICU admissions it is a statement about what this test could detect.

Clinical utility is a different story, and it is measured on a held-out half of each bucket so that no threshold is graded on the admissions it was chosen from. At the MICU threshold the model scores 0.4687 at home against 0.0884 in the SICU (95% CI -0.0592 to 0.1883). Choosing a threshold on a separate half of the SICU and scoring it here recovers +0.3193 to 0.4077 (95% CI on the gain: +0.2433 to +0.4439, which excludes zero).

Two cautions on reading that as a free fix. Choosing a local threshold requires this unit's own labelled outcomes — utility is optimised against them — so it costs local outcome data, not merely a sweep. And the normalised utility score has a cohort-specific, outcome-informed denominator, so part of the movement between units is the metric responding to a septic rate of 10.8% against 4.0%, not the model behaving differently. What the score distribution actually does across the boundary is directly visible: mean predicted risk 0.3415 in the MICU against 0.3352 in the SICU, alerting on 76% against 75% of admissions at the same threshold.

The third bucket is the one it would be convenient to omit: 8,090 admissions where neither unit indicator was recorded — more than either named unit — scoring AUROC 0.7849 at 10.4% prevalence. Dropping them would have made the transfer claim cleaner and the cohort unrepresentative of the hospital it came from.

cohortn_admissionsn_eval_admissionsn_hoursseptic_admissions_pctaurocauroc_loauroc_himean_scorealert_hours_pctalerted_admissions_pctlocal_thresholdutility_frozenutility_frozen_loutility_frozen_hiutility_localutility_local_loutility_local_higain_logain_hi
MICU (held out, same unit)11375694341710.81790.80460.77940.83480.341527.169175.72560.43430.46870.35760.55970.46740.35820.5578-0.01620.0115
SICU (surgical, the transfer target)465023251698214.00000.78100.74770.81830.335221.613374.58060.59930.0884-0.05920.18830.40770.31960.50020.24330.4439
unit not recorded8090404532782710.43260.78490.77270.79820.364532.168275.57480.51520.37730.33470.42200.38100.34550.4212-0.01630.0229

Is ordering behaviour what fails to transfer?

The feature-block ablation showed that 109 features containing no measured value reach most of the full matrix's AUROC. The obvious reading is that the model is partly learning clinical suspicion, which would explain why it transfers poorly to a hospital with different charting habits. That reading is a story that fits the numbers, so it is tested here rather than repeated.

Two models, identical but for the 109 withheld features. With everything, AUROC falls 0.8229 to 0.7868 across the hospital boundary, a gap of 0.0361. Without ordering behaviour, it falls 0.8174 to 0.7921, a gap of 0.0254.

The mechanism story survives its own test. Withholding ordering behaviour shrinks the transfer gap by 0.0107 AUROC (95% CI +0.0004 to +0.0205), so a meaningful part of what fails to cross the boundary is the model's dependence on what the care team chose to measure. The interval is tight against zero, so this is a real effect rather than a large one: it establishes the direction, not the size.

Either way the price of removing them is visible in the first column: hospital A performance drops from 0.8229 to 0.8174. Ordering behaviour is not noise to be regularised away — it is real signal about a real thing, which is precisely why its portability is worth knowing.

modeln_featuresauroc_hospital_aauroc_hospital_btransfer_gaputility_hospital_autility_hospital_b
everything3450.82290.78680.03610.42420.3006
no ordering behaviour2360.81740.79210.02540.41590.3072

Statistical analysis

245 of 331 features separate septic from non-septic admissions at FDR < 0.05 (Benjamini-Hochberg over Welch tests). Tests are run on one summary value per admission, not per ICU hour: hours within a stay are strongly autocorrelated, and treating them as independent inflates significance by orders of magnitude.

featurehedges_gunivariate_aucp_welchq_welch
O2Sat_n_obs0.97810.52663.13e-378.582e-36
MAP_n_obs0.97540.52361.69e-364.054e-35
FiO2_n_obs0.97460.65194.809e-575.274e-55
HR_n_obs0.96880.52049.27e-362.033e-34
Resp_n_obs0.96860.52421.725e-364.054e-35
Bilirubin_direct_hours_since0.94660.58279.552e-050.0001849
iculos0.89870.51743.578e-346.196e-33
Magnesium_n_obs0.82060.60362.4e-417.896e-40
Phosphate_n_obs0.81110.62855.473e-513.001e-49
Calcium_n_obs0.81090.6317.545e-513.546e-49
cum_labs0.79090.58613.708e-391.109e-37
Fibrinogen_hours_since0.74680.53262.775e-076.968e-07

Missingness is not noise

Whether a lab was ordered — ignoring its value entirely — separates the groups on its own. That is why the feature matrix carries recency and ordering-intensity channels rather than imputing them away.

channelorder_rate_septicorder_rate_controlauc_of_orderingq_value
Lactate0.076060.029610.66633.868e-109
BUN0.10270.080880.60261.21e-32
Creatinine0.078050.066040.59983.998e-31
WBC0.091030.07510.57675.388e-19
PTT0.06060.047150.56351.274e-13
Platelets0.073470.065480.54861.425e-08
Fibrinogen0.012360.0068030.53415.273e-12
TroponinI0.00057940.0013150.49510.03014

Logistic regression, interpreted

Refit without penalty on the collinearity-pruned matrix, because Wald standard errors do not apply to penalised coefficients. Odds ratios are per standard deviation of the feature.

featureodds_ratio_per_sdor_ci_lowor_ci_highp_valueq_value
O2Sat_n_obs1.791.6531.9393.362e-462.69e-44
unit_sicu0.66780.62630.71215.968e-352.387e-33
SBP_obs_rate0.76670.73420.80073.532e-339.42e-32
FiO2_obs_rate1.3211.2441.4039.481e-201.896e-18
hosp_adm_time0.91810.8950.94185.275e-118.44e-10
hours_since_any_lab0.80030.74620.85834.352e-105.803e-09
FiO2_n_obs0.82660.77040.88681.133e-071.231e-06
Resp_min61.1631.11.231.231e-071.231e-06
FiO2_dev1.081.0471.1151.274e-061.097e-05
sirs_temp1.1161.0671.1671.371e-061.097e-05
modified_shock_index1.161.0891.2354.347e-063.162e-05
Lactate_n_obs1.1181.0641.1761.083e-057.22e-05

An elastic-net fit (L1-dominant) reduces the matrix to 45 non-zero coefficients — a candidate short list for a bedside score.

featurecoefficientodds_ratio_per_sd
Resp_n_obs0.24151.2732
unit_sicu-0.22060.8020
SBP_obs_rate-0.18220.8334
hours_since_any_lab-0.16800.8454
modified_shock_index0.12731.1358
sirs_temp0.12121.1288
FiO2_obs_rate0.11471.1215
Lactate_obs_rate0.11101.1174
sofa_renal0.10571.1114
FiO2_n_obs0.09731.1022

Model behaviour

Calibration

Class weighting buys ranking quality at the cost of the probability scale. Isotonic maps were fitted on validation and are scored here on the held-out test split — reporting them on validation would be circular, since in-sample isotonic ECE is 0 by construction. Discrimination is essentially unchanged while Brier and ECE move substantially.

modelbrier_rawbrier_calibratedece_rawece_calibratedauroc_rawauroc_calibrated
clinical_rule0.09950.02110.20630.00370.64610.6478
logistic0.17870.02060.33300.00280.79290.7921
xgboost0.17320.01980.36310.00280.81390.8135
gru0.05630.02070.09200.00430.77090.7702

What each model contributes

Blend weights were optimised against validation utility and then frozen. marginal_contribution is what the blend loses if that member is removed — a large weight with a small ablation cost means the member is redundant.

modelweightutility_aloneutility_withoutmarginal_contribution
xgboost0.60220.40170.40560.0232
logistic0.25560.38170.41290.0159
gru0.14220.36430.42350.0053

What drives the booster

featuremean_abs_shap
FiO2_hours_since0.2521
iculos0.1769
unit_sicu0.1141
sirs_score0.0925
shock_index0.0740
log_iculos0.0711
Creatinine_locf0.0685
Lactate_hours_since0.0673
Temp_max60.0556
Temp_locf0.0525
Resp_n_obs0.0406
BUN_locf0.0352

Cross-site shift

78 of 345 features shift materially (PSI > 0.25) between the two hospitals, and admission-level sepsis prevalence falls from 8.8% to 5.7%. The external column in the headline table is the cost of that shift, measured rather than assumed.

featurepsimissing_refmissing_cmpks_stat
HCO3_hours_since4.80760.16730.97970.4002
HCO3_locf4.66440.16730.97970.2932
HCO3_n_obs4.57770.00000.00000.8124
anion_gap4.53510.16920.98050.1522
HCO3_dev4.48150.16730.97970.0722
HCO3_obs_rate4.39070.00000.00000.8137
Chloride_hours_since3.21930.16760.91890.2851
Chloride_dev3.10720.16760.91890.0814
Chloride_n_obs3.09760.00000.00000.7513
Chloride_locf3.04630.16760.91890.0731

Figures

2026-08-29T12:19:22.962191 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Univariate separation between septic and non-septic admissions, FDR-controlled.
2026-08-29T12:19:23.014796 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Features that move most between the two hospital systems.
2026-08-29T12:20:06.989307 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Hyperparameter search, scored on clinical utility rather than AUROC.
2026-08-29T12:20:06.920834 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Split gain versus TreeSHAP attribution for the tuned booster.
2026-08-29T12:19:50.734259 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
ROC and precision-recall on the held-out hospital A test set.
2026-08-29T12:19:50.887780 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Normalised utility against threshold and against alert burden.
2026-08-29T12:19:50.645348 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Reliability before and after isotonic recalibration, on the held-out test split.
2026-08-29T12:19:50.943393 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Hours of warning before clinical onset, for the admissions the model catches.
2026-08-29T12:19:50.987225 image/svg+xml Matplotlib v3.11.1, https://matplotlib.org/
Internal test versus the unseen second hospital.