The authors develop CLARK, an interpretable longitudinal extension of βlatest-valueβ kidney failure risk equations by adding engineered features derived from repeated lab trajectories (baseline levels, variability, linear/nonlinear trends, changepoints, etc.) on top of demographic factors and last pre-index lab values. They train/test on a retrospective CHS/EHR cohort and model the event as initiation of kidney replacement therapy (KRT) after an index date, censoring at death or loss to follow-up.
In the supplied quantitative results, CLARKβs reported average precision (AP) exceeds the locally refit static baseline CSARK in both horizons. The strongest reported gains occur in the eGFR-only setting: for KRT within 2 years, CLARK AP=0.541 vs CSARK AP=0.516 (ΞAP=+0.024; 95% CI for Ξ given as +0.023 to +0.026). For KRT within 5 years, CLARK AP=0.567 vs CSARK AP=0.533 (ΞAP=+0.034; 95% CI +0.033 to +0.035).
The authors also report that calibration (via Brier score) is modestly improved or comparable when adding longitudinal features, and that c-statistics are higher for CLARK.
At KDIGO-aligned intervention thresholds, the authors report net reclassification improvement (NRI) that is generally positive for progressors (NRI_KRT>0) but often negative for non-progressors (NRI_No KRT<0), consistent with improved capture of true progressors at the cost of some false positives. They also report that these patterns are more consistent for 5-year thresholds than 2-year thresholds, with minimal/negative effects at the highest 2-year planning threshold (β₯40% 2-year risk).
In contrast, under quantile-based strategies (intervene on top k% by predicted risk), they report that gains are largely driven by progressor identification (positive NRI_KRT) with NRI_No KRT βsmall and not distinguishable from zero,β and provide an example scale: for top 10% interventions in 5-year prediction, an improvement corresponding to +1.8% (β18 additional KRT cases per 1000 true missed by static model).
A strong disconfirming result would be an external cohort validation showing CLARK (with the same feature engineering and local refits, as appropriate) does not improve discrimination/calibration and does not meaningfully improve progressor-targeting at clinically relevant operating points (KDIGO and capacity-like top-k) compared with the best available latest-value baseline for that external system.
Know what changed, what holds up, and what remains uncertain. Every Friday. No ads.