Souraj Adhikary, Negar Chabi, Andre Mastmeyercs.CV cs.AI cs.LG
Distribution-free risk control adds organ-specific recall guarantees to frozen segmentation. We calibrate per-organ thresholds for an AMOS-trained nnU-Net, audit transfer to RAOS, and estimate local re-certification cost using case-level voxel false-negative rate (FNR). The AMOS control passes, but $7/12$ organs exceed $α{=}0.10$ after transfer; smaller calibration sets can mask exceedances with conservative or vacuous thresholds. Risk-Controlling Prediction Sets (RCPS) give high-probability control of population-mean risk, whereas Conformal Risk Control (CRC) gives weaker expectation control. Both require exchangeability; fixed and global thresholds give no per-organ guarantee. The Waudby--Smith--Ramdas (WSR) betting bound re-certifies six Tier-1 organs with 25 local cases, versus 30--40 for Hoeffding--Bentkus (HB). CRC needs 10--15 but has a heavier individual-case tail. No Tier-2 organ meets our illustrative precision criterion with 25 cases.
Sequential clinical decision-making often involves more than maximizing average efficacy. Clinicians may need to simultaneously optimize clinically relevant tails of the outcome distribution, control treatment-related risk, and choose among multiple treatment options. Existing quantile dynamic treatment regime (DTR) methods capture distributional features of treatment outcomes but remain largely restricted to efficacy-only objectives and binary treatments. To address these limitations, we propose Risk-Aware Quantile Dynamic Treatment Regimes (RQDTR), a unified framework that optimizes a prespecified quantile of the cumulative potential outcome while explicitly incorporating treatment-related risk. We also develop an angle-based formulation for jointly learning decision rules across multiple treatment categories. Our framework includes three interpretable subclasses: efficacy-only quantile learning, constraint-based learning with population-level risk control, and utility-based learning through a composite benefit-risk utility. Theoretically, we establish identification and oracle equivalence, Fisher consistency of the smoothed surrogate, consistency of the estimated regime, and finite-sample performance error rates. Extensive simulation studies and applications to All of Us major depressive disorder and MIMIC-III sepsis data demonstrate that RQDTR improves tail-oriented efficacy and achieves more favorable benefit-risk trade-offs than existing quantile DTR methods.
Ting Yin, Danning Li, Chen Shu +17cs.CV cs.AI cs.LG stat.AP
Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, and subtle morphologic overlap can obscure subtype distinctions. We developed CorePath, a breast-specialized multimodal pathology foundation model fine-tuned from PRISM using 7901 paired CNB whole-slide images and diagnostic reports from two centers. Evaluated across six CNB cohorts and two public breast pathology benchmarks without task-specific retraining, CorePath consistently outperformed PRISM across cancer detection, invasion assessment, and histological subtyping. It achieved weighted area under the receiver operating characteristic curves (AUCs) of 0.9526-0.9735 for five-class CNB histological subtyping across private centers. On public benchmarks, CorePath outperformed leading pathology foundation models, achieving the highest weighted AUCs of 0.7780 for BCNB invasive carcinoma subtyping, 0.8178 for BRACS lesion stratification, and 0.8252 for BRACS fine-grained classification. In report generation, CorePath reduced the overall non-breast hallucinations from 30.1% to 2.8%, demonstrating improved domain fidelity after breast-specific adaptation. CorePath-CRG further combined conformal subtype-confidence gating with Learn-Then-Test risk control to enable selective report release, subtype-level fallback, and deferral. CorePath-CRG achieved zero non-breast hallucinations among released outputs and showed the strongest overall performance in pathologist-validated LLM-based Evaluation Scores and quantitative report-generation metrics across most centers. These results demonstrate that domain-specialized foundation models with statistical risk control offer a promising approach for accurate breast CNB diagnosis and reliable report generation.