Sequential clinical decision-making often involves more than maximizing average efficacy. Clinicians may need to simultaneously optimize clinically relevant tails of the outcome distribution, control treatment-related risk, and choose among multiple treatment options. Existing quantile dynamic treatment regime (DTR) methods capture distributional features of treatment outcomes but remain largely restricted to efficacy-only objectives and binary treatments. To address these limitations, we propose Risk-Aware Quantile Dynamic Treatment Regimes (RQDTR), a unified framework that optimizes a prespecified quantile of the cumulative potential outcome while explicitly incorporating treatment-related risk. We also develop an angle-based formulation for jointly learning decision rules across multiple treatment categories. Our framework includes three interpretable subclasses: efficacy-only quantile learning, constraint-based learning with population-level risk control, and utility-based learning through a composite benefit-risk utility. Theoretically, we establish identification and oracle equivalence, Fisher consistency of the smoothed surrogate, consistency of the estimated regime, and finite-sample performance error rates. Extensive simulation studies and applications to All of Us major depressive disorder and MIMIC-III sepsis data demonstrate that RQDTR improves tail-oriented efficacy and achieves more favorable benefit-risk trade-offs than existing quantile DTR methods.
Emmanuel M. Rockwell, Michael R. Kosorok, Nikki L. B. Freemanstat.ME stat.ML
A central objective of precision medicine is learning optimal dynamic treatment regimes (DTRs) from data. Classification-based methods, like outcome weighted learning (OWL) for single-stage and backward OWL (BOWL) for multi-stage problems, leverage machine learning to directly learn optimal DTRs. However, these methods lack a natural way to quantify uncertainty in treatment decisions at the individual level. In this paper, we extend Bayesian OWL, a Bayesian reformulation of OWL, to the multi-stage setting. We call this method backward Bayesian outcome weighted learning (BBOWL). Like BOWL, our method directly learns an optimal DTR via backward induction, and unlike existing methods, our approach propagates uncertainty backward through the DTR learning process and provides uncertainty quantification of individualized treatment recommendations. We present a theoretical justification of BBOWL and verify its performance via a simulation study.