Conformal Policy Learning with Distribution-Free Safety Guarantees
stat.ME, cs.LG, econ.EM, math.ST, stat.ML, stat.TH
Submitted: 2026-09-15
Updated: 2026-09-15
Code: https://github.com/ying531/conformal-policy-learning
License: http://creativecommons.org/licenses/by/4.0/
The gist: Policy learning aims to determine who should be treated based on individual characteristics.
Terminology
Abstract
Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of ``do no harm.'' In this paper, we propose conformal policy learning (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs.
Sources
- Optimized Conformal Selection: Powerful Selective Inference After Conformity Score Optimization
- Testing for Outliers with Conformal p-values
- Doubly Robust Policy Evaluation and Learning
- Model-free selective inference under covariate shift via weighted conformal p-values
- Cross-Balancing for Data-Informed Design and Efficient Analysis of Observational Studies
- Quantifying Individual Risk for Binary Outcomes
- Safe Individualized Treatment Rules with Controllable Harm Rates
- Safe Policy Learning under Regression Discontinuity Designs with Multiple Cutoffs
Related papers
- Doubly robust inference via calibration
- Bayesian Empirical Bayes: Simultaneous Inference from Probabilistic Symmetries
- Flexible Nonparametric Inference for Causal Effects under the Front-Door Model
- Deployment of AI-Assisted Interventions: Capacity Constraints and Noisy Compliance
- A Survey on Archetypal Analysis
- Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States