Conformal Policy Learning with Distribution-Free Safety Guarantees
TL;DR - Conformal policy learning uses hypothesis tests of counterfactual harm to decide who receives treatment, providing distribution-free safety guarantees in high-stakes settings. It aims to improve welfare while explicitly limiting the probability of treating individuals who would be worse off than under control.
- Treatment is assigned by thresholding conformal p-values built from observable proxies and selective calibration.
- For randomized experiments, CPL offers finite-sample safety at a user-specified level under exchangeability, without outcome-model assumptions.
- With consistent outcome estimation, CPL is asymptotically welfare-optimal subject to the safety constraint.
- For observational studies, learn-then-balance weighting yields doubly robust safety guarantees.