Conformal Policy Learning with Distribution-Free Safety Guarantees
Ranking
Overall
78
Content
95
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - Conformal policy learning uses hypothesis tests of counterfactual harm to decide who receives treatment, providing distribution-free safety guarantees in high-stakes settings. It aims to improve welfare while explicitly limiting the probability of treating individuals who would be worse off than under control.
- Treatment is assigned by thresholding conformal p-values built from observable proxies and selective calibration.
- For randomized experiments, CPL offers finite-sample safety at a user-specified level under exchangeability, without outcome-model assumptions.
- With consistent outcome estimation, CPL is asymptotically welfare-optimal subject to the safety constraint.
- For observational studies, learn-then-balance weighting yields doubly robust safety guarantees.
Sources (1)
Conformal Policy Learning with Distribution-Free Safety Guarantees
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Conformal policy learning uses hypothesis tests of counterfactual harm to decide who receives treatment, providing distribution-free safety guarantees in high-stakes settings. It aims to improve welfare while explicitly limiting the probability of treating individuals who would be worse off than under control.
- Treatment is assigned by thresholding conformal p-values built from observable proxies and selective calibration.
- For randomized experiments, CPL offers finite-sample safety at a user-specified level under exchangeability, without outcome-model assumptions.
- With consistent outcome estimation, CPL is asymptotically welfare-optimal subject to the safety constraint.
- For observational studies, learn-then-balance weighting yields doubly robust safety guarantees.