🛰️ Daily AI Frontier
‹ back to 2026-09-16

Conformal Policy Learning with Distribution-Free Safety Guarantees

arXiv stat.ME Safe Policy Learning Ying Jin, Naoki Egami 2026-09-15

TL;DR - Conformal policy learning uses hypothesis tests of counterfactual harm to decide who receives treatment, providing distribution-free safety guarantees in high-stakes settings. It aims to improve welfare while explicitly limiting the probability of treating individuals who would be worse off than under control.

  • Treatment is assigned by thresholding conformal p-values built from observable proxies and selective calibration.
  • For randomized experiments, CPL offers finite-sample safety at a user-specified level under exchangeability, without outcome-model assumptions.
  • With consistent outcome estimation, CPL is asymptotically welfare-optimal subject to the safety constraint.
  • For observational studies, learn-then-balance weighting yields doubly robust safety guarantees.

view merged work →