🛰️ Daily AI Frontier
‹ back to 2026-08-21

Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

arXiv cs.AI Multimodal & Generative Yansen Han, Shengyi Liao, Yuanxing Zhang, Pengfei Wan, Tao Lin 2026-08-20
Representative image for Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

TL;DR - This paper identifies “manifold drift” as a root cause of reward hacking when preference optimization pushes flow-matching models beyond the pretrained data support. It proposes ThermoDPO and a weighted variant to constrain this drift while improving generation quality.

  • The theory shows that preference updates leave the pretrained manifold when terminal displacement has a nonzero component normal to it.
  • ThermoDPO uses temperature-controlled anchoring on preferred samples, connecting rejection-sampling fine-tuning with FlowDPO.
  • ThermoDPO-weighted addresses weakened optimization signals at low temperatures.
  • It achieves a 0.899 StrictScore on the main toy benchmark and, on SD3.5-M at CFG 4.5, improves OCR by 47.5% and the four-metric average by 16.0%.

view merged work →