Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking
Ranking
Overall
86
Content
95
Popularity
67
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper identifies “manifold drift” as a root cause of reward hacking when preference optimization pushes flow-matching models beyond the pretrained data support. It proposes ThermoDPO and a weighted variant to constrain this drift while improving generation quality.
- The theory shows that preference updates leave the pretrained manifold when terminal displacement has a nonzero component normal to it.
- ThermoDPO uses temperature-controlled anchoring on preferred samples, connecting rejection-sampling fine-tuning with FlowDPO.
- ThermoDPO-weighted addresses weakened optimization signals at low temperatures.
- It achieves a 0.899 StrictScore on the main toy benchmark and, on SD3.5-M at CFG 4.5, improves OCR by 47.5% and the four-metric average by 16.0%.
Sources (1)
Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking
Public signals
Hugging Face upvotes 2 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper identifies “manifold drift” as a root cause of reward hacking when preference optimization pushes flow-matching models beyond the pretrained data support. It proposes ThermoDPO and a weighted variant to constrain this drift while improving generation quality.
- The theory shows that preference updates leave the pretrained manifold when terminal displacement has a nonzero component normal to it.
- ThermoDPO uses temperature-controlled anchoring on preferred samples, connecting rejection-sampling fine-tuning with FlowDPO.
- ThermoDPO-weighted addresses weakened optimization signals at low temperatures.
- It achieves a 0.899 StrictScore on the main toy benchmark and, on SD3.5-M at CFG 4.5, improves OCR by 47.5% and the four-metric average by 16.0%.