MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
Merged summary
TL;DR - MeanFlowNFT adapts DiffusionNFT's forward-process reinforcement learning to MeanFlow average-velocity generators, enabling reward optimization while preserving fast few-step sampling for image and video generation.
- Bridges the mismatch between DiffusionNFT (optimizes instantaneous velocities) and MeanFlow (samples average velocities) by using the MeanFlow identity to build an induced instantaneous-velocity predictor, then applying the DiffusionNFT objective to it.
- Sampling still uses average velocity, preserving MeanFlow's few-step efficiency; the method provably inherits DiffusionNFT's strict policy-improvement guarantee.
- Outperforms prior SOTA RL-tuned few-step generators on most metrics (6 of 8 on SD3.5-M) and can beat multi-step RL-tuned diffusion using only a few steps.
- On Wan 2.1, 4-step MeanFlowNFT reaches a VBench score of 84.33, surpassing 50-step LongCat-Video RL (82.57).
Sources (1)
MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators
TL;DR - MeanFlowNFT adapts DiffusionNFT's forward-process reinforcement learning to MeanFlow average-velocity generators, enabling reward optimization while preserving fast few-step sampling for image and video generation.
- Bridges the mismatch between DiffusionNFT (optimizes instantaneous velocities) and MeanFlow (samples average velocities) by using the MeanFlow identity to build an induced instantaneous-velocity predictor, then applying the DiffusionNFT objective to it.
- Sampling still uses average velocity, preserving MeanFlow's few-step efficiency; the method provably inherits DiffusionNFT's strict policy-improvement guarantee.
- Outperforms prior SOTA RL-tuned few-step generators on most metrics (6 of 8 on SD3.5-M) and can beat multi-step RL-tuned diffusion using only a few steps.
- On Wan 2.1, 4-step MeanFlowNFT reaches a VBench score of 84.33, surpassing 50-step LongCat-Video RL (82.57).