AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
Ranking
Overall
82
Content
95
Popularity
50
Observed public metrics from 1 member.
Merged summary
TL;DR - AV-GRPO is an online diffusion reinforcement-learning framework that decouples joint audio-video optimization into modality-specific subproblems. It improves generation quality, text alignment, and audio-video synchronization while reducing training cost and clarifying reward attribution.
- Modality-anchored rollouts disentangle audio and video learning signals while controlling sample difficulty.
- Trajectory-locked, frozen-tower optimization updates one modality at a time to reduce compute and improve credit assignment.
- Adaptive objectives and perturbation strengths accommodate the differing optimization dynamics of audio and video.
- On JavisBench and VABench, AV-GRPO outperforms LTX-2.3 with both LoRA and full fine-tuning; the accompanying 5DAV dataset supports training across five decoupled dimensions.
Sources (1)
AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
Public signals
Hugging Face upvotes 3
TL;DR - AV-GRPO is an online diffusion reinforcement-learning framework that decouples joint audio-video optimization into modality-specific subproblems. It improves generation quality, text alignment, and audio-video synchronization while reducing training cost and clarifying reward attribution.
- Modality-anchored rollouts disentangle audio and video learning signals while controlling sample difficulty.
- Trajectory-locked, frozen-tower optimization updates one modality at a time to reduce compute and improve credit assignment.
- Adaptive objectives and perturbation strengths accommodate the differing optimization dynamics of audio and video.
- On JavisBench and VABench, AV-GRPO outperforms LTX-2.3 with both LoRA and full fine-tuning; the accompanying 5DAV dataset supports training across five decoupled dimensions.