🛰️ Daily AI Frontier
‹ back to 2026-09-26

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

arXiv cs.CV Multimodal & Generative Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu, Kin-Man Lam, Yuewen Cao 2026-09-24
Representative image for AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

TL;DR - AV-GRPO is an online diffusion reinforcement-learning framework that decouples joint audio-video optimization into modality-specific subproblems. It improves generation quality, text alignment, and audio-video synchronization while reducing training cost and clarifying reward attribution.

  • Modality-anchored rollouts disentangle audio and video learning signals while controlling sample difficulty.
  • Trajectory-locked, frozen-tower optimization updates one modality at a time to reduce compute and improve credit assignment.
  • Adaptive objectives and perturbation strengths accommodate the differing optimization dynamics of audio and video.
  • On JavisBench and VABench, AV-GRPO outperforms LTX-2.3 with both LoRA and full fine-tuning; the accompanying 5DAV dataset supports training across five decoupled dimensions.

view merged work →