🛰️ Daily AI Frontier
‹ back to 2026-09-26

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

Research Multimodal & Generative

Ranking

Overall 82
Content 95
Popularity 50

Observed public metrics from 1 member.

Representative image for AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

Merged summary

TL;DR - AV-GRPO is an online diffusion reinforcement-learning framework that decouples joint audio-video optimization into modality-specific subproblems. It improves generation quality, text alignment, and audio-video synchronization while reducing training cost and clarifying reward attribution.

  • Modality-anchored rollouts disentangle audio and video learning signals while controlling sample difficulty.
  • Trajectory-locked, frozen-tower optimization updates one modality at a time to reduce compute and improve credit assignment.
  • Adaptive objectives and perturbation strengths accommodate the differing optimization dynamics of audio and video.
  • On JavisBench and VABench, AV-GRPO outperforms LTX-2.3 with both LoRA and full fine-tuning; the accompanying 5DAV dataset supports training across five decoupled dimensions.

Sources (1)

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

arXiv cs.CV Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu, Kin-Man Lam, Yuewen Cao 2026-09-24 arXiv:2609.29816
Public signals Hugging Face upvotes 3
Providers: Hugging Face · Upvotes 3 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:03:53.947932 UTC

TL;DR - AV-GRPO is an online diffusion reinforcement-learning framework that decouples joint audio-video optimization into modality-specific subproblems. It improves generation quality, text alignment, and audio-video synchronization while reducing training cost and clarifying reward attribution.

  • Modality-anchored rollouts disentangle audio and video learning signals while controlling sample difficulty.
  • Trajectory-locked, frozen-tower optimization updates one modality at a time to reduce compute and improve credit assignment.
  • Adaptive objectives and perturbation strengths accommodate the differing optimization dynamics of audio and video.
  • On JavisBench and VABench, AV-GRPO outperforms LTX-2.3 with both LoRA and full fine-tuning; the accompanying 5DAV dataset supports training across five decoupled dimensions.
item →