🛰️ Daily AI Frontier
‹ back to 2026-07-31

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

arXiv cs.CV Multimodal & Generative Chongjian Ge, Hanwen Jiang, Tianyu Wang, Jiuxiang Gu, Yiran Xu, Ziwen Chen, Shaoteng Liu, Jing Shi, Yicong Hong, Zefan Cai, Hailin Jin, Hao Tan 2026-07-30
Representative image for Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

TL;DR - Chimera is a hybrid diffusion transformer designed for compute-efficient, long-context image and video generation. Its linear attention, sparse MoE, and scaling recipe enable substantial efficiency gains and zero-shot video-length extrapolation.

  • Combines O(N) Kimi Delta Attention, interleaved global latent attention, and modality-aware local convolutions.
  • HeteroP transfers hyperparameters across heterogeneous modules to derive Chinchilla-style compute-optimal scaling laws.
  • The 11B-parameter model activates 2B parameters and achieves up to 7.3Ă— the compute efficiency of a matched full-attention baseline.
  • It extrapolates from 5-second training clips to 30-second videos, with 6.5% FID degradation in the final five seconds.

view merged work →