🛰️ Daily AI Frontier
‹ back to 2026-07-31

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Research Multimodal & Generative

Ranking

Overall 87
Content 95
Popularity 68

Observed public metrics from 1 member.

Representative image for Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Merged summary

TL;DR - Chimera is a hybrid diffusion transformer designed for compute-efficient, long-context image and video generation. Its linear attention, sparse MoE, and scaling recipe enable substantial efficiency gains and zero-shot video-length extrapolation.

  • Combines O(N) Kimi Delta Attention, interleaved global latent attention, and modality-aware local convolutions.
  • HeteroP transfers hyperparameters across heterogeneous modules to derive Chinchilla-style compute-optimal scaling laws.
  • The 11B-parameter model activates 2B parameters and achieves up to 7.3Ă— the compute efficiency of a matched full-attention baseline.
  • It extrapolates from 5-second training clips to 30-second videos, with 6.5% FID degradation in the final five seconds.

Sources (1)

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

arXiv cs.CV Chongjian Ge, Hanwen Jiang, Tianyu Wang, Jiuxiang Gu, Yiran Xu, Ziwen Chen, Shaoteng Liu, Jing Shi, Yicong Hong, Zefan Cai, Hailin Jin, Hao Tan 2026-07-30 arXiv:2607.28611
Public signals Hugging Face upvotes 22
Providers: Hugging Face · Upvotes 22 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-30 14:28:56.739598 UTC

TL;DR - Chimera is a hybrid diffusion transformer designed for compute-efficient, long-context image and video generation. Its linear attention, sparse MoE, and scaling recipe enable substantial efficiency gains and zero-shot video-length extrapolation.

  • Combines O(N) Kimi Delta Attention, interleaved global latent attention, and modality-aware local convolutions.
  • HeteroP transfers hyperparameters across heterogeneous modules to derive Chinchilla-style compute-optimal scaling laws.
  • The 11B-parameter model activates 2B parameters and achieves up to 7.3Ă— the compute efficiency of a matched full-attention baseline.
  • It extrapolates from 5-second training clips to 30-second videos, with 6.5% FID degradation in the final five seconds.
item →