Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
TL;DR - Chimera is a hybrid diffusion transformer designed for compute-efficient, long-context image and video generation. Its linear attention, sparse MoE, and scaling recipe enable substantial efficiency gains and zero-shot video-length extrapolation.
- Combines O(N) Kimi Delta Attention, interleaved global latent attention, and modality-aware local convolutions.
- HeteroP transfers hyperparameters across heterogeneous modules to derive Chinchilla-style compute-optimal scaling laws.
- The 11B-parameter model activates 2B parameters and achieves up to 7.3Ă— the compute efficiency of a matched full-attention baseline.
- It extrapolates from 5-second training clips to 30-second videos, with 6.5% FID degradation in the final five seconds.