Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Ranking
Overall
87
Content
95
Popularity
68
Observed public metrics from 1 member.
Merged summary
TL;DR - Chimera is a hybrid diffusion transformer designed for compute-efficient, long-context image and video generation. Its linear attention, sparse MoE, and scaling recipe enable substantial efficiency gains and zero-shot video-length extrapolation.
- Combines O(N) Kimi Delta Attention, interleaved global latent attention, and modality-aware local convolutions.
- HeteroP transfers hyperparameters across heterogeneous modules to derive Chinchilla-style compute-optimal scaling laws.
- The 11B-parameter model activates 2B parameters and achieves up to 7.3Ă— the compute efficiency of a matched full-attention baseline.
- It extrapolates from 5-second training clips to 30-second videos, with 6.5% FID degradation in the final five seconds.
Sources (1)
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Public signals
Hugging Face upvotes 22
TL;DR - Chimera is a hybrid diffusion transformer designed for compute-efficient, long-context image and video generation. Its linear attention, sparse MoE, and scaling recipe enable substantial efficiency gains and zero-shot video-length extrapolation.
- Combines O(N) Kimi Delta Attention, interleaved global latent attention, and modality-aware local convolutions.
- HeteroP transfers hyperparameters across heterogeneous modules to derive Chinchilla-style compute-optimal scaling laws.
- The 11B-parameter model activates 2B parameters and achieves up to 7.3Ă— the compute efficiency of a matched full-attention baseline.
- It extrapolates from 5-second training clips to 30-second videos, with 6.5% FID degradation in the final five seconds.