Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Merged summary
TL;DR - Mage-Flow is a compact 4B model family for native-resolution image generation and instruction-based editing. Its tokenizer, diffusion architecture, and CUDA co-design improve training throughput and enable interactive high-resolution inference.
- Mage-VAE cuts tokenization cost by over 10× while retaining reconstruction quality comparable to strong public VAEs.
- Native-resolution packing and CUDA kernel fusion improve end-to-end training throughput by about 2.5×.
- The family includes Base, RL-aligned, and distilled 4-step Turbo variants for generation and editing.
- On one NVIDIA A100 at 1024² resolution, Turbo generates images in 0.59 seconds and edits them in 1.02 seconds.
Sources (1)
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
TL;DR - Mage-Flow is a compact 4B model family for native-resolution image generation and instruction-based editing. Its tokenizer, diffusion architecture, and CUDA co-design improve training throughput and enable interactive high-resolution inference.
- Mage-VAE cuts tokenization cost by over 10× while retaining reconstruction quality comparable to strong public VAEs.
- Native-resolution packing and CUDA kernel fusion improve end-to-end training throughput by about 2.5×.
- The family includes Base, RL-aligned, and distilled 4-step Turbo variants for generation and editing.
- On one NVIDIA A100 at 1024² resolution, Turbo generates images in 0.59 seconds and edits them in 1.02 seconds.