🛰️ Daily AI Frontier
‹ back to 2026-07-22

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

arXiv cs.CV Multimodal & Generative Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xun Guo, Yuxuan Luo, Jiahao Li, Wenxuan Xie, Fanyi Pu, Xiaoyi Zhang, Kaichen Zhang, Zongyu Guo, Tianci Bi, Dongnan Gui, Zhening Liu, Zimo Wen, Zihan Zheng, Senqiao Yang, Xiao Li, Jinglu Wang, Bin Li, Yan Lu 2026-07-21

TL;DR - Mage-Flow is a compact 4B model family for native-resolution image generation and instruction-based editing. Its tokenizer, diffusion architecture, and CUDA co-design improve training throughput and enable interactive high-resolution inference.

  • Mage-VAE cuts tokenization cost by over 10× while retaining reconstruction quality comparable to strong public VAEs.
  • Native-resolution packing and CUDA kernel fusion improve end-to-end training throughput by about 2.5×.
  • The family includes Base, RL-aligned, and distilled 4-step Turbo variants for generation and editing.
  • On one NVIDIA A100 at 1024² resolution, Turbo generates images in 0.59 seconds and edits them in 1.02 seconds.

view merged work →