🛰️ Daily AI Frontier
‹ back to 2026-07-27

InnoText: A Unified Model for Visual Text Generation and Editing

arXiv cs.CV Multimodal & Generative Haowei Liu, Runze He, Jian Lu, Ao Ma, Run Ling, Ke Cao, Jiasong Feng, Wei Feng, Shuo Lu, Yexing Xu, Yun Wang, Jing Wang, Zhanjie Zhang 2026-07-24

TL;DR - InnoText is a unified diffusion-transformer model for generating and editing visual text, including small-scale and Chinese characters. It aims to improve legibility and consistency while replacing separate task-specific pipelines.

  • Uses one DiT-based framework for both visual text generation and editing.
  • Introduces font size-aware modulation and small-character augmentation to improve fine-grained text fidelity.
  • Applies a task-specific region-weighted loss for adaptive optimization.
  • Provides a bilingual English-Chinese dataset spanning varied fonts, sizes, and backgrounds.

view merged work →