🛰️ Daily AI Frontier
‹ back to 2026-08-05

UniWorld-Design: From Pixel Generation to Layer-Native Design

arXiv cs.CV Multimodal & Generative Zongjian Li, Zhiyuan Yan, Chenxu Bai, Chen Chen, Haoxiang Sun, Shaodong Wang, Feize Wu, Shenghai Yuan, Bin Lin, Zheyuan Liu, Yuwei Niu, Li Yuan 2026-08-04
Representative image for UniWorld-Design: From Pixel Generation to Layer-Native Design

TL;DR - UniWorld-Design generates and decomposes images as editable semantic RGBA layers rather than flat pixels, enabling more structured, agent-friendly visual creation and editing.

  • Text-to-RGBA generates standalone transparent assets directly from text.
  • Image-to-Layer produces ordered semantic layers using an image, global instructions, and per-layer prompts.
  • Complete-object layers remain usable when moved or removed, supporting decomposition and targeted extraction.
  • On Crello, it reduced per-layer RGB L1 error by 37% and improved Alpha Soft IoU by 34% relative to Qwen-Image-Layered.

view merged work →