🛰️ Daily AI Frontier
‹ back to 2026-07-22

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

arXiv cs.CV Multimodal & Generative Rahul Sajnani, Yulia Gryaditskaya, Radomír Měch, Srinath Sridhar, Matheus Gadelha 2026-07-21

TL;DR - Appearance pointers enable precise, mask-based regional control of Diffusion Transformers using text or image references. The modality-agnostic method avoids retraining the base model and matches or exceeds modality-specific state-of-the-art approaches across reported metrics.

  • Compact tokens associate appearance cues with user-specified spatial regions.
  • A region correspondence network and spatial aggregation support multiple regional descriptions without substantial token overhead.
  • One model handles localized text and image conditioning through a unified interface.
  • The approach targets controllable materials, object identities, and spatial arrangements in image synthesis.

view merged work →