SenseNova-U1.5: Towards Native Unified Visual Intelligence
TL;DR - SenseNova-U1.5 is an 8B mixture-of-transformers model that unifies visual understanding, reasoning, generation, and editing without separate encoders or VAEs. It suggests that multimodal understanding can transfer directly to complex visual planning and creation in an end-to-end architecture.
- Supports native resolutions up to 4K using spatially coherent patch reconstruction and curated generation and editing data.
- Uses specialized post-training experts for aesthetics, bilingual text rendering, infographics, and editing, consolidated through multi-expert on-policy distillation.
- Reported gains span image fidelity, text rendering, complex composition, multi-reference editing, instruction following, and subject and geometry preservation.
- The authors plan to open-source code for supervised fine-tuning, reinforcement learning, and on-policy distillation.