这个新生图模型有点夯:4K直出的,国产的,开源的!
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - SenseTime open-sourced SenseNova U1.5-Lite-Preview, an 8B-MoT lightweight native unified multimodal model that generates native 4K images and supports precise, instruction-driven editing. It matters because it packages understanding, reasoning, generation, and editing into one small open model aimed at real design workflows rather than one-shot image generation.
- Built on SenseTime's NEO-Unify native unified architecture, modeling language, visual semantics, and pixel generation in a single model covering understanding, reasoning, generation, and editing.
- A redesigned generation head reduces the influence of the visual-token grid and extends training to 4K resolution, cutting grid artifacts, seams, and texture breakage when zoomed in.
- Editing features include reference images, multi-image composition, localized text edits, and targeted edits via red boxes, coordinates, and markers; long/structured prompts are handled with Prompt Enhance.
- Reported benchmarks vs. prior U1: Qwen-Image-Bench 47.14 → 55.20 (with Prompt Enhance), ImgEdit-Bench 3.90 → 4.37, GEdit-Bench EN 7.47 → 8.17 and ZH 7.42 → 8.05, WeEdit overall 6.44. It is an early preview (weak spots: short prompts, small text/portrait detail, aesthetic consistency), with the full U1.5 release promised soon; weights are on GitHub, Hugging Face, and ModelScope.
Sources (1)
这个新生图模型有点夯:4K直出的,国产的,开源的!
TL;DR - SenseTime open-sourced SenseNova U1.5-Lite-Preview, an 8B-MoT lightweight native unified multimodal model that generates native 4K images and supports precise, instruction-driven editing. It matters because it packages understanding, reasoning, generation, and editing into one small open model aimed at real design workflows rather than one-shot image generation.
- Built on SenseTime's NEO-Unify native unified architecture, modeling language, visual semantics, and pixel generation in a single model covering understanding, reasoning, generation, and editing.
- A redesigned generation head reduces the influence of the visual-token grid and extends training to 4K resolution, cutting grid artifacts, seams, and texture breakage when zoomed in.
- Editing features include reference images, multi-image composition, localized text edits, and targeted edits via red boxes, coordinates, and markers; long/structured prompts are handled with Prompt Enhance.
- Reported benchmarks vs. prior U1: Qwen-Image-Bench 47.14 → 55.20 (with Prompt Enhance), ImgEdit-Bench 3.90 → 4.37, GEdit-Bench EN 7.47 → 8.17 and ZH 7.42 → 8.05, WeEdit overall 6.44. It is an early preview (weak spots: short prompts, small text/portrait detail, aesthetic consistency), with the full U1.5 release promised soon; weights are on GitHub, Hugging Face, and ModelScope.