🛰️ Daily AI Frontier
‹ back to 2026-09-08

WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

arXiv cs.CV Multimodal & Generative Hui Zhang, Zongkai Liu, Liqiang Niu, Juntao Liu, Han Li, Zhen Cao, Wenchao Chen, Chengduo Zhao, Fandong Meng 2026-09-04
Representative image for WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing

TL;DR - WeAgent-MMGenEdit is a full-stack framework for knowledge-intensive, agentic image generation and editing that combines multimodal retrieval, evidence management, verification, and post-training. Its compact 30B-total/3B-active policy approaches the reported performance of a 1T-parameter agent.

  • WeAgent-Harness persistently manages retrieved evidence and provides dedicated tools to verify and integrate textual and visual information.
  • Its training pipeline produces 23K supervised trajectories and 14.7K reinforcement-learning tasks with three-layer verifiable checklists.
  • WeBench-MMGenEdit evaluates bilingual, knowledge-intensive image generation and multi-image editing.
  • Two-sided SFT and RL post-training improves both the agent policy and the image-generation backend.

view merged work →