WeAgent-MMGenEdit: A Full-Stack Recipe for Multimodal Agentic Image Generation and Editing
TL;DR - WeAgent-MMGenEdit is a full-stack framework for knowledge-intensive, agentic image generation and editing that combines multimodal retrieval, evidence management, verification, and post-training. Its compact 30B-total/3B-active policy approaches the reported performance of a 1T-parameter agent.
- WeAgent-Harness persistently manages retrieved evidence and provides dedicated tools to verify and integrate textual and visual information.
- Its training pipeline produces 23K supervised trajectories and 14.7K reinforcement-learning tasks with three-layer verifiable checklists.
- WeBench-MMGenEdit evaluates bilingual, knowledge-intensive image generation and multi-image editing.
- Two-sided SFT and RL post-training improves both the agent policy and the image-generation backend.