🛰️ Daily AI Frontier
‹ back to 2026-09-23

深度拆解 MiMo-V2.6:1M 上下文只是表面,2.5 万条轨迹才是底牌

Industry & News LLM Agents

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 深度拆解 MiMo-V2.6:1M 上下文只是表面,2.5 万条轨迹才是底牌

Merged summary

TL;DR - Xiaomi released and open-sourced MiMo-V2.6, a multimodal model family built for long-context agent workloads. Its main advance is a unified, asynchronous reinforcement-learning system that generates 25,088 agent trajectories per update and extracts richer training signals from them.

  • MiMo-V2.6 Pro has 1.02T total/42B active parameters, while Flash has 309B total/15B active parameters; both support 1M-token context and text, image, video, and audio inputs.
  • Pro combines sliding-window and global attention, sparse MoE routing, and speculative decoding to reduce the cost of generating long, tool-using trajectories.
  • Fully asynchronous GRPO separates rollout generation, environment execution, grading, and model updates, avoiding synchronization delays across tasks with widely varying completion times.
  • A unified RL policy spans coding, general agents, vision, and cybersecurity; GRS/GAR rank trajectory quality beyond binary success, while MOPD2 reuses intermediate states from teacher and demonstration trajectories.

Sources (1)

深度拆解 MiMo-V2.6:1M 上下文只是表面,2.5 万条轨迹才是底牌

雷峰网 (AI科技评论) 2026-09-23
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:31.470385 UTC

TL;DR - Xiaomi released and open-sourced MiMo-V2.6, a multimodal model family built for long-context agent workloads. Its main advance is a unified, asynchronous reinforcement-learning system that generates 25,088 agent trajectories per update and extracts richer training signals from them.

  • MiMo-V2.6 Pro has 1.02T total/42B active parameters, while Flash has 309B total/15B active parameters; both support 1M-token context and text, image, video, and audio inputs.
  • Pro combines sliding-window and global attention, sparse MoE routing, and speculative decoding to reduce the cost of generating long, tool-using trajectories.
  • Fully asynchronous GRPO separates rollout generation, environment execution, grading, and model updates, avoiding synchronization delays across tasks with widely varying completion times.
  • A unified RL policy spans coding, general agents, vision, and cybersecurity; GRS/GAR rank trajectory quality beyond binary success, while MOPD2 reuses intermediate states from teacher and demonstration trajectories.
item →