🛰️ Daily AI Frontier
‹ back to 2026-07-27

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

Research LLM Agents

Merged summary

TL;DR - Nanbeige4.2-3B is a compact 3B-parameter agentic model designed for coding, office automation, and complex tool use. It reportedly outperforms substantially larger models on diverse agentic benchmarks while retaining competitive reasoning capabilities.

  • Pretrained from scratch on 28T tokens using a Looped Transformer that reuses its layer stack to increase capacity without adding parameters.
  • Training combines mixed-mode RLHF, length-controlled reasoning RL, and agentic RL with outcome and process rewards.
  • Executable environments, task assets, and agent trajectories were diversified through real-world deployment and large-scale synthesis.
  • Evaluations report stronger agentic performance than Qwen3.5-9B and Gemma4-12B, with OpenClaw results supporting local-assistant use.

Sources (1)

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

arXiv cs.AI Nanbeige Lab, :, Chen Yang, Chengrui Huang, Fufeng Lan, Hanhui Chen, Hao Zhou, Huatong Song, Jiaqi Cao, Jiaying Zhu, Jinlin Niu, Kai Wang, Lisheng Huang, Qiliang Liang, Ran Le, Ruixiang Feng, Shuang Sun, Tao Gu, Tao Zhang, Tianyu Luo, Yang Song, Yun Xing, Yuntao Wen, Ziyao Xu, Zongchao Chen, Zongqiang Li 2026-07-24 arXiv:2607.22083

TL;DR - Nanbeige4.2-3B is a compact 3B-parameter agentic model designed for coding, office automation, and complex tool use. It reportedly outperforms substantially larger models on diverse agentic benchmarks while retaining competitive reasoning capabilities.

  • Pretrained from scratch on 28T tokens using a Looped Transformer that reuses its layer stack to increase capacity without adding parameters.
  • Training combines mixed-mode RLHF, length-controlled reasoning RL, and agentic RL with outcome and process rewards.
  • Executable environments, task assets, and agent trajectories were diversified through real-world deployment and large-scale synthesis.
  • Evaluations report stronger agentic performance than Qwen3.5-9B and Gemma4-12B, with OpenClaw results supporting local-assistant use.
item →