Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode
TL;DR - Nanbeige4.2-3B is a compact 3B-parameter agentic model designed for coding, office automation, and complex tool use. It reportedly outperforms substantially larger models on diverse agentic benchmarks while retaining competitive reasoning capabilities.
- Pretrained from scratch on 28T tokens using a Looped Transformer that reuses its layer stack to increase capacity without adding parameters.
- Training combines mixed-mode RLHF, length-controlled reasoning RL, and agentic RL with outcome and process rewards.
- Executable environments, task assets, and agent trajectories were diversified through real-world deployment and large-scale synthesis.
- Evaluations report stronger agentic performance than Qwen3.5-9B and Gemma4-12B, with OpenClaw results supporting local-assistant use.