从完成任务到自主发现,Agent后训练寻找下一个Scaling Law
TL;DR - At the 2026 Inclusion Conference, researchers and industry practitioners outlined agent post-training as a potential new scaling path: agents should progress from completing tasks to accumulating experience, discovering knowledge, and improving future models. The central challenge is jointly scaling realistic environments, reliable feedback, long-horizon infrastructure, and self-improvement loops.
- Agent scaling depends not only on model size, but also on multi-agent coordination, environment complexity, information management, and mechanisms that preserve and internalize experience.
- Real-world post-training requires interactive professional workflows, traceable outputs, explicit risk boundaries, human handoffs, and robust defenses against reward hacking.
- Long-horizon agents need task decomposition, intermediate feedback, persistent state, and jointly designed harnesses and training objectives rather than simply longer contexts or runtimes.
- Lightweight RL infrastructure such as NVIDIA's Molt and research-agent loops aim to accelerate experimentation, but reliable evaluation and validation of AI-generated discoveries remain major bottlenecks.