🛰️ Daily AI Frontier
‹ back to 2026-08-18

ClawGym II: Exploring Black-Box RL on Agent Harness

arXiv cs.CL LLM Agents Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen 2026-08-17

TL;DR - ClawGym II introduces a black-box reinforcement-learning framework for optimizing agents through opaque, complex harnesses. It enables stable, scalable, and unified training across heterogeneous agent execution systems.

  • Sandbox isolation supports large-scale concurrent rollouts.
  • A serving proxy captures model calls and reconstructs multi-turn trajectories as prefix trees for PPO or GRPO optimization.
  • Mix-harness training jointly optimizes one model through multiple harnesses.
  • Qwen3-30A3B gained 9.98 and 14.81 Pass@1 points on ClawGym-Bench through OpenClaw and Claude Code, respectively.

view merged work →