🛰️ Daily AI Frontier
‹ back to 2026-09-01

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

arXiv cs.CL LLM Agents Jiajun Shi, Siyuan Tao, Yuhao Wu, Zexuan Wang, Jingyuan Zhang, Jiaheng Liu, Xinping Lei, Xinrong Zhang, Siyuan Fang, Zhewen Tan, Tianle Cai, Junhao Fang, Jiameng Huang, Yueyang Wang, Jinkai Liu, Yuxuan Zhang, Jian Yang, Zhoujun Li, Shen Yan, Wenhao Huang, Ge Zhang 2026-08-31
Representative image for S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

TL;DR - S3Gym is an interactive benchmark testing whether LLM agents can explore their behavior, evaluate experience, and convert feedback into better decisions. Results show that self-improvement is task-dependent and can cause negative transfer rather than consistent gains.

  • Evaluates Self-Testing, Self-Judging, and Self-Improvement across seven text-based games with executable verifiers.
  • Compares raw interaction history, score-conditioned summary memory, and parameter training as experience-incorporation methods.
  • Summaries help when experience compresses into reusable strategies, while raw history works better for precise, state-dependent decisions.
  • Parameter training yields substantial gains on some tasks but unstable improvement and severe negative transfer on others.

view merged work →