🛰️ Daily AI Frontier
‹ back to 2026-09-01

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

arXiv cs.CL LLM Agents Yuhan Wang, Zhengxi Lu, Yuchen Yan, Kaitao Song, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen 2026-08-31
Representative image for PaperGym: Rubric-Centered Evolution for Research-Plan Generation

TL;DR - PaperGym converts scientific papers into reinforcement-learning environments for training AI systems to generate research plans, using separate paper sections to derive questions and evaluation rubrics. Its rubric-centered training improves planning benchmarks while reducing criterion leakage.

  • Questions are synthesized from research goals and background, while rubric criteria come from methods and experiments, limiting rewards from simple paraphrasing.
  • Criterion leakage falls to 3.7%, compared with 11.90%–34.10% in existing datasets.
  • Using rubrics first as privileged self-teaching context and then as GRPO rewards improves five-benchmark averages by 4.8–5.6 points across Qwen3 models.
  • The released resources include the PaperGym pipeline, 20,000 training instances, and innovation and experimental-design benchmarks.

view merged work →