🛰️ Daily AI Frontier
‹ back to 2026-09-01

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Research LLM Agents

Ranking

Overall 86
Content 95
Popularity 65

Observed public metrics from 1 member.

Representative image for PaperGym: Rubric-Centered Evolution for Research-Plan Generation

Merged summary

TL;DR - PaperGym converts scientific papers into reinforcement-learning environments for training AI systems to generate research plans, using separate paper sections to derive questions and evaluation rubrics. Its rubric-centered training improves planning benchmarks while reducing criterion leakage.

  • Questions are synthesized from research goals and background, while rubric criteria come from methods and experiments, limiting rewards from simple paraphrasing.
  • Criterion leakage falls to 3.7%, compared with 11.90%–34.10% in existing datasets.
  • Using rubrics first as privileged self-teaching context and then as GRPO rewards improves five-benchmark averages by 4.8–5.6 points across Qwen3 models.
  • The released resources include the PaperGym pipeline, 20,000 training instances, and innovation and experimental-design benchmarks.

Sources (1)

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

arXiv cs.CL Yuhan Wang, Zhengxi Lu, Yuchen Yan, Kaitao Song, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen 2026-08-31 arXiv:2608.31119
Public signals Hugging Face upvotes 32 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 32 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:25:24.072339 UTC

TL;DR - PaperGym converts scientific papers into reinforcement-learning environments for training AI systems to generate research plans, using separate paper sections to derive questions and evaluation rubrics. Its rubric-centered training improves planning benchmarks while reducing criterion leakage.

  • Questions are synthesized from research goals and background, while rubric criteria come from methods and experiments, limiting rewards from simple paraphrasing.
  • Criterion leakage falls to 3.7%, compared with 11.90%–34.10% in existing datasets.
  • Using rubrics first as privileged self-teaching context and then as GRPO rewards improves five-benchmark averages by 4.8–5.6 points across Qwen3 models.
  • The released resources include the PaperGym pipeline, 20,000 training instances, and innovation and experimental-design benchmarks.
item →