PaperGym: Rubric-Centered Evolution for Research-Plan Generation
Ranking
Overall
86
Content
95
Popularity
65
Observed public metrics from 1 member.
Merged summary
TL;DR - PaperGym converts scientific papers into reinforcement-learning environments for training AI systems to generate research plans, using separate paper sections to derive questions and evaluation rubrics. Its rubric-centered training improves planning benchmarks while reducing criterion leakage.
- Questions are synthesized from research goals and background, while rubric criteria come from methods and experiments, limiting rewards from simple paraphrasing.
- Criterion leakage falls to 3.7%, compared with 11.90%–34.10% in existing datasets.
- Using rubrics first as privileged self-teaching context and then as GRPO rewards improves five-benchmark averages by 4.8–5.6 points across Qwen3 models.
- The released resources include the PaperGym pipeline, 20,000 training instances, and innovation and experimental-design benchmarks.
Sources (1)
PaperGym: Rubric-Centered Evolution for Research-Plan Generation
Public signals
Hugging Face upvotes 32 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - PaperGym converts scientific papers into reinforcement-learning environments for training AI systems to generate research plans, using separate paper sections to derive questions and evaluation rubrics. Its rubric-centered training improves planning benchmarks while reducing criterion leakage.
- Questions are synthesized from research goals and background, while rubric criteria come from methods and experiments, limiting rewards from simple paraphrasing.
- Criterion leakage falls to 3.7%, compared with 11.90%–34.10% in existing datasets.
- Using rubrics first as privileged self-teaching context and then as GRPO rewards improves five-benchmark averages by 4.8–5.6 points across Qwen3 models.
- The released resources include the PaperGym pipeline, 20,000 training instances, and innovation and experimental-design benchmarks.