PaperGym: Rubric-Centered Evolution for Research-Plan Generation
TL;DR - PaperGym converts scientific papers into reinforcement-learning environments for training AI systems to generate research plans, using separate paper sections to derive questions and evaluation rubrics. Its rubric-centered training improves planning benchmarks while reducing criterion leakage.
- Questions are synthesized from research goals and background, while rubric criteria come from methods and experiments, limiting rewards from simple paraphrasing.
- Criterion leakage falls to 3.7%, compared with 11.90%–34.10% in existing datasets.
- Using rubrics first as privileged self-teaching context and then as GRPO rewards improves five-benchmark averages by 4.8–5.6 points across Qwen3 models.
- The released resources include the PaperGym pipeline, 20,000 training instances, and innovation and experimental-design benchmarks.