Scaling Scientific Discovery Environments for Turn-Level Agentic RL
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - SciDisco is a framework for training LLM-based scientific discovery agents inside process-verifiable environments, using turn-level RL credit assignment rather than only final-answer rewards. It matters because long-horizon data analysis has lacked environments that can check intermediate analytical progress, not just the final claim.
- SciThèque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress is checkable mid-interaction.
- DAG-grounded trajectory synthesis uses those environments to generate verifier-filtered multi-turn demonstrations for supervised bootstrapping.
- DiscoPO treats the environment itself as the training signal, assigning turn-level credit to actions yielding verifiable analytical evidence.
- The resulting SciDisco-14B model is reported as state-of-the-art on hypothesis-driven scientific data analysis benchmarks; no specific metrics or baselines are given in the abstract.
Sources (1)
Scaling Scientific Discovery Environments for Turn-Level Agentic RL
TL;DR - SciDisco is a framework for training LLM-based scientific discovery agents inside process-verifiable environments, using turn-level RL credit assignment rather than only final-answer rewards. It matters because long-horizon data analysis has lacked environments that can check intermediate analytical progress, not just the final claim.
- SciThèque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress is checkable mid-interaction.
- DAG-grounded trajectory synthesis uses those environments to generate verifier-filtered multi-turn demonstrations for supervised bootstrapping.
- DiscoPO treats the environment itself as the training signal, assigning turn-level credit to actions yielding verifiable analytical evidence.
- The resulting SciDisco-14B model is reported as state-of-the-art on hypothesis-driven scientific data analysis benchmarks; no specific metrics or baselines are given in the abstract.