Training AI Scientists to Replicate Research
TL;DR - Replica is a scalable benchmark and training environment for AI agents that reproduce published research. Post-trained on it, the 27B-parameter Faraday agent outperforms Claude Opus 4.8 and GPT-5.5 on held-out replication tasks.
- Replica frames paper replication as hypothesis-driven, long-horizon scientific work.
- An automatically generated rubric-based judge provides low-noise rewards aligned with human assessments.
- Faraday delegates coding work to coding agents used as tools.
- Rollout analysis suggests Faraday follows a more scientifically principled replication process.