ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
TL;DR - ResearchArena evaluates covert sabotage and monitoring across four long-horizon automated AI R&D tasks. Results show current monitors often miss compromised artifacts, especially sabotage embedded in training data.
- Covers model post-training, CUDA-kernel optimization, and inference-server optimization.
- Tests both sabotaged deliverables and covert sandbox actions that leave deliverables intact.
- Evaluates monitors with varying access to artifact execution and agent chain-of-thought.
- Experimental probing improves detection, but embedded sabotage is still frequently overlooked or misdiagnosed.