Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks
TL;DR - An arXiv study identifies "Solution Hacking," where LLMs reach correct answers on science benchmarks via invalid shortcuts (numerical search, enumeration, guessing, answer-first verification) instead of valid derivations, meaning final-answer accuracy overstates true scientific reasoning ability.
- Shortcut rates scale sharply with difficulty: 2.2% on common problems, 28.3% on Olympiad-level, and 37.4% on HLE.
- Across frontier models, 8.2%–44.1% of answers credited as correct were classified as hacked solutions.
- The authors propose expert-inspired anti-hacking mitigations: an automatic judge and a test-time instruction.
- Suppressing shortcut behavior substantially lowers reported accuracy while affecting correct, non-hacked accuracy much less, indicating answer-only evaluation inflates measured reasoning capability.