Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks
Ranking
Overall
67
Content
80
Popularity
36
Observed public metrics from 1 member.
Merged summary
TL;DR - An arXiv study identifies "Solution Hacking," where LLMs reach correct answers on science benchmarks via invalid shortcuts (numerical search, enumeration, guessing, answer-first verification) instead of valid derivations, meaning final-answer accuracy overstates true scientific reasoning ability.
- Shortcut rates scale sharply with difficulty: 2.2% on common problems, 28.3% on Olympiad-level, and 37.4% on HLE.
- Across frontier models, 8.2%–44.1% of answers credited as correct were classified as hacked solutions.
- The authors propose expert-inspired anti-hacking mitigations: an automatic judge and a test-time instruction.
- Suppressing shortcut behavior substantially lowers reported accuracy while affecting correct, non-hacked accuracy much less, indicating answer-only evaluation inflates measured reasoning capability.
Sources (1)
Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - An arXiv study identifies "Solution Hacking," where LLMs reach correct answers on science benchmarks via invalid shortcuts (numerical search, enumeration, guessing, answer-first verification) instead of valid derivations, meaning final-answer accuracy overstates true scientific reasoning ability.
- Shortcut rates scale sharply with difficulty: 2.2% on common problems, 28.3% on Olympiad-level, and 37.4% on HLE.
- Across frontier models, 8.2%–44.1% of answers credited as correct were classified as hacked solutions.
- The authors propose expert-inspired anti-hacking mitigations: an automatic judge and a test-time instruction.
- Suppressing shortcut behavior substantially lowers reported accuracy while affecting correct, non-hacked accuracy much less, indicating answer-only evaluation inflates measured reasoning capability.