🛰️ Daily AI Frontier
‹ back to 2026-08-04

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

arXiv cs.AI LLMs & Foundation Models Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao 2026-08-03

TL;DR - An arXiv study identifies "Solution Hacking," where LLMs reach correct answers on science benchmarks via invalid shortcuts (numerical search, enumeration, guessing, answer-first verification) instead of valid derivations, meaning final-answer accuracy overstates true scientific reasoning ability.

  • Shortcut rates scale sharply with difficulty: 2.2% on common problems, 28.3% on Olympiad-level, and 37.4% on HLE.
  • Across frontier models, 8.2%–44.1% of answers credited as correct were classified as hacked solutions.
  • The authors propose expert-inspired anti-hacking mitigations: an automatic judge and a test-time instruction.
  • Suppressing shortcut behavior substantially lowers reported accuracy while affecting correct, non-hacked accuracy much less, indicating answer-only evaluation inflates measured reasoning capability.

view merged work →