Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks
AuthorsXuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao
Resources
Many frontier LLMs may get hard science questions right for the wrong reasons, exposing serious flaws in answer-only reasoning benchmarks.
Key results
Shortcut hacking on easy textbook problems.
Shortcut hacking on olympiad-level competition problems.
Shortcut hacking on Humanity’s Last Exam.
Share of correct answers identified as hacked for GPT-4.1.
Agreement of the deployed peer-panel detector with blinded expert labels.
Reported accuracy under the strongest pre-commit anti-hacking instruction.
What the paper found
Researchers at Alibaba Group and Alibaba DAMO Academy identify “solution hacking,” in which an LLM reaches the correct answer through numerical search, enumeration, guessing, formula recall, or answer-first verification instead of performing the reasoning capability a benchmark targets. Auditing 3528 solutions from frontier models including OpenAI’s GPT-5.2, Anthropic’s Claude Opus 4.7, Google DeepMind’s Gemini-3.1-Pro-Preview, and DeepSeek models, they show hacking rises with difficulty: from 2.2% on easy textbook problems to 28.3% on olympiad-level competition problems and 37.4% on Humanity’s Last Exam, or HLE. Among answers marked correct, hacked solutions account for 8.2%-44.1%, with weaker models showing the greatest inflation. The authors define derivation-adjusted accuracy, which credits only correct, non-hacked solutions, and build a majority-vote expert-anchored judge with 75.4% agreement against blinded PhD annotations; it detects 61.2% of expert-confirmed hacks, making reported rates conservative lower bounds. An anti-hacking instruction that permits abstention reduces hard-core reported accuracy from 41.5% to 33.3% while lowering the hack rate from 22.3% to 6.9% and leaving derivation-adjusted accuracy nearly stable. The central conclusion is that answer-only evaluation can substantially overestimate frontier LLM scientific reasoning, especially when problems have searchable or easily verifiable answer spaces.
Original abstract
Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a failure mode in which an LLM reaches the correct answer through invalid shortcuts, such as numerical search, enumeration, guessing, or answer-first verification, without providing a valid task-targeted derivation. We systematically analyze this phenomenon across difficulty levels, scientific domains, and frontier models. Solution hacking increases sharply with benchmark difficulty, from 2.2\% on common problems to 28.3\% on Olympiad-level problems and 37.4\% on HLE. Moreover, 8.2\%-44.1\% of answers credited as correct across frontier models are identified as hacked solutions. We further develop expert-inspired anti-hacking strategies, including an automatic judge and a test-time instruction. The results show that suppressing shortcut behavior substantially reduces reported accuracy while having a smaller effect on correct and non-hacked accuracy. These findings reveal that answer-only evaluation can overestimate the scientific reasoning capabilities of frontier LLMs.
Read the original paperMore in AI Benchmarks
Browse all 45 papers →Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.
EnigmaForge: The Question Is Hidden in the Story
Daniel Eisner
EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.
Video-Index: A Curated Meta-Benchmark for Video Understanding
Enxin Song, Yinuo Xu, Shusheng Yang, Wenhao Chai, Jiatao Gu
Video-Index stress-tests video benchmarks for shortcuts and provides a curated set of harder, more trustworthy evaluation items.