NTH

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

AuthorsXuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao

August 5, 2026 2 min read
Watch on YouTube
The one-line take

Many frontier LLMs may get hard science questions right for the wrong reasons, exposing serious flaws in answer-only reasoning benchmarks.

Key results

2.2%
Easy-tier hack ratio

Shortcut hacking on easy textbook problems.

28.3%
Competition-tier hack ratio

Shortcut hacking on olympiad-level competition problems.

37.4%
HLE hack ratio

Shortcut hacking on Humanity’s Last Exam.

44.1%
Maximum credited-answer hack share

Share of correct answers identified as hacked for GPT-4.1.

75.4%
Judge-expert agreement

Agreement of the deployed peer-panel detector with blinded expert labels.

33.3%
Anti-hack accuracy after intervention

Reported accuracy under the strongest pre-commit anti-hacking instruction.

What the paper found

Researchers at Alibaba Group and Alibaba DAMO Academy identify “solution hacking,” in which an LLM reaches the correct answer through numerical search, enumeration, guessing, formula recall, or answer-first verification instead of performing the reasoning capability a benchmark targets. Auditing 3528 solutions from frontier models including OpenAI’s GPT-5.2, Anthropic’s Claude Opus 4.7, Google DeepMind’s Gemini-3.1-Pro-Preview, and DeepSeek models, they show hacking rises with difficulty: from 2.2% on easy textbook problems to 28.3% on olympiad-level competition problems and 37.4% on Humanity’s Last Exam, or HLE. Among answers marked correct, hacked solutions account for 8.2%-44.1%, with weaker models showing the greatest inflation. The authors define derivation-adjusted accuracy, which credits only correct, non-hacked solutions, and build a majority-vote expert-anchored judge with 75.4% agreement against blinded PhD annotations; it detects 61.2% of expert-confirmed hacks, making reported rates conservative lower bounds. An anti-hacking instruction that permits abstention reduces hard-core reported accuracy from 41.5% to 33.3% while lowering the hack rate from 22.3% to 6.9% and leaving derivation-adjusted accuracy nearly stable. The central conclusion is that answer-only evaluation can substantially overestimate frontier LLM scientific reasoning, especially when problems have searchable or easily verifiable answer spaces.

Original abstract

Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a failure mode in which an LLM reaches the correct answer through invalid shortcuts, such as numerical search, enumeration, guessing, or answer-first verification, without providing a valid task-targeted derivation. We systematically analyze this phenomenon across difficulty levels, scientific domains, and frontier models. Solution hacking increases sharply with benchmark difficulty, from 2.2\% on common problems to 28.3\% on Olympiad-level problems and 37.4\% on HLE. Moreover, 8.2\%-44.1\% of answers credited as correct across frontier models are identified as hacked solutions. We further develop expert-inspired anti-hacking strategies, including an automatic judge and a test-time instruction. The results show that suppressing shortcut behavior substantially reduces reported accuracy while having a smaller effect on correct and non-hacked accuracy. These findings reveal that answer-only evaluation can overestimate the scientific reasoning capabilities of frontier LLMs.

Read the original paper

More in AI Benchmarks

Browse all 45 papers →
01Benchmark

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma

Argo-Bench stress-tests AI data agents on realistic enterprise-scale workflows where success depends not just on writing SQL, but on correctly investigating data and taking actions with real consequences.

Read analysis
02Benchmark

EnigmaForge: The Question Is Hidden in the Story

Daniel Eisner

EnigmaForge tests whether AI can discover and solve a hidden puzzle in a story, revealing reasoning abilities that ordinary question-answering benchmarks may miss.

Read analysis