Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
AuthorsZheng Wu, Chenhao Xue, Shijie Zheng, Yijie Lu, Cheng Yang, Zhuosheng Zhang
This paper shows that LLMs often know the commonsense answer but get distracted by overly salient details, and that better prompting can help them recover it.
Key results
Physically impossible commonsense tasks across four trap dimensions.
State-of-the-art models from Anthropic, OpenAI, Google, DeepSeek, ByteDance, and other labs.
Claude-Opus-4.7’s overall TAR on SaliTrap.
A context-free knowledge probe recovered over 90% of sycophantic-compliance failures.
Increase from 25.9% control TAR to 57.4% with physics-aware priming.
What the paper found
The paper identifies salience bias, a failure mode in which large language models over-prioritize explicit numerical or procedural details and suppress implicit commonsense prerequisites. In the motivating example, Gemini and DeepSeek focus on a car wash being 50 meters away and recommend walking, ignoring that the car itself must be driven there. The authors introduce SaliTrap, a benchmark of 1,145 physically impossible, computation-laden tasks spanning missing prerequisites, environmental mismatch, temporal or physiological violations, and rule mismatch. Across 12 models—including Anthropic’s Claude Opus, OpenAI’s GPT-5, Google’s Gemini, DeepSeek, GLM, Kimi, and ByteDance’s Doubao—every system showed vulnerability: even Claude-Opus-4.7 achieved only 54.8% trap-avoidance accuracy, while 8 of 12 models fell below 30%. The benchmark separates failure to detect a trap from detecting it but complying anyway; GLM-5.1 and Kimi-K2 still complied in 86.2% and 81.8% of trap-aware cases. Increasing the number of injected numerical distractors consistently reduced avoidance and increased chain-of-thought hijacking. However, a context-free probe exposing only the physical contradiction recovered over 90% of sycophantic-compliance failures, indicating knowledge suppression rather than missing commonsense. Lightweight inference-time prompting was effective without retraining: physics-aware priming raised GLM-5.1’s trap-avoidance rate by 31.4%, while reducing hard failures from 27.1% to 0.4%. The authors conclude that commonsense reasoning is often an elicitation problem: salient task framing crowds out knowledge the models already possess.
Original abstract
As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bias: models become easily hijacked by useless explicit distractors (e.g., numerical values), leading them to ignore the implicit physical or commonsense prerequisites of a task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, a high-quality dataset across four trap dimensions. Evaluating 12 state-of-the-art LLMs, we find that all mainstream models suffer significantly from salience bias, with severity scaling with distractor density and detecting the trap often decoupled from actually avoiding it. Crucially, by re-eliciting the same models with the task framing stripped away, we show that this is overwhelmingly a failure of \textbf{knowledge suppression rather than knowledge absence}: a context-free knowledge probe alone recovers over 90\% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors that lure the model into over-compliant, unnecessary computation. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and we release SaliTrap as a testbed for this blind spot. The codes are available at https://github.com/Wuzheng02/SaliTrap.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.