NTH

Superficial Beliefs in LLM Decision-Making

AuthorsGabriel Freedman, Francesca Toni

June 24, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that LLMs can make structured choices for hidden reasons that their own explanations only partly reveal.

Key results

400
training families

source problems used for fitting per theme

100
test families

held-out source problems used for evaluation per theme

80.4%
held-out choice agreement

behavioural model vs observed choices across all main themes

61.0%
direct attribute agreement

direct report vs behavioural-model revealed driver across all main themes

61.3%
score-based attribute agreement

score-based judge vs behavioural-model revealed driver across all main themes

What the paper found

Superficial Beliefs in LLM Decision-Making, from Imperial College London, asks whether models such as OpenAI’s GPT-5-mini and GPT-5-nano, Qwen3-14B, and Ministral-3-14B merely rehearse explanations or instead exhibit a recoverable local decision structure. The authors build a synthetic binary-choice benchmark with four graded attributes across drugs, policy, and software themes, plus two control themes and a structurally different six-attribute hospital cyber-response setting. They fit a binomial logistic behavioural model to 400 training families and evaluate on 100 held-out families, comparing its recovered driver with two explicit elicitation methods: direct response and a score-based judge. Across the main benchmark, the behavioural model matches held-out choices 80.4% of the time, while the explicit direct report names the recovered driver only 61.0% of the time and the score-based judge 61.3% of the time. The same pattern survives perturbation and intervention: score-based judge outputs are more reproducible than direct responses, yet they are not more faithful to the recovered decision driver, and in the control themes irrelevant attributes are selected only 0.2%–1.6% of the time. A targeted occlusion study on GPT-5-mini non-thinking drugs shows that higher-ranked attributes induce larger choice and rationale flips, supporting the claim that these models behave as if guided by probabilistic, attribute-local priorities rather than fully articulated beliefs.

Original abstract

We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which models choose between profiles defined by graded attributes, we compare the attribute a model says mattered most with the attribute that best explains its choice under a behavioural model fit to prior decisions. The behavioural model predicts held-out choices well, showing that model behaviour is systematically related to the visible attributes rather than being random. However, direct self-reports and a separate score-based judge recover the behaviourally inferred driver only partially. The resulting picture is neither one of arbitrary behaviour nor one of fully articulated belief - outputs are structured enough to support prediction, but explicit reasons track the recovered driver only imperfectly. This qualitative pattern persists across prompt-order and sampling perturbations, alternative behavioural models, targeted occlusion analyses, and structurally varied decision settings. We interpret this as evidence for ``superficial belief'' in LLM decision-making: models behave as if guided by probabilistic local priorities over attributes, while having only limited verbal access to the attributes that drive their decisions.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis