Superficial Beliefs in LLM Decision-Making
AuthorsGabriel Freedman, Francesca Toni
Resources
This paper shows that LLMs can make structured choices for hidden reasons that their own explanations only partly reveal.
Key results
source problems used for fitting per theme
held-out source problems used for evaluation per theme
behavioural model vs observed choices across all main themes
direct report vs behavioural-model revealed driver across all main themes
score-based judge vs behavioural-model revealed driver across all main themes
What the paper found
Superficial Beliefs in LLM Decision-Making, from Imperial College London, asks whether models such as OpenAI’s GPT-5-mini and GPT-5-nano, Qwen3-14B, and Ministral-3-14B merely rehearse explanations or instead exhibit a recoverable local decision structure. The authors build a synthetic binary-choice benchmark with four graded attributes across drugs, policy, and software themes, plus two control themes and a structurally different six-attribute hospital cyber-response setting. They fit a binomial logistic behavioural model to 400 training families and evaluate on 100 held-out families, comparing its recovered driver with two explicit elicitation methods: direct response and a score-based judge. Across the main benchmark, the behavioural model matches held-out choices 80.4% of the time, while the explicit direct report names the recovered driver only 61.0% of the time and the score-based judge 61.3% of the time. The same pattern survives perturbation and intervention: score-based judge outputs are more reproducible than direct responses, yet they are not more faithful to the recovered decision driver, and in the control themes irrelevant attributes are selected only 0.2%–1.6% of the time. A targeted occlusion study on GPT-5-mini non-thinking drugs shows that higher-ranked attributes induce larger choice and rationale flips, supporting the claim that these models behave as if guided by probabilistic, attribute-local priorities rather than fully articulated beliefs.
Original abstract
We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which models choose between profiles defined by graded attributes, we compare the attribute a model says mattered most with the attribute that best explains its choice under a behavioural model fit to prior decisions. The behavioural model predicts held-out choices well, showing that model behaviour is systematically related to the visible attributes rather than being random. However, direct self-reports and a separate score-based judge recover the behaviourally inferred driver only partially. The resulting picture is neither one of arbitrary behaviour nor one of fully articulated belief - outputs are structured enough to support prediction, but explicit reasons track the recovered driver only imperfectly. This qualitative pattern persists across prompt-order and sampling perturbations, alternative behavioural models, targeted occlusion analyses, and structurally varied decision settings. We interpret this as evidence for ``superficial belief'' in LLM decision-making: models behave as if guided by probabilistic local priorities over attributes, while having only limited verbal access to the attributes that drive their decisions.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.