Large Language Models Are Overconfident in Their Own Responses
AuthorsMario Sanz-Guerrero, Manuel Mager, Katharina von der Wense
Resources
This paper shows that LLMs are strangely more confident in answers they think are their own, and that a simple prompt trick can make them better calibrated without retraining.
Key results
Models spanning Llama 3.1, Qwen3, and Gemma 3
Average accuracy improvement over base models on MMLU
Average calibration degradation over base models on MMLU
Average probabilistic error increase over base models on MMLU
Largest average raw-confidence increase when the answer is framed as assistant output
What the paper found
In “Large Language Models Are Overconfident in Their Own Responses,” researchers from Johannes Gutenberg University Mainz and the University of Colorado Boulder show that calibration failures in instruction-tuned LLMs are driven less by the chat template alone than by an ownership bias: the same answer is assigned substantially higher confidence when it is framed as the model’s own output than when it is framed as user input. Across 61 open-weight LLMs spanning Llama 3.1, Qwen3, and Gemma 3, evaluated on MMLU with P(True), verbalized percentage, and verbalized linguistic confidence, instruction tuning increased accuracy by 3.7% but worsened ECE by 13.1% and Brier score by 6.5%, and adding the chat template pushed the total ECE increase to 15.8%. When the answer was presented as assistant output, models were up to 26.8% more confident than when the identical answer came from the user, with the largest calibration gaps appearing under verbalized linguistic confidence. This pattern generalized beyond MMLU to GSM8K, TruthfulQA, open-ended MMLU, and even GPT-5.2, where assistant-framed answers remained more miscalibrated across all three elicitation methods. The paper’s practical contribution is a zero-shot inference-time fix: during confidence elicitation, reframe the model’s own answer as user input, which consistently lowers overconfidence and can recover calibration close to, or sometimes better than, base models without retraining.
Original abstract
Prior work has shown that instruction-tuned large language models (LLMs) are less well calibrated than their base pre-trained counterparts. However, little is known about the frequently used chat template's effect on the calibration of conversational LLMs. In this work, we investigate the mechanisms driving this miscalibration by decoupling the effects of the post-training algorithm and the chat format. We find that, while instruction tuning fundamentally harms calibration, the chat template aggravates the issue through an "ownership bias" -- models are significantly more confident in their own answers than in identical answers provided by a user. Extensive experiments across six recent open-weight LLMs, three benchmarks, and three confidence elicitation methods show that models assign up to 26% higher confidence to their own responses. Leveraging this insight, we propose a simple inference-time strategy: framing the model's answer as user input during confidence elicitation. This approach significantly reduces overconfidence and improves calibration by up to 26% without the need for retraining, narrowing the gap between base and instruction-tuned models.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.