NTH

Large Language Models Are Overconfident in Their Own Responses

AuthorsMario Sanz-Guerrero, Manuel Mager, Katharina von der Wense

June 16, 2026 2 min read
Watch on YouTube
The one-line take

This paper shows that LLMs are strangely more confident in answers they think are their own, and that a simple prompt trick can make them better calibrated without retraining.

Key results

61
open-weight LLMs evaluated

Models spanning Llama 3.1, Qwen3, and Gemma 3

3.7%
accuracy gain from instruction tuning

Average accuracy improvement over base models on MMLU

13.1%
ECE increase from instruction tuning

Average calibration degradation over base models on MMLU

6.5%
Brier increase from instruction tuning

Average probabilistic error increase over base models on MMLU

26.8%
assistant-vs-user confidence gap

Largest average raw-confidence increase when the answer is framed as assistant output

What the paper found

In “Large Language Models Are Overconfident in Their Own Responses,” researchers from Johannes Gutenberg University Mainz and the University of Colorado Boulder show that calibration failures in instruction-tuned LLMs are driven less by the chat template alone than by an ownership bias: the same answer is assigned substantially higher confidence when it is framed as the model’s own output than when it is framed as user input. Across 61 open-weight LLMs spanning Llama 3.1, Qwen3, and Gemma 3, evaluated on MMLU with P(True), verbalized percentage, and verbalized linguistic confidence, instruction tuning increased accuracy by 3.7% but worsened ECE by 13.1% and Brier score by 6.5%, and adding the chat template pushed the total ECE increase to 15.8%. When the answer was presented as assistant output, models were up to 26.8% more confident than when the identical answer came from the user, with the largest calibration gaps appearing under verbalized linguistic confidence. This pattern generalized beyond MMLU to GSM8K, TruthfulQA, open-ended MMLU, and even GPT-5.2, where assistant-framed answers remained more miscalibrated across all three elicitation methods. The paper’s practical contribution is a zero-shot inference-time fix: during confidence elicitation, reframe the model’s own answer as user input, which consistently lowers overconfidence and can recover calibration close to, or sometimes better than, base models without retraining.

Original abstract

Prior work has shown that instruction-tuned large language models (LLMs) are less well calibrated than their base pre-trained counterparts. However, little is known about the frequently used chat template's effect on the calibration of conversational LLMs. In this work, we investigate the mechanisms driving this miscalibration by decoupling the effects of the post-training algorithm and the chat format. We find that, while instruction tuning fundamentally harms calibration, the chat template aggravates the issue through an "ownership bias" -- models are significantly more confident in their own answers than in identical answers provided by a user. Extensive experiments across six recent open-weight LLMs, three benchmarks, and three confidence elicitation methods show that models assign up to 26% higher confidence to their own responses. Leveraging this insight, we propose a simple inference-time strategy: framing the model's answer as user input during confidence elicitation. This approach significantly reduces overconfidence and improves calibration by up to 26% without the need for retraining, narrowing the gap between base and instruction-tuned models.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis