What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
AuthorsRaphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho
Resources
The study shows that an LLM's hidden representations can reveal its confidence and evidence use more reliably than its spoken chain of thought.
Key results
Expected calibration error for the layer-21 covariance probe on OpenForesight.
Expected calibration error for EF-8B’s stated probability on the same evaluation.
High-impact evidence-ablation cases where the forecast changed but the chain-of-thought did not.
Spearman correlation between probe shifts and forecast changes across perturbation cases.
Rate at which the probe correctly predicted the direction of behavioral change.
Upper end of generated-token savings from pre-reasoning question routing without measurable accuracy loss.
What the paper found
This paper from Goodfire and Eternis investigates whether an LLM’s internal activations reveal more about confidence and reasoning than its verbal explanations. Using Eternis-Forecaster 8B, or EF-8B, on OpenForesight, the authors train lightweight mean, attention, and covariance-pooling probes over frozen intermediate representations; a layer-21 covariance probe reduces expected calibration error to 0.044, compared with 0.093 for the model’s verbalized probability. The approach also transfers without updating model weights to Zhipu AI’s GLM-4.7-Flash and GLM-4.5-Air, and supports out-of-distribution correctness monitoring for Qwen3-8B on AIME and AMC mathematics. Faithfulness tests remove real news evidence or inject fabricated articles generated with Claude Sonnet. In evidence ablation, 23% of high-impact perturbations changed the forecast while leaving the chain-of-thought unchanged, and written reasoning correlated with behavioral change at only Spearman rho 0.215. A layer-20 attention probe tracked those shifts at rho 0.565 and predicted their direction correctly in 83.6% of cases, including many silent changes. Finally, forced answering with the larger EF-32B shows that confidence and much of the committed answer are fixed before chain-of-thought begins. Routing questions by the entropy of this pre-reasoning answer distribution saves 30% to 47% of generated tokens with no measurable accuracy loss. The central conclusion is that internal representations provide a practical channel for calibration, auditing, and inference-time triage beyond what the model can faithfully say.
Original abstract
Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not faithfully reflect the evidence behind a forecast. We ask whether internal representations offer a more direct window into both. Working with Eternis-Forecaster 8B on OpenForesight, we train representation-pooling probes on intermediate activations and find they achieve substantially better calibration; a result that also holds for GLM-4.7-Flash and GLM-4.5-Air. We then assess CoT faithfulness through evidence ablation and diversionary injection: removing an influential source in the prompt often changes the model's forecast while leaving the reasoning trace untouched. The same probes function as lie detectors: their activations track behavioral shifts far better than the reasoning trace does, and they also predict the direction of change in 84% of cases, including when the CoT conceals the perturbation's influence. Finally, forced answering reveals that forecasts are largely fixed before reasoning begins: a single pre-reasoning pass recovers the committed answer and confidence, and routing questions by the spread of this pre-set answer distribution saves 30-47% of generated tokens, with no loss of accuracy. Together, these results establish probing internal representations as a practical tool for calibrating, auditing, and triaging language model forecasters and reasoning models more broadly.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.