NTH

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

AuthorsPatrik Wolf, Thomas Kleine Buening, Andreas Krause, Celestine Mendler-Dünner

July 17, 2026 3 min read
Watch on YouTube
The one-line take

The study shows that language models often know how subgroups differ but fail to combine that knowledge into accurate population-level answers.

Key results

3.8M
ACS reference population

Individuals represented in the 2024 American Community Survey evaluation.

0.33
GPT-5.4 ACS split consistency

Fraction of split checks satisfied on the ACS income task.

1.00
GPT-5.4 ACS order consistency

Fraction of order checks satisfied on the ACS income task.

0.70
Sonnet 4.6 WVS split consistency

Average split consistency across five World Values Survey questions and two countries.

0.90
Sonnet 4.6 WVS order consistency

Average order consistency across five World Values Survey questions and two countries.

83.7%
Qwen3.6 Plus micro-to-macro win rate

Share of GlobalOpinionQA question-country cases improved over direct prompting.

What the paper found

Researchers from the Max Planck Institute for Intelligent Systems, ETH Zürich, ELLIS, and the Tübingen AI Center test whether language models behave like coherent conditional probability estimators. Their method uses binary conditioning trees: models estimate distributions for increasingly specific subpopulations, then researchers recombine those estimates with the law of total probability and compare them with direct population-level answers. On the 2024 American Community Survey, covering 3.8M individuals, this reveals the “macro fallacy”: subgroup estimates aggregated upward are often more accurate than direct aggregate prompts. For example, GPT-5.4’s ACS income alignment error falls from 0.14 at the root to 0.03 after level-2 aggregation. The authors also introduce reference-free split consistency and order consistency tests, evaluated at tolerance 0.02. GPT-5.4 achieves only 0.33 split consistency but 1.00 order consistency on the ACS income task, while Anthropic’s Sonnet 4.6 reaches 0.70 and 0.90 average split and order consistency on five World Values Survey questions across Canada and Indonesia. The failures generalize to opinion modeling and synthetic forecasting, and do not systematically disappear in stronger models, including GPT, Grok, DeepSeek, and Qwen systems. A lightweight “micro-to-macro” prompt, which asks models to reason through subpopulations before answering, partially recovers the benefit; Qwen3.6 Plus improves over direct prompting on 83.7% of GlobalOpinionQA cases. The paper argues that statistical self-consistency is a distinct, unsaturated evaluation axis complementary to accuracy and human-distribution alignment.

Original abstract

In-context learning is commonly interpreted as a form of conditional inference, in which the prompt specifies a context and the model's output is treated as an estimate of the corresponding conditional distribution. If this interpretation holds, then LLM estimates should satisfy basic probabilistic identities. In particular, the law of total probability asserts that prior-weighted conditional distributions aggregate into population-level marginals over any valid partition of the population. In this work, we investigate to what extent LLM estimates adhere to this self-consistency principle. We use binary trees as an evaluation scaffold to recursively partition a population into increasingly fine-grained subpopulations. We then prompt LLMs with verbalized subpopulation descriptions in context, aggregate the resulting estimates back into population-level estimates, and compare them across partitions of varying granularity. Applying this protocol across problem domains and state-of-the-art frontier models, we show widespread violations of basic consistency properties. An in-depth study of persona prompting reveals a pattern we call the macro fallacy: estimates reconstructed from more fine-grained subpopulation responses are often better aligned with human reference data than direct population-level estimates. This effect persists across variations in tree structure and estimation task, and can be partially recovered through implicit prompting. Together, these findings suggest that models possess relevant subpopulation knowledge but do not reliably propagate it into aggregate estimates. This gap establishes statistical self-consistency as an unsaturated, reference-free criterion for evaluating LLMs.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis