NTH

Why Large Language Models Fail at Tabular Prediction

AuthorsMarta Garnelo, Wojciech M. Czarnecki

August 6, 2026 3 min read
Watch on YouTube
The one-line take

LLMs can reason impressively over many modalities, but this study finds that their tabular prediction ability collapses as the number of input dimensions grows.

Key results

31
Dimensionality benchmark sweep

Number of benchmark datasets used in the random-projection dimensionality experiment

9
Methods compared

Claude was the only method whose accuracy declined as dimensionality increased

91.6%
2D Gaussian-process agreement

Grid-prediction agreement between Claude and the best standardized Matérn Gaussian process

64.8%
Best high-dimensional agreement

Maximum prediction agreement between Claude and any of 252 configured classical models

0.64
Noise-model improvement

Largest agreement gain, in percentage points, from tuned dimension-dependent label noise

14
Unfaithful explanations

Of 20 two-dimensional tasks where the model’s explanation failed to match the data

What the paper found

Marta Garnelo and Wojciech M. Czarnecki investigate why generic language models struggle with tabular classification when prompted in pure inference mode: one generation over complete training and test tables, without tools, retrieval, agents, or fine-tuning. Using Anthropic’s Claude Opus 4.6, with an additional two-dimensional probe on Qwen3-235B-A22B, they test five explanations: class overlap, flattened CSV structure, numeric tokenisation, the number of test rows per query, and feature dimensionality. Controlled interventions reject the first four: Claude can copy a target column even at 60 features, and rounding numbers or splitting test batches does not rescue accuracy. The decisive result comes from random projections of 31 benchmark datasets: Claude is the only one of 9 methods whose accuracy declines as dimensionality increases, while classical baselines remain flat or improve; by roughly 16 dimensions, performance is often at or below majority-class guessing. Behaviourally, the model resembles a local distance-based learner in two dimensions: a standardized Matérn Gaussian process agrees with 91.6% of its grid predictions, and 1-nearest-neighbour reaches 91.0%. In higher dimensions, none of 252 configured classical models reproduces its predictions beyond 64.8% agreement, and dimension-dependent noise improves agreement by at most 0.64 percentage points. A memorisation probe also exposes contamination in familiar datasets, while explanations fail to match the data in 14 of 20 tasks. The authors conclude that prompt formatting is not the core problem; generic LLM in-context learning dissolves as tables become wide, supporting purpose-built tabular foundation models such as TabPFN while leaving the internal mechanism unresolved.

Original abstract

Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. This gap is the founding premise of the fast-growing field of tabular foundation models, but the question of why generic LLMs fail has remained open. We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning - and systematically evaluate five hypotheses for the failure: (a) an inability to handle noisy or non-linearly-separable data; (b) the linearised CSV format obscuring column structure; (c) the tokenisation of numeric values; (d) the number of test points classified per query; and (e) the dimensionality of the input. Controlled experiments falsify (a)-(d). Dimensionality, in contrast, is decisive: sweeping random linear projections of thirty-one benchmark datasets, the LLM is the only method among nine whose accuracy decreases as dimensionality grows, while every classical baseline stays flat or improves. A behavioural comparison against 252 configured classical models finds that in two dimensions the LLM predicts like a local, distance-based method (up to 91.6% grid agreement), but in higher dimensions no classical model - even when augmented with tuned, dimension-dependent noise - reproduces its predictions. We do not claim to have identified the internal mechanism; our results show, more modestly, that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis