Why Large Language Models Fail at Tabular Prediction
AuthorsMarta Garnelo, Wojciech M. Czarnecki
Resources
LLMs can reason impressively over many modalities, but this study finds that their tabular prediction ability collapses as the number of input dimensions grows.
Key results
Number of benchmark datasets used in the random-projection dimensionality experiment
Claude was the only method whose accuracy declined as dimensionality increased
Grid-prediction agreement between Claude and the best standardized Matérn Gaussian process
Maximum prediction agreement between Claude and any of 252 configured classical models
Largest agreement gain, in percentage points, from tuned dimension-dependent label noise
Of 20 two-dimensional tasks where the model’s explanation failed to match the data
What the paper found
Marta Garnelo and Wojciech M. Czarnecki investigate why generic language models struggle with tabular classification when prompted in pure inference mode: one generation over complete training and test tables, without tools, retrieval, agents, or fine-tuning. Using Anthropic’s Claude Opus 4.6, with an additional two-dimensional probe on Qwen3-235B-A22B, they test five explanations: class overlap, flattened CSV structure, numeric tokenisation, the number of test rows per query, and feature dimensionality. Controlled interventions reject the first four: Claude can copy a target column even at 60 features, and rounding numbers or splitting test batches does not rescue accuracy. The decisive result comes from random projections of 31 benchmark datasets: Claude is the only one of 9 methods whose accuracy declines as dimensionality increases, while classical baselines remain flat or improve; by roughly 16 dimensions, performance is often at or below majority-class guessing. Behaviourally, the model resembles a local distance-based learner in two dimensions: a standardized Matérn Gaussian process agrees with 91.6% of its grid predictions, and 1-nearest-neighbour reaches 91.0%. In higher dimensions, none of 252 configured classical models reproduces its predictions beyond 64.8% agreement, and dimension-dependent noise improves agreement by at most 0.64 percentage points. A memorisation probe also exposes contamination in familiar datasets, while explanations fail to match the data in 14 of 20 tasks. The authors conclude that prompt formatting is not the core problem; generic LLM in-context learning dissolves as tables become wide, supporting purpose-built tabular foundation models such as TabPFN while leaving the internal mechanism unresolved.
Original abstract
Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. This gap is the founding premise of the fast-growing field of tabular foundation models, but the question of why generic LLMs fail has remained open. We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning - and systematically evaluate five hypotheses for the failure: (a) an inability to handle noisy or non-linearly-separable data; (b) the linearised CSV format obscuring column structure; (c) the tokenisation of numeric values; (d) the number of test points classified per query; and (e) the dimensionality of the input. Controlled experiments falsify (a)-(d). Dimensionality, in contrast, is decisive: sweeping random linear projections of thirty-one benchmark datasets, the LLM is the only method among nine whose accuracy decreases as dimensionality grows, while every classical baseline stays flat or improves. A behavioural comparison against 252 configured classical models finds that in two dimensions the LLM predicts like a local, distance-based method (up to 91.6% grid agreement), but in higher dimensions no classical model - even when augmented with tuned, dimension-dependent noise - reproduces its predictions. We do not claim to have identified the internal mechanism; our results show, more modestly, that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.