NTH

A Mechanistic Study of Tabular Foundation Models

AuthorsMarin Biloš, James T. Wilson, Anderson Schneider, Yuriy Nevmyvaka

May 28, 2026 2 min read
Watch on YouTube
The one-line take

This paper peels back the internals of tabular foundation models to show how different architectures make predictions, why they become permutation-invariant, and how they fail under targeted attacks.

Key results

49 classification and 10 regression datasets
Evaluation suite

The paper audits TabPFNv2, TabICLv2, and Mitra across the full benchmark suite used for the mechanistic study.

mean maximum attention weight 0.925
TabPFNv2 row-attention sharpness

At L9, TabPFNv2 attention becomes highly concentrated on a small set of context rows.

Pearson r=0.89
TabPFNv2 vote fidelity

An attention-weighted label vote at L9 closely tracks TabPFNv2 native predicted probabilities.

0.874 to 0.488
TabPFNv2 uniform-attention ablation

Replacing the L9 attention pattern with uniform attention collapses accuracy to the majority-class baseline.

0.854 versus 0.864 native accuracy
TabICLv2 prototype readout

A nearest-class prototype decoder on TabICLv2’s final representation nearly matches the model’s end-to-end classification accuracy.

What the paper found

This paper reverse-engineers three tabular foundation models, TabPFNv2, TabICLv2, and Mitra, on 49 classification and 10 regression datasets and shows that similar benchmark accuracy hides fundamentally different inference algorithms. TabPFNv2 and Mitra realize a late similarity-based readout: at layer 9, TabPFNv2’s query-row attention becomes sharply peaked, with mean maximum attention weight 0.925, and an attention-weighted label vote matches native probabilities with Pearson r=0.89; forcing uniform attention drops accuracy from 0.874 to 0.488. Mitra follows the same vote family with r=0.89 but is not identical, because its shared per-feature embedding and label slot make column permutation exact by construction. TabICLv2 is mechanistically different: most class structure is already readable at the column-embedding output, and a parameter-free class-prototype decoder on the final representation reaches 0.854 versus 0.864 native accuracy, while kNN on the final representation reaches 0.856. Cross-backbone readout transplantation fails catastrophically, losing 33 to 40 percentage points, showing the readout is jointly tuned to each backbone. The authors also identify cheap symmetry fixes: zeroing TabPFNv2’s positional matrix W or removing RoPE from TabICLv2 makes column invariance exact with no accuracy loss, and a One-vs-All wrapper yields exact class-order invariance. Finally, mechanism-grounded attacks confirm the inferred circuits: hub poisoning cuts accuracy by 3.3 to 3.9 points on TabPFNv2/TabICLv2, rank warp costs 8.0 to 10.1 points on both but barely affects Mitra, and SVD burial causes up to 8.3-point drops, directly linking model failures to their learned similarity or prototype geometry.

Original abstract

Tabular foundation models with different architectures converge in accuracy across a range of classification and regression tasks. This raises questions a leaderboard cannot answer: (i) whether the models execute the same in-context algorithm, (ii) where row, column, and class-permutation invariances originate, and (iii) how robust they are under perturbations engineered against the inferred mechanism. We characterize all three. The model families realize qualitatively distinct similarity-based readouts: from an attention-weighted vote over context labels to a class-conditional mean readout, each confirmed by causal intervention. We find that the representation collapse highlighted in prior work is not a practical concern for them. Each model's permutation invariances trace to specific positional parameters whose removal preserves accuracy and makes approximate invariance exact. Perturbations engineered against each readout reproduce predicted failure modes; hub and rank attacks isolate them from refit baselines. Together these results give a mechanistic account of contemporary tabular foundation models and identify which inductive biases govern both their accuracy and characteristic failures.

Read the original paper

More in Foundation Models

Browse all 47 papers →
01Foundation Model

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.

Read analysis
02Foundation Model

TabFM: A Zero-Shot Foundation Model for Tabular Data

Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das

TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.

Read analysis
03Foundation Model

When Do Biological Reasoning Models Use Their Biological Inputs?

Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik

The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.

Read analysis