Xiaomi-TabLDM: A Tabular Foundation Model Technical Report
AuthorsXiaomi-TabLDM Team, :, Penghui Wang, Wei Liu, Hong Wang, Chengyue Huang, Yuxi Sun, Zirui Wang, Hongming Huang, Quan Wang, Chunxiao Liu, Erli Meng, Bin Wang
Resources
Xiaomi-TabLDM is a synthetic-data-trained tabular foundation model that aims to deliver strong, efficient classification and regression without task-specific fine-tuning.
Key results
Second-place Elo score on the 13-dataset TabArena regression suite.
Best average rank across 33 OpenML-CTR23 regression datasets.
Less training time on TabArena regression than TabFM.
Less prediction time on TabArena regression than TabFM.
Total parameters in the released Xiaomi-TabLDM regressor.
Share of 2,240 controlled synthetic datasets where Xiaomi-TabLDM achieved the best R2 score.
What the paper found
Xiaomi-TabLDM is a tabular foundation model for classification and regression that predicts unseen datasets through in-context learning, without task-specific fine-tuning. Unlike models such as TabPFN-3 and TabFM, it is pretrained entirely on synthetic tasks generated from structural causal models, using a three-stage curriculum that reaches contexts of up to 60,000 samples. Its architecture combines dual-stream feature grouping, Set Transformer column encoders, lightweight Attention Residual connections, query-aware scalable softmax, and sparse Mixture-of-Experts layers, while test-time scaling combines shuffled feature subsets and preprocessing views through nonnegative least squares. Across TALENT, TabArena, BCCO, and OpenML-CTR23, its strongest results are in regression: it achieves an Elo score of 1900 on the 13-dataset TabArena regression suite, ranking second behind TabFM, and an average rank of 3.03 on OpenML-CTR23, ranking first across 33 regression datasets. The TabArena regression result comes with 82% less training time and 68% less prediction time than TabFM, while outperforming comparable models including TabPFN-3. The released regressor contains 71.08M total parameters, with 62.68M active in a forward pass because of sparse routing. In a controlled robustness study spanning 2,240 synthetic datasets and fourteen noise conditions, Xiaomi-TabLDM achieved the best R2 score on 70.3% of datasets, indicating that its advantage persists under diverse exogenous noise.
Original abstract
We introduce Xiaomi-TabLDM, a tabular large data foundation model for classification and regression via in-context learning, which delivers superior prediction accuracy without requiring task-specific fine-tuning. Pretrained exclusively on synthetic data generated from structural causal models (SCMs), our model enables more flexible context utilization and more efficient capacity scaling. i) A new performance standard. Strong regression performance across benchmarks: Xiaomi-TabLDM ranks 1st on OpenML-CTR23 and 2nd on regression across TALENT, TabArena, and BCCO, demonstrating consistently strong regression performance across four complementary benchmark suites. Favorable performance--efficiency trade-off: Xiaomi-TabLDM combines strong predictive performance with substantially lower computational cost. For example, on TabArena regression, it achieves the second-highest Elo while using 82% less training time and 68% less prediction time than the top-ranked TabFM. ii) Large-scale synthetic pretraining. Xiaomi-TabLDM expands the coverage and diversity of synthetic tabular data used for pretraining. We also adopt a three-stage training strategy together with dual-stream feature grouping, lightweight Attention Residual, and sparse Mixture-of-Experts, enabling Xiaomi-TabLDM to learn richer feature interactions and expert specialization across diverse tabular tasks. iii) Test-time scaling. Xiaomi-TabLDM further extends tabular prediction through test-time compute scaling, where allocating additional computation at inference time consistently improves predictive performance over the base model.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.