TabPFN-3.5: Technical Report
AuthorsBenjamin Jäger, Nick Erickson, Léo Grinsztajn, Felix Birkel, Klemens Flöge, Oscar Key, Kürşat Kaya, Jonas Kübler, Adèle Frankel, Tobias Schröder, Anurag Garg, Jan Hendrik Metzen, David Salinas, Simon Bing, Kristina Collins, Tuana Çelik, Vahid Balazadeh, Lydia Sidhoum, Tomás Pereda, Brendan Roof, Andrej Tschalzev, Siyuan Guo, Philipp Singer, Lennart Purucker, Jake Robertson, Marie Salmon, Philipp Jund, Jerry Chen, Diana Kriuchkova, Arthur Cahu, Eliott Kalfon, Adrian Hayler, Georg Grab, Vitor Monteiro, Lilly Wehrhahn, Dominik Safaric, Clara Cornu, Alan Arazi, Rylee Grace, Simone Alessi, Mihir Manium, Bernhard Schölkopf, Yann LeCun, Madelon Hulsebos, Sauraj Gambhir, Noah Hollmann, Frank Hutter
Resources
TabPFN-3.5 advances tabular foundation models with stronger multimodal and real-world data handling, faster inference, and scalable reasoning modes.
Key results
TabPFN-3.5 ranks 1st of 89 methods on TabArena.
Mean win rate against the strongest other tabular foundation model across seven benchmarks.
Recommended maximum training-table size.
Parameter count of TabPFN-3.5.
TabPFN-3.5-Fast runs up to 3 times faster than TabPFN-3.
TabPFN-3.5 mean rank on CRPS across ScoringBench.
What the paper found
TabPFN-3.5 is a tabular foundation model designed to extend in-context prediction beyond clean i.i.d. datasets to temporal and grouped splits, high-cardinality categorical variables, wide tables, strings, text, images, relational data, and predictive distributions. On TabArena, it ranks 1st of 89 methods, while the broader TabPFN-3.5 family ranks first across seven benchmarks, with mean win rates of 78% against other tabular foundation models and 89% against non-foundation models such as XGBoost, CatBoost, and tuned MLPs. The model supports up to 1M rows and 6K features with 220M parameters, using Fourier value encodings, in-context ECDF features, shared classification and regression representations, and a synthetic prior targeted at grouped, wide, and high-cardinality data. Its transformer width increases from 512 to 1024 dimensions, while grouped-query attention keeps cached prediction latency close to TabPFN-3. TabPFN-3.5-Fast runs up to 3 times faster than TabPFN-3, and TabPFN-3.5-Thinking scales inference-time computation without LLMs or external data. On ScoringBench, TabPFN-3.5 achieves a 2.85 mean rank and improves over TabPFN-3 on 85 of 101 datasets, demonstrating gains not only in point prediction but also in calibrated predictive distributions. The Plus variants add native text and date handling and FP8 attention for deployment-oriented multimodal inference.
Original abstract
We introduce TabPFN-3.5, our new flagship Tabular Foundation Model. It significantly outperforms its predecessor, TabPFN-3, and all existing baselines across a broad range of tabular problems. TabPFN-3.5 sets a new state of the art on standard tabular prediction in TabArena, and extends it to the data practitioners encounter in practice: non-i.i.d. data with temporal or grouped splits, tables with strings, text and images, high-cardinality categorical features, and wide tables with many features. These gains carry over to our task-specific harnesses: state of the art on relational data and stronger time-series forecasting. For faster inference, our variant TabPFN-3.5-Fast runs up to 3x faster than TabPFN-3 while keeping most of the accuracy gains. In addition, we upgrade TabPFN-3.5-Plus, expanding our multimodal capabilities with advanced text and date handling alongside proprietary inference optimizations. Finally, we release a new version of our Thinking mode, TabPFN-3.5-Thinking, which scales inference-time computation to push the state of the art further. It benefits from our stronger base model and from inference-time improvements that make it up to 12x faster than TabPFN-3-Thinking.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.