EXAONE Tabular 1.0 : Technical Report
AuthorsMoonjung Eo, Min-Kook Suh, Hye-Seung Cho, Jiwon Kim, Seoyoon Kim, Sangjun Nam, Soonyoung Lee
Resources
EXAONE Tabular shows that a compact, causally synthetic-pretrained foundation model can deliver strong tabular predictions without fine-tuning, at far lower cost than much larger alternatives.
Key results
Parameter count of the EXAONE Tabular classification checkpoint.
Parameter count of the EXAONE Tabular regression checkpoint.
Default inference time in seconds per 1,000 samples.
Mean classification accuracy across the BCCO benchmark.
Mean regression R² across BCCO datasets.
Mean regression R² across TALENT datasets.
What the paper found
EXAONE Tabular is a compact tabular foundation-model family for classification and regression that predicts through in-context learning, without dataset-specific gradient updates. Its central innovation, the Cross-Axis Summary Transformer, or CAST, preserves cell-level representations across 12 Transformer layers by interleaving feature-axis attention within each row with support-conditioned item-axis attention within each feature, using item-summary and feature-summary tokens. The separately trained models contain 20.81M parameters for classification and 21.11M for regression, and are pretrained entirely on synthetic structural-causal-model, or SCM, tasks spanning nonlinear mechanisms, categorical variables, noise, and missingness. Missing values are handled natively rather than imputed, while regression predicts 999 conditional quantiles to support both point estimates and uncertainty distributions. On TabArena, the default classifier ranks first overall, and the model runs at 0.605 seconds per 1,000 samples; its regression performance reaches the regime of Google’s 1.64B-parameter TabFM at roughly 1/11 the inference cost. EXAONE Tabular ranks second in classification and first in regression on BCCO and TALENT: on BCCO it records 0.792 mean accuracy and 0.799 mean R², while on TALENT it achieves 0.736 mean R². On ScoringBench, it leads mean-rank evaluations for R², RMSE, and CRPS, outperforming alternatives including TabPFN-3, TabICLv2, XGBoost, LightGBM, and CatBoost in the reported comparisons.
Original abstract
EXAONE Tabular is a compact tabular foundation model family for classification and regression via in-context learning, producing predictions without dataset-specific gradient updates. Pretrained exclusively on a synthetic structural-causal-model (SCM) prior, its central contribution is an architecture-centered redesign of tabular in-context learning. Rather than compressing features into a fixed row embedding before a separate row-level learner, EXAONE Tabular interleaves feature-axis attention within each item with support-conditioned item-axis attention within each feature at every Transformer layer, mediated by item-summary and feature-summary tokens. Across four public benchmarks, EXAONE Tabular combines strong predictive performance with high efficiency. On TabArena, its 20.81M-parameter classification model ranks first overall, surpassing tuned ensembles and 4-hour AutoML pipelines, while regression reaches the performance regime of the 1.64B-parameter TabFM at roughly 1/11 the inference cost. On BCCO and TALENT, EXAONE Tabular ranks second in classification and first in regression. On ScoringBench, it achieves the best mean rank for both point-estimation and predictive-distribution quality, leading the $R^2$, RMSE, and CRPS evaluations. Together, these results establish EXAONE Tabular as a state-of-the-art compact tabular foundation model family, combining strong predictive performance across classification, point regression, and probabilistic regression with an efficient model design.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.