Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees
AuthorsAditya Tanna, Nassim Bouarour, Mohamed Bouadi, Vinay kumar Sankarapu, Pratinav Seth
Resources
This paper shows how to shrink slow tabular foundation models into fast CPU-friendly decision-tree ensembles, preserving most of the accuracy while cutting inference to milliseconds.
Key results
Distilling TabICLv2 into XGBoost across 153 classification datasets yields a macro-mean ROC-AUC of 0.882.
The TabICLv2→XGBoost student retains 96.5% of the teacher’s AUC.
The distilled XGBoost student runs on CPU in 1.9 ms, compared with 151 ms for the teacher.
The paper reports a 79× speedup for TabICLv2→XGBoost relative to the teacher.
TabICLv2→XGBoost beats a tuned CatBoost baseline on 51% of the 153 benchmark datasets.
Across tasks with 21 or fewer features, distillation improves CatBoost by 0.011 AUC on average.
What the paper found
Pocket Foundation Models shows that tabular foundation models such as TabICLv2, TabPFNv2.6, LimiX, and Orion-MSP can be compressed into CPU-ready gradient-boosted tree students without most of their accuracy loss, but only if the teacher is labeled out of fold. The key novelty is that in-context learning teachers leak labels when they score examples from their own training context, collapsing soft targets to near-one-hot vectors; stratified 5-fold out-of-fold labeling restores usable inter-class structure for distillation. Across 153 classification datasets from TALENT, OpenML-CC18, TabZilla, and TabArena, TabICLv2→XGBoost reaches 0.882 macro-mean ROC-AUC, retaining 96.5% of the teacher’s AUC while running in 1.9 ms on CPU, versus 151 ms on GPU for the teacher, a 79× speedup and 38× to 860× across all teacher-student pairs. This student beats a tuned CatBoost baseline on 51% of datasets with a Wilcoxon p-value of 0.0008. The strongest teacher always produces the strongest student, so teacher ranking transfers exactly to student ranking, and gains concentrate on low-dimensional tasks with 21 or fewer features, where distillation improves CatBoost by 0.011 AUC on average; above 21 features the gain falls to 0.001. Multi-teacher averaging helps MLP students by +0.006 AUC, but is negligible for tree students, and distillation fails when the teacher itself underperforms CatBoost on high-dimensional data.
Original abstract
A fraud scorer needs to answer in under 2 ms. The best tabular foundation models (TFMs) take 151-1,275 ms on GPU. We close this gap by distilling the TFM offline into an XGBoost or CatBoost student that runs natively on CPU. The central obstacle is specific to in-context learning (ICL) teachers: they leak labels when scoring their own training set, so the soft targets collapse to near-one-hot vectors with no inter-class structure left to distill. Stratified out-of-fold (OOF) teacher labeling prevents this. Across 153 classification datasets drawn from TALENT, OpenML-CC18, TabZilla, and TabArena, distilling TabICLv2 into XGBoost gives 0.882 macro-mean AUC (96.5% of teacher AUC) at 1.9 ms on CPU, a 38x to 860x speedup across teacher-student pairs with a statistically significant edge over a tuned CatBoost baseline (Wilcoxon p = 0.0008; 51% win rate). Four further findings: teacher rank transfers exactly to student rank; gains concentrate on low-dimensional data (< 21 features: +0.011 over CatBoost vs. >21 features: +0.001); multi-teacher averaging helps MLP students (+0.006, p = 0.003) but adds less than 0.001 for tree students; and on high-dimensional tasks where the teacher itself trails CatBoost, distillation makes things worse rather than better. The full pipeline is open-sourced as part of the TabTune library.
Read the original paper