When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning
AuthorsDaehwan Kim, Haejun Chung, Ikbeom Jang
This paper improves self-supervised learning on medical tabular data by adaptively refining feature discretization during training, leading to better representations without hand-tuned binning.
Key results
medical tabular datasets in the benchmark
epochs without improvement before Feature-Wise Plateau Trigger refinement
linear-probing average rank for Adaptive Binning
linear-probing average rank for fixed-binning BinRecon under masking
best linear-probing result on Heart Failure Clinical Records
What the paper found
This paper from Hanyang University and Hankuk University of Foreign Studies proposes Adaptive Binning, a self-supervised learning pretext for medical tabular data that replaces fixed global quantile discretization with a training-adaptive, feature-wise coarse-to-fine curriculum. The method specifies when to refine with a Feature-Wise Plateau Trigger after 5 epochs of saturation, where to refine with Dispersion-Informed Gain-based Splitting using the product of variance reduction and representation-space dispersion gain, and how to supervise with HORD, which combines categorical cross-entropy with ordinal soft-target cross-entropy plus mean-variance regularization. On a benchmark of 8 public medical tabular datasets spanning binary, nominal and ordinal classification, and regression, the method achieves the best average linear-probing rank, 1.50, versus 6.31 for fixed-binning BinRecon under matched masking, and it remains strong after fine-tuning with MLP, ResNet, TabNet, FT-Transformer, and T2G-Former. The strongest linear-probe result reaches 96.88% accuracy on Heart Failure Clinical Records and 70.51 RMSE improvement on Parkinsons Telemonitoring? no, 70.51 is classification accuracy on Maternal Health Risk; importantly, the method’s gains persist without dataset-specific discretization tuning, suggesting that learning-coupled discretization is a transferable inductive bias rather than a probe-specific artifact.
Original abstract
Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because reliable labels often require costly expert adjudication, even though structured clinical variables are routinely available in tabular form. Self-supervised learning can leverage these unlabeled tables, and recent binning-based pretexts offer a promising inductive bias, but existing objectives fix a single global quantile discretization and apply feature-agnostic supervision. We propose Adaptive Binning, a training-adaptive discretization pretext for tabular SSL that couples discretization to learning through a feature-wise coarse-to-fine curriculum. Motivated by the spectral bias of neural networks and the principles of curriculum learning, our method progressively refines discretization per feature upon plateau detection and selects representation-aware splits to jointly improve value-space concentration and representation-space coherence. A heterogeneity-aware objective unifies categorical reconstruction with ordinal supervision for numerical features, and experiments on public medical tabular datasets under unified evaluation protocols show consistent gains for linear probing and fine-tuning without dataset-specific discretization tuning. We further introduce a medical tabular SSL benchmark with standardized protocols to support reproducible progress in this underexplored domain. Our code is available at https://github.com/labhai/Adaptive-Binning.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.