UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation
AuthorsYohann Perron, Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu
UniverSat makes one transformer backbone work across many Earth-observation sensors, resolutions, and modalities by learning a shared patch embedding through self-supervision.
Key results
Number of heterogeneous datasets used for self-supervised pretraining
Number of sensors covered in pretraining
Spatial resolution range in meters
Temporal depth range across datasets
Spectral width range across inputs
What the paper found
UniverSat, developed by researchers at LIGM, IGN, CNES, and École des Ponts, is a resolution- and modality-agnostic Earth observation transformer that replaces the fixed ViT patch projector with a Universal Patch Encoder built from linear-complexity axial cross-attention. The model embeds arbitrary spatio-temporal-spectral patches from optical, radar, hyperspectral, and elevation sensors into a shared latent space without input resampling, then fuses co-registered modalities and reconstructs dense feature maps at a user-specified output resolution. Pretraining uses a multimodal self-supervised objective combining cross-modal contrast and latent multimodal masked modeling, with aggressive masking that removes about 90% of input atoms. A single model is trained on 7 datasets spanning 13 sensors, covering 0.1 to 300 m GSD, 1 to 150 timestamps, and 1 to 396 channels. In probing experiments, UniverSat-B reaches 94.5 accuracy on BrickKiln, 80.1 mIoU on Sen1Floods11, 47.9 mIoU on PASTIS-R, and 41.1 mIoU on AI4Farms, while using only 9K supervised parameters for segmentation probes versus 33M to 47M for decoder-based baselines. On SpectralEarth, it surpasses DOFA without being trained on EnMAP, and the ablation study shows that removing the Universal Patch Encoder drops PASTIS performance from 32.9 to 21.5 mIoU, confirming that the shared encoder and skip-connected high-resolution pathway are central to the model’s transferability.
Original abstract
Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfer to Earth Observation (EO), where input modalities, scales, and resolutions vary widely. We introduce UniverSat, a ViT-style backbone built around a Universal Patch Encoder that maps patches from arbitrary spatial, spectral, and temporal resolutions, and from both optical and non-optical sensors, into a shared embedding space with a shared set of weights. This enables training a single model on heterogeneous multimodal corpora via self-supervision, yielding robust, sensor-agnostic spatial features. We validate this approach with strong results across classification and segmentation on standard EO benchmarks from GeoBench, PANGEABench, and SpectralEarth. Our code and models are available at https://github.com/gastruc/UniverSat.
Read the original paperMore in Self-Supervised Learning
Browse all 22 papers →Self-Play Pretraining with Zero Data
Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine
A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.
Strategically Diverse Sampling for Self-Training
Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata
Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
TT-VidT pretrains video models to focus on motion while preserving appearance, achieving strong action-recognition results with substantially lower compute.