NTH

UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation

AuthorsYohann Perron, Guillaume Astruc, Nicolas Gonthier, Clement Mallet, Loic Landrieu

June 24, 2026 2 min read
Watch on YouTube
The one-line take

UniverSat makes one transformer backbone work across many Earth-observation sensors, resolutions, and modalities by learning a shared patch embedding through self-supervision.

Key results

7
training datasets

Number of heterogeneous datasets used for self-supervised pretraining

13
sensors

Number of sensors covered in pretraining

0.1-300
ground sampling distance range

Spatial resolution range in meters

1-150
timestamps range

Temporal depth range across datasets

1-396
spectral channels range

Spectral width range across inputs

What the paper found

UniverSat, developed by researchers at LIGM, IGN, CNES, and École des Ponts, is a resolution- and modality-agnostic Earth observation transformer that replaces the fixed ViT patch projector with a Universal Patch Encoder built from linear-complexity axial cross-attention. The model embeds arbitrary spatio-temporal-spectral patches from optical, radar, hyperspectral, and elevation sensors into a shared latent space without input resampling, then fuses co-registered modalities and reconstructs dense feature maps at a user-specified output resolution. Pretraining uses a multimodal self-supervised objective combining cross-modal contrast and latent multimodal masked modeling, with aggressive masking that removes about 90% of input atoms. A single model is trained on 7 datasets spanning 13 sensors, covering 0.1 to 300 m GSD, 1 to 150 timestamps, and 1 to 396 channels. In probing experiments, UniverSat-B reaches 94.5 accuracy on BrickKiln, 80.1 mIoU on Sen1Floods11, 47.9 mIoU on PASTIS-R, and 41.1 mIoU on AI4Farms, while using only 9K supervised parameters for segmentation probes versus 33M to 47M for decoder-based baselines. On SpectralEarth, it surpasses DOFA without being trained on EnMAP, and the ablation study shows that removing the Universal Patch Encoder drops PASTIS performance from 32.9 to 21.5 mIoU, confirming that the shared encoder and skip-connected high-resolution pathway are central to the model’s transferability.

Original abstract

Vision Transformers (ViT) dominate computer vision. However, their reliance on rigid patch projectors hinders transfer to Earth Observation (EO), where input modalities, scales, and resolutions vary widely. We introduce UniverSat, a ViT-style backbone built around a Universal Patch Encoder that maps patches from arbitrary spatial, spectral, and temporal resolutions, and from both optical and non-optical sensors, into a shared embedding space with a shared set of weights. This enables training a single model on heterogeneous multimodal corpora via self-supervision, yielding robust, sensor-agnostic spatial features. We validate this approach with strong results across classification and segmentation on standard EO benchmarks from GeoBench, PANGEABench, and SpectralEarth. Our code and models are available at https://github.com/gastruc/UniverSat.

Read the original paper

More in Self-Supervised Learning

Browse all 22 papers →
01Self Supervised

Self-Play Pretraining with Zero Data

Aditya Cowsik, Kfir Dolev, Michael Y. Li, G. Bruno De Luca, Nourya Cohen, Noah D. Goodman, Yoav Levine

A learner and an RL-powered program generator teach each other from scratch, producing synthetic data that enables surprisingly meaningful transfer to natural datasets.

Read analysis
02Self Supervised

Strategically Diverse Sampling for Self-Training

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Instead of training LLMs on many similar correct answers, this work shows that exposing them to diverse problem-solving strategies—even imperfect ones—can produce stronger models.

Read analysis