NTH
AI research

SpectralEarth-FM: Bringing Hyperspectral Imagery into Multimodal Earth Observation Pretraining

AuthorsNassim Ait Ali Braham, Aaron Banze, Conrad M. Albrecht, Julien Mairal, Jocelyn Chanussot, Xiao Xiang Zhu

May 26, 2026 2 min read
Watch on YouTube
The one-line take

This work builds a large multimodal Earth observation foundation model that learns from hyperspectral, multispectral, radar, and thermal data together, aiming to improve both hyperspectral and general geospatial AI tasks.

Key results

about 2 million globally distributed locations, 25 million georeferenced patches, and more than 40 TB of data
SpectralEarth-MM dataset scale

The pretraining corpus co-locates hyperspectral, multispectral, radar, and land surface temperature observations across global footprints.

1.4
HSI benchmark average rank

SpectralEarth-FM-a achieved the best overall average rank on ten hyperspectral downstream benchmarks under the SpectralEarth protocol.

2.5
HSI benchmark average rank, single-branch

SpectralEarth-FM-s performed strongly on the same ten hyperspectral benchmarks, ranking behind the multi-branch variant but ahead of specialized baselines.

3.43
PANGAEA average rank

On the PANGAEA general EO suite, the multimodal variant ranked first overall across the evaluated tasks.

What the paper found

SpectralEarth-FM is the first multisensor Earth observation foundation model to bring spaceborne hyperspectral imagery into joint pretraining with multispectral imagery, synthetic aperture radar, and land surface temperature. Its key novelty is a hierarchical transformer that preserves hyperspectral structure through sensor-specific spectral tokenization, then fuses modalities with projected attention before a shared Hiera trunk. The model is pretrained on SpectralEarth-MM, a new corpus of about 2 million globally distributed locations, 25 million georeferenced patches, and more than 40 TB of data built from EnMAP, EMIT, and DESIS hyperspectral missions aligned with Sentinel-2, Landsat-8/9 OLI, Landsat LST, and Sentinel-1. Pretraining uses a JEPA-style latent prediction objective, with an EMA teacher producing stop-gradient targets from two global multimodal views and a student matching those targets from global and sensor-dropped local crops; a SIGReg penalty prevents representation collapse. On ten hyperspectral benchmarks under the SpectralEarth protocol, SpectralEarth-FM-a achieves the best overall average rank, 1.4, and SpectralEarth-FM-s reaches 2.5, beating specialized hyperspectral baselines such as Spectral ViT-L, HyperSIGMA, DOFA, and CARL. On the PANGAEA general EO suite, the multimodal variant also ranks first overall with an average rank of 3.43, showing that hyperspectral pretraining can improve transfer without sacrificing performance on standard optical and radar tasks.

Original abstract

Earth observation (EO) foundation models (FMs) are increasingly trained on multisensor data, spanning multispectral imagery (MSI), synthetic aperture radar (SAR), and derived geospatial layers, but hyperspectral imagery (HSI) remains underrepresented. Conversely, existing hyperspectral FMs are trained on HSI alone, leaving joint pretraining and fusion of HSI with co-located EO sensors unexplored. We introduce SpectralEarth-FM, a hierarchical transformer for multisensor EO input with heterogeneous spectral dimensionality. The architecture combines spectral tokenization for hyperspectral inputs, sensor-specific encoders, a cross-sensor fusion module, and a shared hierarchical encoder, enabling joint processing of HSI and lower-channel observations. To pretrain SpectralEarth-FM, we curate SpectralEarth-MM, a dataset that co-locates HSI from three spaceborne sensors (EnMAP, EMIT, DESIS) with Sentinel-2, Landsat-8/9 optical imagery, Landsat land surface temperature (LST), and Sentinel-1 SAR, over common geographic footprints. It comprises approximately 2M globally distributed locations, 25M georeferenced patches, and over 40TB of data. Pretraining uses a Joint-Embedding Predictive Architecture (JEPA)-style objective that matches representations between global views and single-sensor local views from the same location. We evaluate SpectralEarth-FM on hyperspectral downstream tasks and standard EO benchmarks following the PANGAEA protocol, achieving state-of-the-art results across both evaluation settings.

Read the original paper