TESSERA v2: Scaling Pixel-wise Earth Foundation Models
AuthorsZhengpeng Feng, Sadiq Jaffer, Ira Shokar, Jovana Knezevic, Mark Elvers, Clement Atzberger, Robin Young, Aneesh Naik, Niall Robinson, Andrew Blake, David Coomes, Anil Madhavapeddy, Srinivasan Keshav
TESSERA v2 shows how to scale Earth-observation foundation models more effectively, revealing that downstream results—not pretraining loss—should guide model selection and that bigger encoders plus distillation can produce compact, highly competitive embeddings.
Key results
Controlled BARLOW TWINS pretraining runs evaluated on downstream tasks.
NVIDIA GH200 superchips used for the scaling sweep.
Additional compute required when selecting models by pretraining loss rather than downstream performance.
TESSERA v2-1B-M score across the 29-task benchmark.
Parameter count of the TESSERA v2-1B-M student.
Performance retained by the 16-dimensional prefix relative to the 128-dimensional embedding.
What the paper found
TESSERA v2, from the University of Cambridge, NVIDIA, and dClimate Labs, presents a downstream-driven scaling study for pixel-wise Earth-observation foundation models. The team ran 395 controlled BARLOW TWINS pretraining experiments on 1,024 NVIDIA GH200 superchips, evaluating every model across 15 downstream tasks. Pretraining loss was a weak proxy for utility, and selecting models by loss required 254% more compute to reach the same downstream score. Instead, the compute-optimal recipe scales encoder capacity and training data together while keeping the projector fixed. Applying this rule, the researchers trained a 1B Sentinel-1/2 temporal teacher and distilled it into compact students, including the 21M-parameter TESSERA v2-1B-M. Across a 29-task benchmark, this student achieved a composite score of 0.611, outperforming open and proprietary systems including TESSERA v1, AlphaEarth, and OlmoEarth. TESSERA v2 also combines knowledge distillation with Matryoshka representation learning: a 16-dimensional embedding prefix retains 92% of the full 128-dimensional performance while using 1/8 of the storage. The result is an analysis-ready embedding family that trades model capacity, storage, and accuracy without retraining, with global annual Sentinel-1/2 embeddings intended for delivery through GeoTessera.
Original abstract
Pixel-wise Earth-observation (EO) foundation models are now achieving state-of-the-art performance via generated spatial embeddings. However, how these models scale and how best to spend a pretraining budget remain poorly understood. We present the largest controlled scaling study for EO to date: 395 training runs on 1,024 GH200 superchips within a fixed pixel-wise Barlow Twins family, each evaluated on 15 downstream tasks. We find that pretraining loss barely predicts downstream performance (|Pearson r| < 0.2), so selecting models by loss wastes a large share of the compute. We also find that, as the training budget grows, the encoder and the data should grow together while the projector stays fixed, which gives a simple rule for allocating compute. Using this rule, we train a family of pixel-wise models (0.5B and 1B, with a 2B model in training) and distill them into compact students for embeddings-as-data deployment. The 21-million-parameter distilled TESSERA v2-1B-M in aggregate outperforms all open and proprietary models tested, some of which are orders of magnitude larger. These students produce Matryoshka representations that are inexpensive to serve: a 16-dimensional prefix keeps 92% of the full 128-dimensional performance at 1/8 of the storage. Upon completion of training we plan to release v2 global embeddings covering 2017-2025. Together, these results give a concrete, empirically grounded recipe for scaling pixel-wise EO foundation models: train large encoders, select by downstream performance, and distil into flexible student models. All code will be released at https://github.com/ucam-eo/tessera.
Read the original paperMore in Foundation Models
Browse all 47 papers →How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI-generated web text can help language models at first, but beyond a tipping point it degrades performance on human writing, making data filtering and separate evaluation increasingly important.
TabFM: A Zero-Shot Foundation Model for Tabular Data
Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan, Rajat Sen, Yichen Zhou, Deqing Fu, Samet Oymak, Abhimanyu Das
TabFM is a large synthetic-data-trained model that aims to make accurate tabular predictions instantly, without retraining for each new dataset.
When Do Biological Reasoning Models Use Their Biological Inputs?
Ada Fang, Nikitha Thoduguli, Lukas Fesser, Hanlin Zhang, Sham M. Kakade, Marinka Zitnik
The study finds that many biological reasoning systems appear to succeed without meaningfully using the biological inputs they were designed to reason over.