NTH

Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings

AuthorsEvelyn Ma, Rama Kumar Pasumarthi, Kishwar Shafin, Mandar Sharma, Mimi Sun, Hamed Sadeghi, Dav M. Ebengo, Mbulayi Onesime, Rouslan Solomakhin, John Wamburu, William Ogallo, Aisha Walcott-Bryant, Sanxing Chen, Arbaaz Muslim, Yael Mayer, Ronald Ho, Roy Lee, Ruth Alcantara, Abdoulaye Diack, Monica Bharel, Lambert Rosique, Jeremy Amez-Droz, Christopher Haire, James Manyika, Yossi Matias, Niv Efron, Gautam Prasad, Shravya Shetty

September 6, 2026 3 min read
Watch on YouTube
The one-line take

PPE is an autonomous geospatial AI pipeline that finds the right data and models itself to make better predictions about health, disasters, food security, and disease spread.

Key results

76.8%
CDC health mean R²

Mean performance across 21 CDC health indicators, compared with 60.0% for a manual expert pipeline.

66.1%
Nigeria food-security downscaling R²

PPE performance when downscaling from state to local-government-area resolution, compared with 31.5% baseline R².

83.3%
Ebola hotspot Recall@10

PPE identified 15 of 18 newly invaded health zones across five weekly forecasts.

330
PDFM embedding dimension

Dimensionality of the Population Dynamics Foundation Model embedding used for socioeconomic and demographic representation.

64
AlphaEarth embedding dimension

Dimensionality of the satellite-derived AlphaEarth embedding used for ecological and land-use features.

What the paper found

The Planetary Prediction Engine, developed within Google Earth AI, converts a natural-language geospatial question into a complete predictive workflow. LLM orchestrators infer whether the task is spatial regression, super-resolution downscaling, spatial transmission, or epidemiological nowcasting; discover and rank data from Data Commons, Google Earth Engine, open-web repositories, and sources such as Gemini-processed news; then apply leakage controls, multimodal feature fusion, and automated model search across regularized linear models, gradient boosting, XGBoost, and multilayer perceptrons. Its core representation combines the 330-dimensional Population Dynamics Foundation Model, or PDFM, with 64-dimensional AlphaEarth satellite embeddings and dynamically selected covariates such as mobility, vegetation, food prices, and infrastructure. Across benchmarks, PPE reaches 76.8% mean R² on 21 CDC health indicators versus 60.0% for a manual expert pipeline, and 66.1% versus 31.5% when downscaling Nigerian food-security indicators from state to local-government-area resolution. During the 2026 DRC Bundibugyo Ebola outbreak, it achieves 83.3% Recall@10, identifying 15 of 18 newly invaded health zones across five weekly forecasts, a 10.3-percentage-point gain over the approximately 73% public baseline. The system’s novelty is joint optimization of data selection and model configuration, supported by feature gates, split-isolated imputation, spatial validation, overfitting guards, and self-correction, although the study also finds that high-resolution satellite features can introduce noise in some cross-scale prediction tasks.

Original abstract

Addressing critical global challenges, from food security and disaster risk to disease outbreaks and socio-economic vulnerability, demands high-fidelity geospatial modeling. However, building predictive planetary models remains bottlenecked by a fragmented data ecosystem, requiring manual data retrieval, multimodal data curation and fusion along with iterative model selection. We present the Planetary Prediction Engine (PPE), an autonomous AI system that executes this end-to-end workflow directly from natural-language queries. PPE synthesizes multimodal datasets on the fly, retrieving spatiotemporally relevant covariates across open-web and Earth observation platforms (Data Commons, Google Earth Engine) and fusing them with geospatial foundation model embeddings (PDFM, AlphaEarth). Simultaneously, it searches over task-tailored model architecture families with automated overfitting guards. Across diverse tasks, geographies, and scientific domains, PPE consistently outperforms state-of-the-art or manually tuned expert baselines. For US spatial regression, PPE improves mean $R^2$ across 21 CDC health indicators (76.8% vs. 60.0%), FEMA national risk indices (64.9% vs. 60.0%), and the Social Vulnerability Index (66.2% vs. 58.6%). For spatial downscaling in data-scarce settings, PPE integrates localized proxies to double baseline accuracy in Nigerian food security indicators ($R^2$ of 66.1% vs. 31.5%). For epidemiological nowcasting of the 2026 DRC Bundibugyo Ebola outbreak, PPE achieves a Recall@10 of 83.3% (identifying 15 of 18 newly invaded health zones across five weekly forecasts), a +10.3 percentage-point improvement over the public state-of-the-art modeling (~73%). By combining autonomous multimodal planetary data discovery with targeted model optimization, PPE lowers the technical barrier to planetary-scale analytics, enabling rapid, customized, expert-level deployment.

Read the original paper

More in AI for Science

Browse all 43 papers →
01Scientific Ai

AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution

Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli

An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.

Read analysis
03Scientific Ai

EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights

Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig

EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.

Read analysis