WeatherNext 3: Increasing resolution and performance of global weather models with raw observations
AuthorsStephan Rasp, Boris Babenko, Dominic Masters, Andrew El-Kadi, Samier Merchant, Guy Shalev, Ilan Price, Fred Zyda, Remi Lam, Sasha Shysheya, Matthew Willson, Stratis Markou, Shreya Agrawal, Suhani Vora, Mohammed Alewi Hassen, Sunny Mak, Tom R. Andersson, Megan Bela, Akib Uddin, Nofar Peled Levi, Ben Gaiarin, Ferran Alet, Aaron Bell, Peter Battaglia, Alvaro Sanchez-Gonzalez
Resources
WeatherNext 3 brings AI weather forecasting closer to operational models by combining raw satellite and station observations with hourly, high-resolution global predictions.
Key results
Degrees of spatial resolution for hourly single-level forecasts
Ensemble members generated for 15-day forecasts
Maximum reduction in 2-meter temperature CRPS versus WeatherNext 2
Maximum early-lead CRPS reduction against IMERG
Average first-week upper-level CRPS improvement over AIFS ENS v2
What the paper found
Google DeepMind’s WeatherNext 3 advances global AI weather forecasting by ingesting raw observations rather than relying solely on analysis fields. Built on an encode-process-decode architecture with Functional Generative Networks, it combines ERA5 and HRES-fc0-5 analyses with an 11-channel geostationary satellite mosaic, PARDIG and IMERG precipitation estimates, METAR and Mesonet station reports, and IBTrACS cyclone data. The model produces new forecasts every hour, predicts single-level variables at 0.1° resolution with hourly time steps, and generates 15-day, 64-member probabilistic ensembles optimized with the continuous ranked probability score. A continuous station head uses latent interpolation plus elevation and land-sea metadata to query 2-meter temperature and dewpoint at arbitrary locations and times, including stations withheld from training. On 2024 evaluations, station-head 2-meter temperature CRPS improved by up to 30% over WeatherNext 2, while PARDIG precipitation reduced CRPS by up to 60% against IMERG, 30% against MRMS, and 10% against rain gauges at early lead times. Hourly initialization produced a 2–3-hour effective lead-time gain for rapidly evolving precipitation. In a six-week 2026 real-time comparison, WeatherNext 3 improved upper-level CRPS by roughly 10% over ECMWF’s AIFS ENS v2 during the first forecast week. The remaining limitations include hexagonal spatial artifacts, temporal discontinuities at six-hour boundaries, and under-dispersed tropical-cyclone ensembles.
Original abstract
State-of-the-art AI weather models have shown impressive medium-range forecast skill and computational efficiency, but suffer two key shortcomings: their forecasts have lower spatial and temporal resolution than the best physics-based models and they are exclusively initialized with and trained on analysis data. As a result, they cannot directly make use of observations, and any biases in the analysis are inherited by the forecast. WeatherNext 3 addresses these shortcomings and establishes a new state-of-the-art for probabilistic medium-range forecasting skill. First, WeatherNext 3 generates new forecasts every hour (rather than every 6 hours like traditional global models) by ingesting low-latency geostationary satellite data. Second, WeatherNext 3's temporal and spatial resolution are on par with physics-based global models, with hourly time steps and 0.1 degree resolution for single-level variables, including solar radiation and cloud cover. Third, WeatherNext 3 moves beyond traditional analysis variables by learning to predict satellite-derived precipitation estimates, as well as tropical cyclone and station observations. Modelling sparse station data allows WeatherNext 3 to make 2m temperature and dewpoint predictions at any location and time, conditioned on local geographical features, with substantially lower error than competing global models, even when evaluated against unseen stations. Together, WeatherNext 3's capabilities move operational AI-based weather forecasting beyond emulating the traditionally distinct stages of data assimilation, forecasting and post-processing, which helps to further push the frontier of performance and granularity for global weather prediction.
Read the original paperMore in AI for Science
Browse all 43 papers →AI-guided high-throughput discovery of iridium- and ruthenium-free palladium-oxide catalysts for durable acidic oxygen evolution
Ken J. Jenewein, Faezeh Habib Zadeh, Xiaoxiao Wang, Gustavo Malkomes, Huafan Zhang, Natalie Page, Jae Jin Bang, Peter J. Santiago, Karla V. Contreras, Katherine K. Li, Allison Perna, Lorena M. Britton, Fahrettin Kilic, Kevin J. Cruse, Armin Taheri, Krishnanand Mallayya, Harley Quinn, Rebecca A. Durr, Peter A. Beaucage, John M. Gregoire, Rafael Gómez-Bombarelli
An AI-guided robotic lab discovered palladium-based catalysts that could make acidic water electrolysis more durable while reducing dependence on scarce iridium and ruthenium.
Discovery of radio emission from the exoplanet $β$ Pictoris b
Kevin N. Ortiz Ceballos, Edo Berger, Yvette Cendes
Astronomers have detected radio auroras from β Pictoris b, revealing that this distant giant planet has a magnetic field at least 1.25 kilogauss strong.
EurekaBench: Measuring Agentic Ability to Discover New Scientific Insights
Jiayi Geng, Zhengxuan Wu, Kevin S. Chen, Seungone Kim, Joseph Janssen, Zora Zhiruo Wang, Bhupalee Kalita, Runtian Gao, Aaron Ho, Andrew Oakleigh Nelson, Olexandr Isayev, Francisco Villaescusa-Navarro, Ching-Yao Lai, Howard Chen, Graham Neubig
EurekaBench tests whether AI agents can move beyond accurate prediction to uncover mechanisms and insights that genuinely advance scientific understanding.