WeatherNext 3
Vintage: 2026-09. Primary evidence is Google's Sep 3, 2026 Keyword / DeepMind blog plus the fetched developers.google.com WeatherNext 3 model card, as hydrated in 2026-09-03-weathernext-3-deepmind-google-research-sep-3-2026. Official @GoogleDeepMind X matches those text claims. Root + 5 km-resolution posts have untranscribed video (
x_video: true); claims here are from fetched blog/docs and post text only — do not treat video demos as grain. The paper PDF URL is known; the PDF body was not fetched. Brightband's live board was not snapshot-captured.
WeatherNext 3
One-line summary: DeepMind + Google Research's Sep 3, 2026 flagship global weather AI — live geostationary satellite mosaics as a direct hourly input, multi-resolution output down to ~5 km for station-trained temperature/dew point, rolled into Search / Gemini / Maps / Earth Engine plus BigQuery / GCS. Accuracy "50%" lines disagree across surfaces; do not flatten.
What it is
Google's Keyword / DeepMind blog dated Sep 03, 2026 presents WeatherNext 3 as DeepMind + Google Research's flagship global weather model: it learns from real-time observations, uses raw satellite data to produce a forecast every hour at high resolution, and claims it is "the most advanced and accurate global weather model to date, according to independent live evaluations by Brightband." Prefer Sep 3 for announced / productized. The developers model card's Release Date field says August 2026 — keep that as the docs table value, not as a second announce date.
This page records issuer-hydrated claims from the fetched blog, the fetched developers.google.com model card, and official X text. It does not re-file same-day gemini-3-8-flash, k2-horizon, qwen-3-8-max-0902, mostik-ai, or claude-fable-5-1 permalinks. It does not cite the held OpenAI video-only post or the Sep 1 world-labs-atlas pointer named as out-of-window in the source.
Why it matters to this thread
Frontier model releases and AI applied beyond coding (multimodality / real-world observation) are in-scope. This is the first dated WeatherNext / weather-foundation-model page in the thread. Live satellite mosaics as an hourly init are a lab datapoint on ai-real-world-data-grounding — a vertical weather application, not a confirmation of Marshall's general "large earth models" thesis.
Key facts (from 2026-09-03-weathernext-3-deepmind-google-research-sep-3-2026)
Product announcement (Sep 3)
- From 2026-09-03-weathernext-3-deepmind-google-research-sep-3-2026 (Google Keyword / DeepMind blog, Sep 03, 2026): WeatherNext 3 learns from real-time observations, uses raw satellite data to produce a forecast every hour at high resolution, and claims it is "the most advanced and accurate global weather model to date, according to independent live evaluations by Brightband."
- From the same source (same blog): product surfaces named as starting that day — Google Search, Gemini app, Google Maps, Google Maps Platform Weather API, Google Earth Engine. Developer path named: BigQuery, Earth Engine, bulk download from Google Cloud Storage.
- From the same source (@GoogleDeepMind X, 2026-09-03, 2095528012791902536): root announce with @GoogleResearch — learns from real-world real-time observations. That post has video; the claim here is post text only.
- From the same source (2095528025978765787): powers Search / Gemini / Maps / Maps Platform; BigQuery, Earth Engine, GCS — links the blog.
Architecture and resolution (issuer)
- From the same source (blog): Functional Generative Network (FGN) mesh transformer that ingests live 1-hour geostationary satellite mosaics plus traditional historical analysis, and outputs dense gridded fields, discrete cyclone tracks, and station-level sparse coordinates.
- From the same source (blog, vs WeatherNext 2): surface temperature/moisture visualized at 5 km, other surface vars at 10 km, atmospheric vars (e.g. wind) at 25 km, described as "roughly five times sharper" than WeatherNext 2's 25 km / 6-hour grid.
- From the same source (@GoogleDeepMind, 2095528015778164838): hourly refresh vs six-hour traditional.
- From the same source (2095528019062304828): station training + 5× temp resolution 25 km → 5 km. That post has video; the claim here is post text only.
Fetched developers model card operational specs (same source; docs table):
| Attribute | Docs value |
|---|---|
| Release Date field | August 2026 (docs table; product blog + DeepMind X are 2026-09-03) |
| Spatial | 0.05° (~5 km) stations · 0.1° (~10 km) gridded surface · 0.25° (~25 km) pressure levels |
| Temporal | 1-hour timesteps |
| Horizon | 15 days (360 h) on 6-hourly cycles; 48 h on interim hourly runs |
| Init frequency | Every hour (24/day) |
| Ensemble | 64 members |
| Architecture | FGN mesh transformer |
| Inputs | Live geostationary satellite mosaics + ECMWF HRES analysis |
| Training data | ERA5 / HRES-fc0, IMERG, station observations, geostationary mosaics |
Paper PDF linked from the model card: https://storage.googleapis.com/deepmind-media/papers/weathernext_3.pdf. URL present on the docs page; PDF body not fetched — cite the link, do not invent paper abstract claims.
Precipitation — four "50% / CRPS" lines, not one number
Keep these as separate issuer claims. Do not collapse into a single 50%.
- From the same source (blog, medium-range precip CRPS, early lead times): up to 60% vs IMERG, 30% vs MRMS, 10% vs rain gauges.
- From the same source (blog, consumer product copy for day-or-more-ahead planning): up to 50% more accurate precipitation forecasts, largest gains where forecasts were historically less reliable.
- From the same source (developers model card): trains against ECMWF reanalysis, NASA IMERG, and Google's satellite-radar precip reanalysis; "up to a 50% reduction in Brier score and CRPS compared to numerical weather prediction baselines when evaluated against IMERG observations."
- From the same source (@GoogleDeepMind, 2095528022208086478): up to 50% reduction in precip error.
Clean-energy variables (issuer)
- From the same source (blog): 100-meter wind (turbine-height), high-res cloud cover, sun radiation / solar irradiance for solar farms.
- From the same source (developers model card):
wind_speed_100m, SSRD/FDIR solar components, and cloud-layer fractions on the 0.1° grid.
Brightband ranking (issuer / advocate, board not captured)
- From the same source (blog): "most advanced and accurate … according to independent live evaluations by Brightband."
- From the same source (@FerranAlet, 2026-09-03, 2095537054004183150): claims best global model on Brightband's live leaderboard and links https://owb.brightband.com/. Author-adjacent, not lab primary. Treat as issuer/advocate claim until a dated leaderboard snapshot is filed. Do not invent ranks or scores.
What this source does not establish
- No fetched paper body. The PDF URL is known; contents were not fetched. No paper numbers.
- No Brightband board snapshot.
owb.brightband.comwas not captured this pass. "Best / most accurate" stays contested. - Videos untranscribed. DeepMind root (2095528012791902536) and 5 km-resolution (2095528019062304828) posts are demo only until
/transcribe-clipping. - Not a re-file of gemini-3-8-flash, k2-horizon, qwen-3-8-max-0902, mostik-ai, or claude-fable-5-1.
- Held out of this source: multi-vendor outage chatter; @OpenAI video-only; Sep 1 world-labs-atlas pointer.
- No ticker, 8-K, or stock-market tag.
Contradictions / tensions
- Date label: docs model card says Release Date August 2026; public blog + DeepMind X are 2026-09-03. Prefer Sep 3 for announced / productized; keep August as the docs table value. Not reconciled.
- "50%" is not one number: blog product copy ("50% more accurate precip" for users), X ("50% reduction in error"), docs ("50% reduction in Brier score and CRPS vs NWP baselines on IMERG"), and blog research CRPS (60% / 30% / 10% vs IMERG / MRMS / gauges). Do not collapse.
- Brightband "best / most accurate": blog + Ferran Alet assert it; live board was not snapshot-captured. Flag as contested.
- Vertical vs general grounding: live satellite-init weather is a lab product, not evidence that Earth-observation data lifts general frontier capability. See ai-real-world-data-grounding.
Open questions
- What does the unfetched paper add (methods, tables, independent eval protocol) that the blog and model card omit?
- Does a dated Brightband board scrape support, narrow, or contradict the "best / most accurate" line?
- Are the four precip "50% / CRPS" surfaces the same metric family, or four different scores?
- What do the untranscribed DeepMind videos show that the text posts do not?