← Back to Blog

Can AI weather forecasts be trusted during extreme weather, or are traditional models more reliable?

Can AI weather forecasts be trusted during extreme weather, or are traditional models more reliable?

AI weather forecasts can provide valuable guidance during extreme weather – sometimes highly accurate guidance. But no evidence supports trusting AI or traditional models across every situation. Reliability depends on the hazard, the forecast detail, the lead time and which specific model produced the number. Safety decisions should follow official warnings built from multiple sources.

There Is No Single Winner for Every Extreme Weather Event

Cyclone track and cyclone intensity are different forecasting tasks. A model that places a storm's landfall correctly can still underestimate peak wind speed by a category. A heatwave forecast can identify the right region and the right week while missing the record temperature at one station by 3°C.

The reason this matters: extreme weather forecast reliability isn't one number. Predicting whether a dangerous event may occur, where and when it arrives, how intense it becomes, and what happens at your specific location – these are four separate problems. A model can solve the first well and the fourth poorly. Collapsing all four into a single accuracy claim obscures exactly the gaps that cause harm.

How AI and Physics-Based Weather Models Produce a Forecast

The difference in how each model type generates a forecast explains a great deal about where each struggles.

Traditional Models Simulate Atmospheric Processes

Physics-based numerical weather prediction models calculate how temperature, moisture, pressure, wind and other variables change according to physical equations applied across a three-dimensional grid of the atmosphere. Because the equations encode atmospheric behaviour directly, these models can in principle simulate conditions that have never occurred in exactly the same form before – the physics applies regardless of whether the situation appears in any training dataset.

They are not automatically correct. Performance is limited by imperfect initial observations, the resolution of the computational grid, and the need to approximate small-scale processes that don't fit within that grid. A model constrained by coarse resolution will smooth out features that determine local intensity.

AI Models Learn From Weather Data

GenCast, developed by Google DeepMind, generates an ensemble of 50 forecasts at 15-day range in around 8 minutes on a single chip – roughly 1,000 times less energy than conventional systems for an equivalent task. That speed and efficiency come from identifying patterns in decades of historical observations and reanalysis data rather than calculating physical equations from scratch at each step.

GenCast has greater skill than ENS on 97.4% of 1,320 evaluated targets across lead times of 1 to 15 days. Different AI systems use different architectures, training datasets and objectives. GraphCast, Pangu-Weather, Fuxi and AIFS are not interchangeable – comparing results from one doesn't predict performance from another.

Extreme Weather Exposes Weaknesses Hidden by Average Scores

A model's aggregate performance across all weather conditions doesn't reveal how it handles the cases that matter most. Record-breaking events sit at the edge of what any model has learned, and that edge is precisely where the consequences of error are highest.

Record-Breaking Conditions May Fall Outside the Training Range

AI models tend to underestimate both the frequency and intensity of record-breaking events, and they underpredict hot records and overestimate cold records with growing errors for larger record exceedance. The mechanism is straightforward: a model trained to minimise average errors across all weather conditions learns to weight familiar outcomes. When a temperature or wind value exceeds anything in its training period, the model has no direct example to draw on – and pulling the prediction toward the familiar range produces a smaller average error than reproducing the full intensity.

Small but Intense Features Can Be Smoothed Out

Global forecast scoring systems tend to reward a broadly correct, moderately intense prediction more than a highly concentrated event placed slightly off target. A narrow rain band that falls 20 kilometres from its forecast position scores worse than a smooth, weaker rain area covering the whole region – even if the narrow band is the more physically accurate representation.

The result is smoother precipitation fields, weaker wind maxima and less distinct local extremes in model output. Severe thunderstorms, narrow bands of heavy rain and damaging gusts in a small area are all affected by this. Traditional global models face the same constraint at coarse resolution; the issue is not unique to AI, but it matters for any system evaluated on broad aggregate metrics.

Recent Studies Produce Different Answers for Good Reasons

Apparently conflicting conclusions about AI versus physics-based forecasting often reflect differences in what was actually tested – the systems evaluated, the benchmark chosen, the test period, the lead times examined and how extreme events were defined.

Physics-Based Models Led in a Test of Record-Breaking Extremes

A Science Advances study tested three leading AI weather models – Google DeepMind's GraphCast, Huawei's Pangu-Weather and Shanghai's Fuxi – against the European Centre for Medium-Range Weather Forecasts' physics model on roughly 246,000 record-breaking heat, cold and wind events from 2018 and 2020. For record-breaking weather extremes, the physics-based numerical model HRES from ECMWF consistently outperformed all three AI models across nearly all lead times.

The scope matters. The study provides evidence about those specific models tested against those specific events. It doesn't establish that every physics-based system outperforms every AI system for hurricanes, floods, rainfall totals, or future extremes beyond the test period.

Other AI Systems Perform Strongly on Some Extreme-Weather Measures

GenCast has greater skill than ENS on 97.4% of 1,320 targets evaluated, and better predicts extreme weather, tropical cyclones, and wind power production. When tested on extreme heat and cold, and high wind speeds, GenCast consistently outperformed ENS.

These results don't invalidate the Science Advances findings. GenCast is a probabilistic ensemble system compared against ENS, ECMWF's operational ensemble. The Science Advances study evaluated deterministic AI models against HRES, the high-resolution deterministic physics-based forecast. Different systems, different benchmarks, different performance measures – the studies are measuring related but distinct questions.

Track current conditions and forecast updates for your location on MeteoFlow as an extreme event develops.

Reliability Depends on the Hazard and Forecast Detail

Matching the right model type to the right task gives a more useful answer than declaring an overall winner. Different extreme weather forecast tasks involve different atmospheric scales, lead times and physical mechanisms.

Tropical Cyclone Track and Intensity

AI models have demonstrated strong skill predicting the broad movement and probable track of tropical cyclones. ECMWF's AIFS outperforms state-of-the-art physics-based models for tropical cyclone tracks, with gains of up to 20%. GenCast outperformed the ECMWF ensemble on predicting cyclone trajectories across all lead times from one to five days. Fast ensemble generation allows many possible track scenarios to be produced quickly, improving the representation of track uncertainty before landfall.

Intensity is a different question. Rapid intensification – when a storm's maximum winds increase by 35 knots or more within 24 hours – remains among the hardest forecasting problems for any system. Maximum wind speeds, storm inner structure and local landfall effects are harder to predict accurately regardless of model type. Track skill and intensity skill don't move together.

Heavy Rain, Flooding and Severe Thunderstorms

Local rainfall totals and flash-flood risk depend on processes that operate at scales smaller than any global model grid. Where exactly a rain band sets up, how terrain channels moisture, whether soil is already saturated – these details can shift the difference between a flood event and a wet afternoon. A model may correctly identify that heavy rain is likely over a region while placing the most intense cells 30 kilometres from where they actually form.

An official flood warning typically combines atmospheric model output with real-time radar data, river gauge measurements and hydrological models. The weather forecast is one input into that process, not the sole basis for the warning. A forecast showing high precipitation probability and a flood warning are related but distinct products.

Extreme Heat, Cold and Wind

The large-scale pattern – where a blocking high sits, which direction warm air is advecting from – often becomes predictable several days out even when the peak temperature at any single station remains uncertain. Broad and local are different things, and forecasts can get one right while missing the other.

A forecast that identifies a multi-day heatwave correctly but shows 41°C where the observed maximum hits 44°C is not simply "a bit low." That 3°C gap determines whether a heat health alert reaches its highest category, whether certain infrastructure thresholds activate, and whether outdoor work in exposed conditions must legally stop. The regional pattern was right. The number that drives the decision was wrong.

How Meteorologists Decide Whether an Extreme Forecast Is Credible

Professional forecasters rarely rely on a single model run. When several independent systems – AI and physics-based, deterministic and ensemble – repeatedly indicate the same high-impact scenario across consecutive updates, that agreement increases confidence meaningfully. The scenario is consistent with current observations, physically plausible, and stable across different methods reaching it independently.

When one model produces an isolated extreme outcome that disappears in the next update, and no other system supports it, the appropriate response is heightened monitoring rather than either dismissal or full preparation. The outlier may be noise; it may be an early signal. Tracking whether it reappears in subsequent runs is more informative than treating any single update as definitive.

Agreement across models increases confidence but doesn't guarantee the outcome. Disagreement is information about uncertainty – a reason to communicate a wider range of scenarios to decision-makers, not a reason to set the threat aside.

AI and Traditional Models Work Best as a Combined System

ECMWF launched AIFS into operational use on 25 February 2025, running side by side with its traditional physics-based Integrated Forecasting System. The AIFS generates forecasts approximately 1,000 times more energy-efficiently than conventional systems while the physics-based IFS provides the initial conditions and reanalysis data that trained AIFS in the first place.

The practical division is already in operation: AI generates forecasts and ensembles quickly, identifies patterns across large datasets, and supports frequent updates. Physics-based models provide atmospheric simulation that can represent conditions outside the historical record. Meteorologists interpret the output of both, connect it with local observations, and make the judgements that determine when a warning is issued.

Use MeteoFlow for Updates and Official Sources for Warnings

As an extreme event develops, conditions and forecasts both change – sometimes within hours. MeteoFlow's local weather forecast shows current conditions, temperature, wind, precipitation probability and short-term outlook for any selected location, updated as new model data arrives. Checking the same location across several updates shows whether the expected pattern is holding or shifting.

Follow the forecast for your location on MeteoFlow and check official warnings when extreme weather is developing nearby.

FAQ

Are official weather warnings created by AI?

Official warnings are issued by national meteorological services using a combination of model output – both AI and physics-based – observational data and expert meteorologist judgement. AI contributes to the analysis and forecast products that inform warnings, but the decision to issue a warning goes through professional review rather than automated model output alone.

Does a faster forecast automatically provide an earlier warning?

Speed of forecast generation and timing of an official warning are different things. A warning requires sufficient confidence that a dangerous event will occur – confidence that builds as models converge across consecutive updates and observations support the expected development. A faster model that reaches an uncertain conclusion earlier doesn't advance the warning timeline.

Can an AI weather forecast show how uncertain its prediction is?

Probabilistic AI systems such as GenCast generate ensembles of possible outcomes, making the spread of scenarios visible. Deterministic AI models produce a single forecast without explicit uncertainty information. The greater the spread of results across an ensemble, the greater the uncertainty in the forecast – this applies equally to AI and physics-based ensemble systems.

Should users choose a weather service based on the model it uses?

Most consumer weather services combine output from several models and apply post-processing before displaying a forecast. The underlying model matters less than update frequency, location precision, and whether the service links clearly to official warnings for hazardous conditions. A service that explains what it's showing and when it was last updated is more useful than one that claims a single model as its only source.