AI models fed predicted ocean temperatures rival a physics model on hurricane seasons

AI atmosphere models driven by forecast ocean temperatures can match or beat a leading physics-based model for seasonal storm counts in the Atlantic and eastern Pacific, but not everywhere, and some of that skill may come from compensating errors.

EarthArXiv 2 min read Preprint

Why it matters

Seasonal storm outlooks usually need costly physics-based models. The results suggest cheaper AI models could add useful forecasts alongside them. The authors also warn that high skill scores alone do not show a model gets the physics right, so trust should depend on more than the scores.

What they did

The authors ran seasonal hindcasts for 1993–2025 covering July–November in three basins: western North Pacific, eastern North Pacific and North Atlantic. They compared GFDL-SPEAR, a coupled physics model, with a high-resolution physics model (HiFLOR-S) and with two AI models, ACE2 and NeuralGCM. The AI models were forced by predicted sea surface temperature and sea ice from SPEAR and from five other forecast centres. They scored basin totals and grid-point maps with rank correlation, and also checked the large-scale conditions that drive storm genesis.

Key findings

  • At zero-month lead in the North Atlantic, six of seven AI-based configurations reach correlations of 0.60–0.76, versus 0.59 for SPEAR.
  • At the grid-point scale, nearly every AI-based configuration exceeds SPEAR’s significant-skill area in the eastern North Pacific and North Atlantic, by up to 20–25 percentage points.
  • NeuralGCM forced by SPEAR SST kept significant skill at 4-month lead in the western and eastern North Pacific, where SPEAR did not (0.42 and 0.69 versus 0.27 and 0.36).
  • SPEAR stays unmatched for accumulated cyclone energy in the western North Pacific, and ACE2 produces almost no hurricane-strength storms.
  • Some AI skill may come from compensating errors. For example, one configuration (N CMCC) picked the wrong main genesis driver in the North Atlantic yet still scored well.

Limitations

  • There was no control run with persisted SST anomalies, so the benefit of using predicted SST over persisted SST is not quantified.
  • The 33-year record is short and the top configurations differ by only 0.05–0.15 in correlation, so rankings are indicative only. The environment diagnostics are correlational, not causal.
  • The AI models’ storm detection thresholds were tuned empirically. NeuralGCM could not run for January and February starts, and only one deterministic version of each AI model was used.

Glossary

  • NeuralGCM: A hybrid atmosphere model that pairs a physics-based dynamical core with neural networks for small-scale processes.
  • ACE2: A deep-learning atmosphere emulator trained on reanalysis data that steps forward in 6-hour increments.
  • Hindcast: A forecast run for past years, used to test how well a method would have done.
  • Rank correlation (RCOR): A measure of how well forecasts reproduce the observed ordering of years from quietest to most active.

Original paper

Skillful Seasonal Tropical Cyclone Prediction with AI Weather--Climate Models Forced by Dynamically Predicted SSTs

Hiroyuki Murakami, Baoqiang Xiang

EarthArXiv · 22 September 2026

Read the original paper Licence: see terms · doi:10.31223/x5880n

This paper is a preprint. It has not been peer reviewed, and its results may change.

AI-generated summary of the original article; changes were made. Check the original before relying on it.