LSTM beats Transformers on standard hydrology tasks, but attention models win on harder ones

In a benchmark of 11 Transformer-type models against an LSTM, the LSTM was best at regression and short forecasts, while attention-based models did better at long-horizon autoregression and zero-shot forecasting.

Hydrology and Earth System Sciences 2 min read Peer-reviewed

Bar charts show how much each model's error differs from the LSTM baseline at 1, 7, 30 and 60 day horizons, with green bars for models that beat LSTM and blue for models that do worse.
Figure 3 from Liu et al. (2025), CC BY 4.0. Resized from the original.

Why it matters

Hydrologists choosing a deep learning model can see that newer attention-based designs do not automatically beat LSTM on routine tasks. They may help when predicting far ahead from past data alone, at unmonitored sites, or at flow extremes. Pre-trained models that need no local training look promising for data-poor regions, though the test was small.

What they did

The authors built one framework that handles data processing, model management and task running. They tested 11 attention-based models and an LSTM baseline on five daily datasets: US streamflow, global streamflow, soil moisture, snow water equivalent and dissolved oxygen. Tasks ranged from regression and forecasting to autoregression, spatial cross-validation and zero-shot forecasting. The zero-shot test used pre-trained language models and time series models on seven randomly chosen basins, with 90 d of history as input.

Key findings

  • In regression, LSTM was best on four of five datasets. On global streamflow its median KGE was 0.75, 0.11 above the best Transformer-type model.
  • For snow water equivalent, the Non-stationary Transformer slightly beat LSTM (KGE 0.88 versus 0.87).
  • In autoregression at 7 d and longer, LSTM fell sharply (KGE 0.15 at 7 d, -0.03 at 30 d). Attention models held up better, for example Pyraformer at 0.32 for 7 d.
  • In spatial cross-validation on CAMELS, LSTM’s KGE dropped from 0.80 to 0.62, while Crossformer’s went from 0.73 to 0.63.
  • In zero-shot tests at 7 d, TimeGPT reached KGE 0.68 against 0.50 for a trained LSTM, but it still missed peak flows.

Limitations

  • Zero-shot tests used only seven randomly chosen basins. The authors say language models do not truly capture hydrologic relationships, and they scored poorly on regression and forecasting.
  • PatchTST and TimesNet were run only for autoregression because training time grew too much with basin count. LSTM used hyperparameters from earlier studies, while the Transformer models were tuned on CAMELS.
  • Even the best attention model scored below 0.5 in KGE and R2 at long horizons, so long-range prediction stays weak.

Glossary

  • LSTM: Long Short-Term Memory, a recurrent neural network that uses gates to decide what to keep or forget across a time series.
  • Zero-shot forecasting: Making predictions with a model that was never trained on the target data.
  • KGE: Kling–Gupta Efficiency, a score combining correlation, bias and variability; 1 is perfect.
  • Autoregression: Predicting future values from past values of the same variable only, with no weather inputs.

Original paper

From RNNs to Transformers: benchmarking deep learning architectures for hydrologic prediction

Jiangtao Liu, Chaopeng Shen, Fearghal O'Donncha, Yalan Song, Wei Zhi, Hylke E. Beck, Tadd Bindas, Nicholas Kraabel, Kathryn Lawson

Hydrology and Earth System Sciences · 1 December 2025

Read the original paper Licence: see terms · doi:10.5194/hess-29-6811-2025

AI-generated summary of the original article; changes were made. Check the original before relying on it.