Stand-alone LSTMs cannot predict river flows above a limit set below their training peak

When fed extreme design rainfall, a stand-alone LSTM runoff model hits a ceiling well below flows it saw in training, and its runoff share falls as rain rises, while a hybrid model scales more sensibly.

Hydrology and Earth System Sciences 2 min read Peer-reviewed

Comparison of LSTM, hybrid and HBV simulated discharge for the 1 d events with the highest runoff in each of 25 catchments, showing the LSTM flattening out near its limit.
Figure B2 from Baste et al. (2025), CC BY 4.0. Resized from the original.

Why it matters

Deep learning models are widely used for rainfall–runoff modeling, but flood design and risk work needs credible behavior beyond past records. The study shows a stand-alone LSTM can underestimate extreme floods and respond in ways that conflict with hydrological expectation. Hybrid models with a conceptual core give more plausible extrapolation, though larger networks and more data help the LSTM only partly.

What they did

The authors trained LSTM ensembles on 196 Swiss catchments (CAMELS-CH), and a hybrid model that combines an LSTM with a conceptual HBV model. They took 201 high-rain days in 25 catchments and replaced the rain with 1, 3 and 5 d design precipitation for return periods of 50, 100 and 300 years. They compared the simulated discharge of both models and measured how saturated the LSTM’s internal units became. They also tested a larger LSTM and training on CAMELS-CH plus CAMELS-US.

Key findings

  • The LSTM ensemble has a theoretical prediction limit of 73 mm d−1, below the training maximum of 183 mm d−1. Its highest simulated value in the experiments was 60 mm d−1, even with rain of up to 1000 mm d−1.
  • For the three events with the most runoff, LSTM discharge rose on average 6 % from the 50-year to the 300-year event, while precipitation rose 39 %. The hybrid model rose 51 %.
  • LSTM runoff coefficients fell as rain intensified, contrary to hydrological expectation. The hybrid model’s response was roughly linear.
  • LSTM cell states never fully saturated, especially for 1 d events. The input and forget gates mostly kept the extreme rain from reaching the cell states.
  • A 256-unit LSTM trained on both datasets raised the design limit to 110 mm d−1, still far below the training maxima.

Limitations

  • Only precipitation was changed, so the link between rain and other inputs such as temperature is broken. The authors say this limits how deep the analysis can go.
  • Design precipitation is valid at station points but the models were trained on catchment averages, and only 25 catchments had a nearby station.
  • The experiments cannot say whether rare extremes, the squared-error loss or noisy data cause the underestimation.

Glossary

  • LSTM: Long short-term memory network, a neural network with gated memory cells for time series.
  • Theoretical prediction limit: The highest output an LSTM can produce, set by the weights of its final linear layer.
  • Hybrid model: A model where an LSTM sets the parameters of conceptual HBV hydrological models that compute the discharge.
  • Runoff coefficient: The share of rainfall that becomes streamflow.

Original paper

Unveiling the limits of deep learning models in hydrological extrapolation tasks

Sanika Baste, Daniel Klotz, Eduardo Acuña Espinoza, Andras Bardossy, Ralf Loritz

Hydrology and Earth System Sciences · 3 November 2025

Read the original paper Licence: see terms · doi:10.5194/hess-29-5871-2025

AI-generated summary of the original article; changes were made. Check the original before relying on it.