Feeding past streamflow into an LSTM sharply improves Western U.S. streamflow estimates
Giving an LSTM recent streamflow observations raises its accuracy most at the daily scale, while snow observations help only at the monthly scale, mainly in snow-dominated basins.
Why it matters
Water managers in the arid Western U.S. need accurate short- and long-range streamflow forecasts. This simple approach adds new observations as inputs and needs no heavy data assimilation. The results suggest it could support automated forecasting, but only retrospective tests were done.
What they did
The authors trained LSTM models on 646 Western U.S. basins using daily weather forcings and basin attributes, with data from 1983 to 2002. They tested a data integration version that also receives streamflow or snow water equivalent (SWE) from earlier time steps. They used lags of 1–10 days and 1–6 months. They scored the models on 2003–2022 with Kling-Gupta Efficiency (KGE) and related metrics, using observed weather rather than forecasts.
Key findings
- The baseline LSTM already had a median KGE of 0.80 at both daily and monthly scales.
- Adding 1 d lagged streamflow raised the daily median KGE to 0.96. With a 10 d lag it was still 0.89.
- At the monthly scale, 1-month lagged streamflow raised the median KGE from 0.80 to 0.86.
- Daily SWE integration did not improve KGE. Monthly SWE with a 1-month lag raised the median KGE to 0.82 across all basins.
- SWE helped more in snow-dominated basins during the April to July snowmelt season. The overall ranking was daily streamflow, then monthly streamflow, then monthly SWE, then daily SWE.
Limitations
- The tests used observed weather inputs instead of real forecasts, so the gains are likely an upper bound for real forecasting.
- The models give single deterministic estimates, and uncertainty from inputs and training data was not analyzed.
- Very dry southern basins with flash floods saw little or no improvement.
Glossary
- Data integration (DI): Feeding past observations directly into the model as extra inputs so it learns how to use them.
- LSTM: Long short-term memory network, a neural network built to learn from long sequences of data.
- Kling-Gupta Efficiency (KGE): A score combining correlation, bias and variability errors, where 1 is perfect.
- Snow water equivalent (SWE): The amount of water held in the snowpack.
Original paper
Improving streamflow simulation through machine learning-powered data integration and its potential for forecasting in the Western U.S.
Hydrology and Earth System Sciences · 21 October 2025
AI-generated summary of the original article; changes were made. Check the original before relying on it.