Checking for missing content type metadata ...

Checking for non-preferred file/folder path names (may take a long time depending on the number of files/folders) ...

Data for the study Comprehensive evaluation on the robustness of deep learning based rainfall-runoff models under climate change scenarios


An older version of this resource http://www.hydroshare.org/resource/65432d098e604986b19517378e7bf118 is available.
Authors:
Owners: This resource does not have an owner who is an active HydroShare user. Contact CUAHSI (help@cuahsi.org) for information on this resource.
Type: Resource
Storage: The size of this resource is 962.1 MB
Created: Sep 18, 2026 at 10:45 p.m. (UTC)
Last updated: Sep 23, 2026 at 6 p.m. (UTC) (Metadata update)
Published date: Sep 23, 2026 at 6 p.m. (UTC)
DOI: 10.4211/hs.9ff1c91ad7234bb8a8edbdde4e30bff3
Citation: See how to cite this resource
Sharing Status: Published
Views: 139
Downloads: 9
+1 Votes: Be the first one to 
 this.
Comments: No comments (yet)

Abstract

Deep learning rainfall-runoff models have shown the ability to reproduce the observed hydrograph exceeding the performance of physics-based models, but their reliability outside the training distribution is still an open question. Most robustness evaluation have focused on a single architecture (LSTM), and have used controlled forcing perturbations (temperature offset or precipitation scaling) to probe the non-stationarity condition. In this study, we evaluate the robustness of five deep learning architectures (LSTM, minLSTM, minGRU, MLP-Mixer, Transformer) against a re-calibrated physics-based benchmark (SAC-SMA + Snow-17) across 554 CAMELS catchments, driven by a six-Global Change Model (GCM), four-Shared Socioeconomic Pathway (SSP) NEX-GDDP-CMIP6 and three-future periods ensemble. The results show that the deep learning models outperform the physics-based benchmark in terms of accuracy, but their robustness is limited under non-stationary conditions. The LSTM and minGRU architectures show the best performance among the deep learning models, while the MLP-Mixer and minLSTM architectures show the worst performance. We found that the predictive skill does not imply robustness: the LSTM, the most accurate architecture, is the least robust. Transformer, the most robust on annual volumes, fails to reproduce the intra-annual redistribution of streamflow, diverging from the physics-based benchmark in both directions, compensating that difference at annual basis. Architecture, not GCMs, SSPs nor period, is the dominant factor of divergence (45 \% of the evaluated catchments). The results highlight the need for a more comprehensive evaluation of deep learning models under non-stationary conditions, and the importance of considering both accuracy and robustness in model selection.

Subject Keywords

Coverage

Spatial

Coordinate System/Geographic Projection:
WGS 84 EPSG:4326
Coordinate Units:
Decimal degrees
North Latitude
49.3500°
East Longitude
-66.9500°
South Latitude
24.7400°
West Longitude
-124.7800°

Temporal

Start Date:
End Date:

Content

README.md

Robustness of deep-learning rainfall–runoff models under non-stationary forcing

Data release: five deep-learning (DL) architectures benchmarked against SAC-SMA + Snow-17 across CAMELS catchments (CONUS), under a global-warming ensemble of six NEX-GDDP-CMIP6 GCMs, four SSP scenarios and three periods.

Contents

Path Files Description
KGE_models/ 6 Per-catchment test-period skill metrics, one CSV per model.
Model_files/ 20 + 671 Trained DL checkpoints and configs; calibrated SAC-SMA parameter sets.
Results/ 10 PGW analysis tables: Λ, PBIAS, robustness ratings, factor attribution, water balance.
Results/signatures_raw/ 648 Raw streamflow signatures the Results/ tables derive from.

Dimensions

Dimension Values
Models sacsma (SAC-SMA), lstm (LSTM), minlstm (minLSTM), mingru (minGRU), mlpmixer (MLP-Mixer), transformer (Transformer)
GCMs ACCESS-CM2, CanESM5, GFDL-ESM4, GISS-E2-1-G, MPI-ESM1-2-HR, MRI-ESM2-0
Scenarios historical, ssp126, ssp245, ssp370, ssp585
Periods p1 2015–2045, p2 2046–2070, p3 2071–2100; historical 1952–2014
Catchments 554 in all data tables; 671 in Model_files/SAC-SMA_parameters/ only
Gauge IDs USGS, zero-padded 8-character strings (01013500)

Units

Quantity Unit
Streamflow, precipitation mm day⁻¹
Temperature, temperature change °C
Λ, precip_factor, η²_H dimensionless
PBIAS percentage points
Day of year 1–366

Definitions

Term Definition
Λ (lambda) φ(q_scenario) / φ(q_historical), same model and GCM. φ = period mean or DOY climatology.
PBIAS 100 × (Λ_DL − Λ_SAC-SMA), same catchment / GCM / scenario / period.
Robustness Very robust |PBIAS| < 5; Robust 5–10; Marginal robust 10–15; Non-robust ≥ 15.

Provenance

Item Value
Forcings CAMELS (Daymet / Maurer / NLDAS); NEX-GDDP-CMIP6, empirical quantile mapping
DL training PyTorch, 200 epochs, MSE, 34 inputs (19 static attributes + 15 daily forcings)
DL splits train 1999-10-01–2008-09-30, validation 1980-10-01–1989-09-30, test 1989-10-01–1999-09-30
SAC-SMA calibration SCE-UA via SPOTpy, independent per catchment

Related Resources

This resource updates and replaces a previous version Tapia Araya, A., Andrew, B. (2026). Data for the study Robustness assessment of deep learning-based rainfall-runoff models, HydroShare, http://www.hydroshare.org/resource/65432d098e604986b19517378e7bf118

How to Cite

Tapia Araya, A., Andrew, B. (2026). Data for the study Comprehensive evaluation on the robustness of deep learning based rainfall-runoff models under climate change scenarios, HydroShare, https://doi.org/10.4211/hs.9ff1c91ad7234bb8a8edbdde4e30bff3

This resource is shared under the Creative Commons Attribution CC BY.

http://creativecommons.org/licenses/by/4.0/
CC-BY

Comments

There are currently no comments

New Comment

required