Checking for missing content type metadata ...
This resource contains content types with missing metadata required to make it public or discoverable. Show missing content type metadata.
Click on the edit button ( ) below to edit this resource.
Checking for non-preferred file/folder path names (may take a long time depending on the number of files/folders) ...
This resource contains some files/folders that have non-preferred characters in their name. Show non-conforming files/folders.
This resource contains content types with files that need to be updated to match with metadata changes. Show content type files that need updating.
Data for the study Comprehensive evaluation on the robustness of deep learning based rainfall-runoff models under climate change scenarios
| Authors: |
|
|
|---|---|---|
| Owners: |
|
This resource does not have an owner who is an active HydroShare user. Contact CUAHSI (help@cuahsi.org) for information on this resource. |
| Type: | Resource | |
| Storage: | The size of this resource is 962.1 MB | |
| Created: | Sep 18, 2026 at 10:45 p.m. (UTC) | |
| Last updated: | Sep 23, 2026 at 6 p.m. (UTC) (Metadata update) | |
| Published date: | Sep 23, 2026 at 6 p.m. (UTC) | |
| DOI: | 10.4211/hs.9ff1c91ad7234bb8a8edbdde4e30bff3 | |
| Citation: | See how to cite this resource |
| Sharing Status: | Published |
|---|---|
| Views: | 139 |
| Downloads: | 9 |
| +1 Votes: | Be the first one to this. |
| Comments: | No comments (yet) |
Abstract
Deep learning rainfall-runoff models have shown the ability to reproduce the observed hydrograph exceeding the performance of physics-based models, but their reliability outside the training distribution is still an open question. Most robustness evaluation have focused on a single architecture (LSTM), and have used controlled forcing perturbations (temperature offset or precipitation scaling) to probe the non-stationarity condition. In this study, we evaluate the robustness of five deep learning architectures (LSTM, minLSTM, minGRU, MLP-Mixer, Transformer) against a re-calibrated physics-based benchmark (SAC-SMA + Snow-17) across 554 CAMELS catchments, driven by a six-Global Change Model (GCM), four-Shared Socioeconomic Pathway (SSP) NEX-GDDP-CMIP6 and three-future periods ensemble. The results show that the deep learning models outperform the physics-based benchmark in terms of accuracy, but their robustness is limited under non-stationary conditions. The LSTM and minGRU architectures show the best performance among the deep learning models, while the MLP-Mixer and minLSTM architectures show the worst performance. We found that the predictive skill does not imply robustness: the LSTM, the most accurate architecture, is the least robust. Transformer, the most robust on annual volumes, fails to reproduce the intra-annual redistribution of streamflow, diverging from the physics-based benchmark in both directions, compensating that difference at annual basis. Architecture, not GCMs, SSPs nor period, is the dominant factor of divergence (45 \% of the evaluated catchments). The results highlight the need for a more comprehensive evaluation of deep learning models under non-stationary conditions, and the importance of considering both accuracy and robustness in model selection.
Subject Keywords
Coverage
Spatial
Temporal
| Start Date: | |
|---|---|
| End Date: |
Content
README.md
Robustness of deep-learning rainfall–runoff models under non-stationary forcing
Data release: five deep-learning (DL) architectures benchmarked against SAC-SMA + Snow-17 across CAMELS catchments (CONUS), under a global-warming ensemble of six NEX-GDDP-CMIP6 GCMs, four SSP scenarios and three periods.
Contents
| Path | Files | Description |
|---|---|---|
KGE_models/ |
6 | Per-catchment test-period skill metrics, one CSV per model. |
Model_files/ |
20 + 671 | Trained DL checkpoints and configs; calibrated SAC-SMA parameter sets. |
Results/ |
10 | PGW analysis tables: Λ, PBIAS, robustness ratings, factor attribution, water balance. |
Results/signatures_raw/ |
648 | Raw streamflow signatures the Results/ tables derive from. |
Dimensions
| Dimension | Values |
|---|---|
| Models | sacsma (SAC-SMA), lstm (LSTM), minlstm (minLSTM), mingru (minGRU), mlpmixer (MLP-Mixer), transformer (Transformer) |
| GCMs | ACCESS-CM2, CanESM5, GFDL-ESM4, GISS-E2-1-G, MPI-ESM1-2-HR, MRI-ESM2-0 |
| Scenarios | historical, ssp126, ssp245, ssp370, ssp585 |
| Periods | p1 2015–2045, p2 2046–2070, p3 2071–2100; historical 1952–2014 |
| Catchments | 554 in all data tables; 671 in Model_files/SAC-SMA_parameters/ only |
| Gauge IDs | USGS, zero-padded 8-character strings (01013500) |
Units
| Quantity | Unit |
|---|---|
| Streamflow, precipitation | mm day⁻¹ |
| Temperature, temperature change | °C |
Λ, precip_factor, η²_H |
dimensionless |
| PBIAS | percentage points |
| Day of year | 1–366 |
Definitions
| Term | Definition |
|---|---|
Λ (lambda) |
φ(q_scenario) / φ(q_historical), same model and GCM. φ = period mean or DOY climatology. |
| PBIAS | 100 × (Λ_DL − Λ_SAC-SMA), same catchment / GCM / scenario / period. |
| Robustness | Very robust |PBIAS| < 5; Robust 5–10; Marginal robust 10–15; Non-robust ≥ 15. |
Provenance
| Item | Value |
|---|---|
| Forcings | CAMELS (Daymet / Maurer / NLDAS); NEX-GDDP-CMIP6, empirical quantile mapping |
| DL training | PyTorch, 200 epochs, MSE, 34 inputs (19 static attributes + 15 daily forcings) |
| DL splits | train 1999-10-01–2008-09-30, validation 1980-10-01–1989-09-30, test 1989-10-01–1999-09-30 |
| SAC-SMA calibration | SCE-UA via SPOTpy, independent per catchment |
Related Resources
| This resource updates and replaces a previous version | Tapia Araya, A., Andrew, B. (2026). Data for the study Robustness assessment of deep learning-based rainfall-runoff models, HydroShare, http://www.hydroshare.org/resource/65432d098e604986b19517378e7bf118 |
How to Cite
This resource is shared under the Creative Commons Attribution CC BY.
http://creativecommons.org/licenses/by/4.0/
Comments
There are currently no comments
New Comment