Skip to content Where Legends Are Made
Cooperative Institute for Research to Operations in Hydrology

CIROH Training and Developers Conference 2026 Abstract

Authors: Mirce Morales-Velazquez, Beverley C. Wemple, Kristen L. Underwood, Donna M. Rizzo, Shaurya Swami, Garnet Williams, Patrick Clemins, Noah B. Beckage, Andrew Schroth – University of Vermont; James Shanley – United States Geological Survey; Zachary Suriano – Western Kentucky University  

Title:  Drivers of poor National Water Model streamflow estimation performance during flood events in mountainous catchments across the Northeastern US 

Presentation Type: Poster Presentation 

Abstract:  The National Water Model (NWM) is an important tool for delivering operational streamflow forecasts to communities across the United States. However, performance varies substantially in space and time. Gaining insight into the primary mechanisms driving model errors and their sources, whether stemming from antecedent conditions, atmospheric forcings, or model physical representations, is paramount for enhancing NWM’s performance and informing the continued development of the NextGen framework. In this study, we conducted an analysis using machine learning to identify the key drivers of poor NWM streamflow performance across 128 mountainous basins in the northeastern United States. We analyzed events from 2019 to 2024 with peak streamflow corresponding to an Annual Exceedance Probability (AEP) less than or equal to 50% (i.e. floods of 2-yr recurrence intervals or greater). For these events, we extracted simulated streamflow data from the NWM Analysis and Assimilation Extended configuration without data assimilation (Extended AnA no DA) and the Medium-Range forecast products. We then computed the percentage error in peak streamflow (

) by comparing the simulated and observed flows at the USGS streamgages, along with a suite of features intended to explain model performance. These explanatory features included atmospheric forcing from the NWM (e.g., accumulated precipitation, mean rain rate, temperature, specific humidity) at multiple time lags prior to observed peak flow, snowpack conditions sourced from SNODAS, storm type, and basin physiographic characteristics from NHDPlus v2 dataset. Feature selection was performed using a tandem evolutionary algorithm (TEVA), and the selected features were subsequently used to train a random forest classifier to differentiate peak flow underestimation or overestimation for poorly performing events (defined as 

). The results indicate that poor performance in the Extended AnA no DA configuration is primarily driven by deficiencies in the representation of land-atmosphere mass and energy exchanges, and to a lesser extent, by precipitation forcing errors. In contrast, poor performance in the Medium-Range forecasts is dominated by errors in the precipitation forcings. These findings highlight the need for bias correction of atmospheric input variables beyond precipitation in the Extended AnA no DA configuration, particularly for processes influencing snowmelt. Furthermore, improving precipitation forecasts remains critical for enhancing medium-range streamflow forecasts. Enhancements to the Extended AnA no DA would also improve initial conditions for Medium-Range forecasts, fostering a comprehensive improvement in NWM performance. This work highlights key components of both the NWM and the NextGen framework that warrant prioritization for performance optimization.