environment
Groundwater's Overlooked Role in Drought Prediction
A look at 2025-2026 USGS research showing how overlooked groundwater-surface water interactions drive prediction bias in national hydrologic models and reshape how scientists classify drought.
When most people picture a drought, they picture a sky that refuses to rain. But the U.S. Geological Survey's Water Resources Mission Area has spent the past two years quietly making the case that this picture is incomplete, and possibly misleading for anyone trying to forecast how droughts will unfold. A cluster of 2025-2026 USGS studies argues that the real fault line running through the nation's drought forecasts is underground: the exchange of water between aquifers and streams, a process most national-scale hydrologic models still treat as an afterthought. That gap, researchers say, is quietly distorting which watersheds get flagged as high-risk and which get missed entirely.
The invisible variable in national water models
Groundwater and surface water are not two separate systems. They physically overlap at what hydrologists call the groundwater/surface-water interface, where streams recharge aquifers and aquifers, in turn, sustain streamflow during dry stretches, sometimes for decades or centuries before that water resurfaces. This slow, subsurface transfer is exactly what keeps some rivers flowing through a rainless summer while others go dry within weeks. It is also, according to USGS scientists, one of the least faithfully represented processes in the National Water Model and other continental-scale forecasting systems, which were built primarily to route surface runoff rather than resolve subsurface storage and release.
That mismatch matters because drought forecasting increasingly depends on machine learning models trained to reproduce the behavior of these larger hydrologic systems. If the underlying physics of groundwater discharge is poorly captured, the errors do not distribute evenly. They cluster in specific types of watersheds, at specific times of year, producing systematic prediction bias rather than random noise.
A roadmap for finding where the models go wrong
A 2026 USGS study led by Ryan van der Heijden and colleagues set out to map exactly where that bias hides. The team analyzed daily streamflow from 797 reference-quality streamgages across the contiguous United States, using a calibrated digital filter to isolate the baseflow component, the portion of a river's flow sustained by groundwater discharge, from each gage's record. Hierarchical clustering of those monthly baseflow index signatures identified seven distinct regimes capturing how groundwater contributions vary regionally and seasonally.
The researchers then linked those regimes directly to the performance of the National Water Model, describing their goal in the study itself: "Understanding how groundwater-surface water interactions shape streamflow variability is critical for diagnosing low flow behavior and prediction bias in continental scale hydrologic models." Their conclusion was that watersheds dominated by strong, stable groundwater contributions behave very differently, and are modeled with very different degrees of accuracy, than flashier, runoff-dominated basins. By tying model error back to physically interpretable baseflow regimes rather than treating it as an unexplained residual, the team argues it becomes possible to generate testable hypotheses about exactly which physical processes a national model is failing to represent, rather than simply flagging that it has failed.
From drought driver to drought consequence
A related strand of USGS research, led by Corinne Casey Bowers with Jared David Smith, Hedeff Essaid, and John C. Hammond, tackled a companion problem: even when a model captures streamflow reasonably well, connecting the atmospheric and hydrologic drivers of a drought to its eventual severity and downstream consequences is its own analytical challenge. Focused on the Delaware River Basin, the team built what they describe as a two-step data-driven approach. Using their words, "we use machine learning feature selection techniques to transform hydrologic model outputs into calculated indices and determine which features are most predictive of streamflow drought duration and severity," before using those selected features to cluster individual drought events between 1985 and 2015.
That clustering step is where drought typology comes in. Rather than treating drought as a single phenomenon that simply varies in intensity, the Delaware Basin analysis sorted events into distinct types based on their driving mechanisms, seasonal timing, and links to water-management metrics such as reservoir operations. A drought triggered by a sharp precipitation deficit in a groundwater-poor headwater catchment behaves, both physically and statistically, nothing like a slow-building deficit in a basin buffered by large aquifer storage. Lumping them together in a single national model, the research implies, obscures the very distinctions that would let water managers anticipate which kind of event they are facing and how it is likely to evolve.
Where deep learning models actually break down
A third piece of the puzzle, published in Water Resources Research by Ali Dadkhah and coauthors, tested this idea directly against a USGS long short-term memory (LSTM) deep learning model used to forecast streamflow drought across the Colorado River Basin. Using data from 384 streamgages, the researchers clustered catchments two ways: first by static physical and climatological attributes, and second by their streamflow drought signatures over time, then examined how each clustering scheme related to the model's error patterns.
The results were specific enough to be actionable. Elevation, degree of streamflow regulation, baseflow contribution, catchment aridity, and drainage area emerged as the attributes most associated with model performance, and catchments with pronounced seasonal peak runoff between January and June generally saw better forecasts than those without it. Perhaps the most striking finding was procedural: a simple Random Forest classifier, trained only on physical and climatological catchment attributes, could predict whether the LSTM model would perform well or poorly with an F1 score of 0.72. The authors singled out one variable, low degree of flow regulation, as a particularly reliable indicator of stronger model performance, an implicit acknowledgment that dams, diversions, and groundwater pumping introduce exactly the kind of human-managed complexity that pattern-matching machine learning struggles to learn from limited historical records.
A structural gap the agency is trying to close
These three studies converge on a shared diagnosis that the USGS itself has now formalized as an open research priority. A 2026 Mendenhall Research Fellowship opportunity notice states plainly that "a critical gap in national scale hydrologic modeling efforts is the inability to integrate the surface water stream network with subsurface hydrogeology and groundwater transport," and calls for work coupling land-surface hydrology models with groundwater-transport codes such as MODFLOW, the agency's long-standing modular groundwater simulator, or ParFlow, at basin or national scale. That coupling is technically difficult: surface hydrology models operate on fast timescales measured in hours and days, while aquifer systems evolve over months, years, and sometimes centuries, and reconciling the two within a single computational framework remains an active engineering problem, not a solved one.
In the meantime, the agency has been building narrower, more targeted tools that at least surface the symptom even if the underlying models cannot yet fully resolve the cause. River DroughtCast, one of several web applications now maintained by the Water Resources Mission Area, delivers weekly streamflow drought forecasts by combining real-time gage data with statistical and modeled projections, giving water managers a near-term view even in basins where the deeper groundwater physics is still poorly constrained. The agency has also leaned on its century-long groundwater monitoring network, the Climate Response Network, whose oldest well in Woodgate, New York, passed its hundredth consecutive year of continuous data collection in 2026, to provide the kind of long-baseline observational record that any future coupled model will need for calibration and validation.
Why this matters beyond hydrology
The practical stakes of this research extend well past academic model-fitting. Drought early-warning systems, agricultural water allocation decisions, reservoir operating rules, and wildfire-risk assessments all lean, directly or indirectly, on the outputs of the same continental-scale hydrologic models this research is now probing for bias. If a model systematically underestimates how long a groundwater-buffered basin can sustain flow, planners may over-restrict water use in places with more resilience than the forecast suggests. If it overestimates that buffering in a flashier, runoff-dominated catchment, the opposite failure occurs: a drought arrives faster and harder than anyone was told to expect.
What makes this line of USGS research distinctive is its refusal to treat model error as a black box to be minimized statistically. By anchoring machine learning diagnostics to physically meaningful watershed attributes, baseflow regimes, flow regulation, aquifer storage, catchment aridity, the researchers are effectively building a map of where and why national drought forecasts can be trusted, and where they cannot yet be. That map, still being drawn one basin and one streamgage record at a time, may end up mattering as much for climate resilience planning as any single improvement to the models themselves.