When a storm dumps large amounts of water on a catchment, how can we predict how much will reach the river, how quickly and what discharge it will reach? That is the fundamental question of hydrological modelling, the discipline that turns rainfall data into forecasts of discharge and water level. Behind every flood warning are mathematical models simulating the complex journey of water from the moment it falls as rain until it flows down a channel. This article explains the key concepts in accessible terms.

The basic problem: turning rainfall into runoff

Not all the rain falling on a catchment ends up in the river. Some infiltrates the soil, some is intercepted by vegetation, some fills small depressions in the ground, and some evaporates directly. Only the fraction not absorbed by any of these processes becomes surface runoff and flows downhill towards streams and rivers.

The relationship between the rain that falls and the discharge that results at a point on the river is known as the rainfall-runoff transformation, and it is the core of hydrological modelling. It depends on many factors: soil type, vegetation cover, slope, antecedent soil moisture, rainfall intensity and duration, and how urbanised the catchment is.

Lumped versus distributed models

Two broad philosophies exist in hydrological modelling, differing in how they represent spatial variability.

Lumped models treat the catchment as a homogeneous unit characterised by average parameters. Rainfall is assumed uniform across the whole catchment and the response is computed globally. They are simpler, need less data and run fast, but they lose the information about exactly where it rains hardest or where the soil is more permeable.

Distributed models divide the catchment into cells or sub-basins, each with its own soil, vegetation, slope and rainfall. They compute runoff cell by cell and route the flow through the drainage network. They are more realistic but demand far more information — digital elevation models, soil maps, gridded rainfall — and far more computing power.

Semi-distributed models: In practice many operational models take a middle path: they divide the catchment into sub-basins, each modelled in lumped form, and then route the resulting discharges through the river network. That offers a good balance between realism and practicality. HEC-HMS, widely used in Spain, works this way.

The SCS-CN method: how much rain becomes runoff

One of the most widely used methods for estimating effective rainfall — the part that generates runoff — is the Soil Conservation Service Curve Number method, developed by the US Department of Agriculture in the 1950s and still standard in professional practice.

The method rests on a dimensionless parameter, the Curve Number, ranging from 0 to 100, that reflects the catchment’s runoff potential. A high Curve Number close to 100 indicates a very impermeable catchment where almost all rain becomes runoff; a low one indicates high infiltration capacity.

It depends on three main factors:

  • Hydrological soil group. Soils are classified into four groups, A to D, by infiltration rate. Group A, deep sands, infiltrates a great deal; group D, impermeable clays, very little.
  • Land use and vegetation cover. Dense forest generates far less runoff than an asphalt car park on the same soil type.
  • Antecedent moisture condition. Soil already saturated by earlier rain generates much more runoff than dry soil. The method allows for three prior moisture conditions: dry, normal and wet.

In the method’s fundamental equation, accumulated runoff depends on accumulated rainfall minus an initial abstraction, divided by that same quantity plus the maximum potential retention, which is derived directly from the Curve Number.

A worked example: A forested catchment on sandy soil might have a Curve Number of 55, meaning 100 mm of rain would generate about 18 mm of runoff. The same rain over a sealed urban area with a Curve Number of 98 would generate 94 mm. Urbanisation multiplies runoff by a factor of five or more.

The Clark unit hydrograph: when the water arrives

The SCS-CN method tells us how much rain becomes runoff, but not when that runoff reaches the control point on the river. For that we need a unit hydrograph, a model describing how the catchment’s response to a pulse of rain is distributed over time.

The Clark unit hydrograph is one of the most used in Spain. It rests on the time-area concept: it divides the catchment into bands of equal travel time to the outlet, called isochrones, and calculates what fraction of the catchment area contributes to discharge in each time interval. It then applies a linear reservoir to simulate the storage and attenuation the catchment exerts on the flow.

The Clark model has two parameters:

  • Time of concentration: how long water takes to travel from the furthest point of the catchment to the outlet, depending on the length, slope and roughness of the path.
  • Storage coefficient: the attenuation the terrain exerts on the flood wave. Flat catchments with wide floodplains have a high value and attenuate strongly; steep, incised catchments have a low value and respond fast and sharply.

Routing the flood: the Muskingum method

Once runoff has been generated in each sub-basin, we need to simulate how the flood wave travels along the river reaches connecting them. This process is called flood routing.

The Muskingum method, originally developed for the Muskingum river in Ohio, is the most widespread routing method in hydrological models. It simulates two fundamental effects the channel exerts on a flood wave:

  • Translation. The wave moves downstream at a certain speed, so there is a delay between the flood passing one point and reaching the next.
  • Attenuation. The channel and its banks temporarily store water as the flood passes, reducing the peak and lengthening the wave. Natural floodplains are extraordinarily effective at this.

The method uses two parameters: the travel time of the reach, in hours, and a weighting parameter between 0 and 0.5 that controls how much attenuation occurs. Its Muskingum-Cunge variant allows both to be estimated from the geometry and roughness of the channel, without empirical calibration.

Calibration: fitting the model to reality

Every hydrological model needs calibration: its parameters must be adjusted by comparing the model output, the simulated discharge, against real gauged data. Without calibration, a model can produce results far removed from reality.

The calibration process involves:

  • Selecting historical events for which both rainfall and observed discharge data exist.
  • Running the model with the rainfall data as input.
  • Comparing simulated hydrographs against observed ones.
  • Iteratively adjusting the parameters until the simulation reproduces the observations reasonably well.
  • Validating the model against events not used in calibration, to confirm the parameters are robust and not overfitted.
Uncertainty is unavoidable: Even a well-calibrated model carries uncertainty. Rainfall data contains errors, since a gauge network cannot capture all the spatial variability; parameters are simplifications of complex processes; and catchment conditions change over time through urbanisation, fires and land-use change. Recognising and quantifying that uncertainty matters as much as the model itself.

Ensemble forecasts and handling uncertainty

To cope with uncertainty, modern practice uses ensemble forecasting. Instead of running the model with a single weather forecast, it is run with multiple rainfall scenarios drawn from probabilistic weather forecasts, such as the ECMWF ensemble with its 51 members. That produces a spread of possible future hydrographs, each with an associated probability.

Rather than saying «the peak will be 500 m³/s», an ensemble forecast says there is an 80 % probability of exceeding 300 m³/s, a 50 % probability of exceeding 500 m³/s and a 10 % probability of exceeding 800 m³/s. That probabilistic information is far more useful for civil protection decisions than a single deterministic number.

Real-time updating

Where observed discharge data is available in real time from SAIH gauging stations, hydrological models can be updated continuously. If observed discharge diverges from the simulation, the model corrects its parameters or its internal states — soil moisture levels, stored volumes — to match what is actually happening. This technique, known as data assimilation, significantly improves forecast accuracy over the following hours.

Models used in Spain

Several hydrological models are in operational use in Spain:

  • SIMPA. A distributed model developed by CEDEX to assess water resources in Spain. It operates at monthly scale and is used mainly for long-term water planning, not real-time flood forecasting.
  • HEC-HMS. Developed by the US Army Corps of Engineers, probably the most widely used flood simulation model in the world. It is semi-distributed, incorporates all the methods described here and is free to use. Many Spanish flood studies are carried out with it.
  • MIKE. Developed by the Danish Hydraulic Institute, a one-dimensional hydrodynamic model that solves the full Saint-Venant equations for open channel flow. It is used when channel hydraulics must be modelled in detail: water levels, velocities, overtopping, the effect of obstacles. Several Spanish river basin authorities run operational MIKE models.
WhatAWeather and simplified modelling: WhatAWeather integrates discharge forecasts from the Open-Meteo Flood API, which uses the GloFAS model — the Global Flood Awareness System of the European Centre for Medium-Range Weather Forecasts. GloFAS provides probabilistic discharge forecasts for the world’s major rivers, which WhatAWeather presents as flood risk warnings several days ahead.

The future: machine learning in flood forecasting

In recent years machine learning has arrived in hydrology in force. Models based on recurrent neural networks, particularly LSTM networks, have shown a remarkable ability to predict discharge from rainfall, temperature, soil moisture and previous discharge.

Recent studies show that LSTM models can match or beat traditional hydrological models calibrated on individual catchments, with the added advantage of learning generalisable patterns from many catchments at once. Google’s flood forecasting project, which uses these techniques, has produced promising results at global scale.

Machine learning models have limitations too:

  • Black box behaviour. It is hard to interpret why the model makes a particular prediction, which complicates error detection and undermines the confidence of the people who must act on it.
  • Data dependence. They need long, high-quality historical series for training. In poorly gauged catchments they can perform worse than physically based models.
  • Extrapolation. They can fail on extreme events outside the range of the training data — precisely the most dangerous events.

Coupling weather and hydrology

The most promising frontier in flood forecasting is coupling meteorological and hydrological models. Instead of using a rainfall forecast as a static input, the most advanced systems run the whole chain: global atmospheric model, then a high-resolution regional weather model, then a catchment hydrological model, then a hydrodynamic channel model, and finally flood maps.

That full chain makes it possible to move from a forecast of «heavy rain» to a forecast of «this street will be under 80 cm of water in six hours», which is the information that actually supports civil protection decisions.

Hydrological modelling, with its classical methods and its new computational tools, remains the indispensable bridge between observing the weather and protecting lives and property from floods. Its continued improvement, fed by better data, greater computing power and new artificial intelligence techniques, is one of the pillars of our adaptation to a changing climate.