Case study · Pakistan · Live system
Glacial lake outburst flood early warning
Nine glaciers across Gilgit-Baltistan and Chitral, refreshed weekly from open satellite archives. A precursor that warns weeks ahead — and a clear statement of the days when it cannot see anything at all.
documented bursts warned, across two independent lakes
maximum validated lead time before a burst
structures inside the routed corridors of the seven GB glaciers
routed arrival time to the first settlement at the fastest site
The hazard, in the record rather than the abstract
Pakistan's documented outburst-flood record holds 151 events, and 140 of them — 93% — are in Gilgit-Baltistan. The Hunza basin alone accounts for 79, more than half the national total. Half of all events are ice-dammed, and that share has climbed across successive decades, from 29% through 35% to 40%.
That concentration is what makes the problem tractable. A small number of specific valleys carry most of the risk, and in those valleys the mechanism is usually the same: a lake fills behind an ice dam over weeks, then releases in hours.
A lake that takes weeks to fill can be watched. That is the whole basis of the system — and the reason the model that came before it had to be thrown away.
The model that failed, and why we published it
The project's original predictor was a recurrent neural network trained on weather forcing. Displayed in-sample, it looked adequate — a ROC-AUC around 0.67. But it was trained on the full record including event years, and only masked a year for display, which is not a test of anything.
We built a leave-one-year-out harness: retrain on every year except the held-out one, refit the calibration and alarm threshold on those years too, then predict the year the model has never seen. Under that test the network scored an out-of-sample ROC-AUC of 0.120 — well below chance — against a best single-feature baseline of 0.584. It warned 0 of 15 documented events.
A pooled regional model fared no better: trained on one glacier and tested on another it scored 0.339, because the mechanisms differ — one site is an ice-dam system, the other a surge system — and meteorological forcing alone does not transfer between them.
So we published the failure and rebuilt around the physical observable that actually precedes these events: the lake itself.
What replaced it
The precursor combines lake area against that site's own historical capacity with the rate at which it is filling, forward-filled between sparse observations with a decay so a stale high reading cannot masquerade as a current one. Validated against documented bursts at two independent lakes:
| Site | Mechanism | ROC-AUC | False-alarm rate | Events warned |
|---|---|---|---|---|
| Shisper | Ice-dammed lake, Hunza | 0.742 | 0.042 | 3 of 7 |
| Khurdopin / Virjerab | Surge-dammed lake, Shimshal | 0.866 | 0.183 | 4 of 4 |
| Weather neural model (replaced) | Meteorological forcing | 0.120 | — | 0 of 15 |
Seven of eleven bursts, at two lakes with different damming mechanisms, with lead times reaching 21 days and a false-alarm rate at Shisper of 4.2%. Not a solved problem — four bursts were missed, and the reasons are documented site by site — but a measurable, reproducible improvement over the alternative, which warned nothing.
Terrain decides how much that warning is worth
Lead time is only useful against travel time. The terrain engine routes each glacier's drainage to the first inhabited structure downstream and measures the valley width at 50 m above the channel bed — the height a flood actually occupies, rather than the summer stream.
| Glacier | Routed arrival to first settlement | Corridor width at +50 m |
|---|---|---|
| Minapin | 28.5 min | 375 m |
| Hispar | 28.8 min | 206 m |
| Passu | 31.6 min | 750 m |
| Shisper | 35.6 min | 319 m |
| Darkut | 39.9 min | 694 m |
| Biafo | 57.7 min | 356 m |
| Baltoro | 108.5 min | 750 m |
For context, the corridor that carried Nepal's lethal August 2026 flood was 180–480 m wide, measured the same way, and it needed 60 km of that confinement to reach people. Five of these seven valleys sit inside or below that band — and here the first settlements are 7 to 17 km from the ice. Confinement that took 60 km to become lethal in Nepal has a far shorter distance to work with in the Karakoram.
An alert that knows what season it is
In July 2026 the alert engine fired a critical warning for a newly detected impounded lake, then repeated a "stay evacuated" reminder weekly for a month. The lake was real. It was also entirely ordinary: that lake has filled to 14–19 ha every summer since 2016. The engine had been reading the routine spring refill as a lake forming for the first time, because through winter the snow and ice cover makes the water read as zero area.
Four fixes went in, and the design principle behind them is the important part:
- A seasonal baseline comparing each reading to all prior-year observations within three weeks of the same day of year — and returning "unavailable" rather than "normal" when the record is too thin, so a short history can never suppress an alert.
- Hysteresis on the rating thresholds, because a precursor sitting on a cut point flip-flops between states from noise alone.
- Radar as a floor, never a ceiling — a Sentinel-1 confirmation can only ever hold a lake's rating up, never reduce it.
- Observation age beside every rating. A low reading taken forty days ago is an absence of information, not an all-clear, and the two must not look alike on a screen.
Seasonality may gate novelty, but it must never gate danger. Back-testing a seasonal filter against seven documented bursts at one glacier showed that only one of the seven exceeded its own seasonal ninetieth percentile in the run-up. Applying the filter to escalations would have suppressed six real warnings.
So it is applied only to "this lake is new" and "this lake is growing". Escalations always fire; they are merely annotated with the seasonal context. The current elevated site reads 16.07 ha against a 15.65 ha median for the same week across nine prior years — a ratio of 1.03. Loaded as usual for late melt season, with 8,537 structures downstream. That caveat is given the same weight on the page as the alarm.
Where the system is blind, stated plainly
Optical satellites cannot see through cloud, and an early-warning system that does not publish its own gaps is not an early-warning system.
- 85% of regional events fall between May and September. In season the median gap between usable observations is 3 to 10 days, which is comfortable against a 21-day lead.
- The tail is the problem. Ninetieth-percentile gaps run 20 to 35 days, and the worst in-season gaps at two sites reach 115 days. Fifteen of ninety in-season intervals there exceed the entire warning lead. Success has to be defined as "worst in-season gap under 21 days", never as an average.
- Two of the fastest-onset valleys have no lake to watch. The sites with the narrowest corridor and the shortest arrival time have no mapped impounding lake, so they have no leading indicator — only velocity and weather. That is stated as a coverage hole, not papered over.
- One monitored site has no detectable lake at all, and is reported as an honest not-applicable rather than assigned a fabricated reading.
Where the method does not transfer
Extending the work into Chitral produced the most consequential negative result of the programme. Routing 334 glacier traces at native 30 m resolution — a coarser grid silently destroys the analysis, because the side valleys are narrower than three 90 m cells — showed that the dominant mechanism there is the englacial water pocket, not a surface lake. Four of eleven documented events outside Gilgit-Baltistan are water-pocket releases, including the two that struck the villages in question.
A water pocket sits inside or beneath the ice. There is no surface lake to watch, so the precursor validated in Gilgit-Baltistan does not transfer to Chitral. It also finally explained why one monitored site had been reading zero lake area for years: correct boundary, wrong mechanism. We said so in the briefing rather than extend the product over a region it cannot serve.
Method: Copernicus Sentinel-2 L2A lake-area series 2016–present with per-scene cloud masking and processing-baseline reflectance correction; Sentinel-1 for cloud-independent confirmation; ITS_LIVE surface velocity; Copernicus GLO-30 for D8 routing, corridor width and arrival time; open building footprints for exposure; HMAGLOFDB for documented events; RGI 7.0 for glacier outlines. Validation by leave-one-year-out retraining and event back-testing. Exposure is reported as structures, never converted to population without a supplied household figure. Contains modified Copernicus Sentinel data 2016–2026.
Related
The flood that set the benchmark.
Our reconstruction of the August 2026 Nepal–Tibet event — including the claim we withdrew.