Forecast verification · Australian Alps
A warm rain-on-snow event. The five verified resort stations recorded 56–86 mm of precipitation over 48 hours, of which 7.6–17.0 mm of water equivalent fell as snow — 16% pooled — while the resorts themselves reported 4–23 cm of new snow across the two mornings. Every one of the eight models called a snowier phase mix than fell. Decomposing that error shows it was almost entirely a phase error, not a rainfall one: the models broadly knew how much water was coming, and were wrong about what it would fall as.
Why the snow didn't come
Wet-bulb zero height inferred half-hourly from the Victorian station network (five stations, 1295–1849 m; mean R² 0.917 across the 81 precipitating steps). Rust marks the level above the highest lift-served terrain in the country; blue below it. The bars share the colouring: Victorian mean precipitation per half hour. The dashed lines are what the models forecast the level would do: 8 models, each its last initialisation before the window opened (8–14 h ahead of it), averaged across the same four Victorian sites as the observed curve. Three are drawn at first; the switches add the rest, and the figure beside each is that model's hourly level error from the table below.
The moisture and the cold air arrived in the wrong order. The level climbed through the 9th and peaked at 2436 m at 02:30 on the 10th, while the rain was at its heaviest — Thredbo Top Station was recording rain at 1957 m in gusts to 130 km/h. Every one of the 8 runs drawn here put that peak lower than it was: at the hour of the observed 2436 m they sat between 1979 m (ICON) and 2362 m (UKMO), 74 to 457 m low — and every metre of that gap is terrain the model had catching snow while it was catching rain. It then dropped roughly a kilometre in six hours, by which time little was left to fall. Precipitation-weighted across the event the level sat at 2068 m, above every summit in the country. New South Wales has three stations, but Thredbo Village reports at 09:00 and 15:00 only, leaving two with half-hourly thermometry — below the four the regression requires — so the curve is Victorian and inferred, not measured.
Observations
Event totals by station, midnight to midnight AEST. Gauge snow is water equivalent under the +1.0 °C wet-bulb phase call — no station measures phase or snowfall depth directly, so no gauge-derived centimetre figure is offered. The resort column is the resorts' own reported new snow, shown for context (see note).
| Station | Elevation | Precipitation, mm | As snow, mm water | Snow fraction | Resort reported, cm | Peak gust, km/h |
|---|---|---|---|---|---|---|
| Mount Buller VIC | 1707 m | 56.0 | 12.6 | 22% | 9 (1+8) | 100 |
| Perisher Valley NSW | 1738 m | 63.8 | 7.6 | 12% | 15 (5+10) | 67 |
| Falls Creek VIC | 1765 m | 70.2 | 8.4 | 12% | 12 (2+10) | 96 |
| Mount Hotham VIC | 1849 m | 77.8 | 17.0 | 22% | 23 (10+13) | 120 |
| Thredbo Top Station NSW | 1957 m | 85.8 | 11.2 | 13% | 6 (1+5) | 130 |
| Mount Hotham Airport VICbelow the lifts | 1295 m | 38.4 | 1.8 | 5% | — | 57 |
| Thredbo Village NSWtotal only | 1380 m | 18.2 | — | — | — | — |
| Mount Baw Baw VICtotal only | 1561 m | 39.0 | — | — | 4 (4+0) | 67 |
Pooled across the five verified stations, 56.8 mm of 353.6 mm fell as snow (16.1%); the unweighted mean of the five is 16.3%. Mount Hotham Airport sits below lift-served terrain and is measured but not verified against — including it gives a regional 14.9% across 6 stations. Mount Buller and Falls Creek each have a gap on the morning of the 10th reconstructed from BoM's published products (4.4 and 9.0 mm; see method); no other verified gauge takes more than 5% of its event total in any one half-hour.
Two gauges give a total but not a split. Thredbo Village and Mount Baw Baw each carry a stretch of precipitation longer than the scoring will phase, so neither can say when its water fell. Both keep their totals and are scored on rainfall alongside the rest; both are left out of the snow fraction, and their thermometers stay in the snow-level regression. Station by station →
Resort-reported new snow is context, not verification data. Figures are the resorts' own morning reports summed over the 10 and 11 August report days (the two ~7 am cycles covering this window) from the Alpine Weather Dashboard's official-report store. They measure settled centimetres on a stake at a site the resort chooses, on a window offset several hours from this one — a different quantity, instrument and clock from the gauge columns, and no water-to-depth conversion between them is attempted. Read together they agree on the shape of the event: modest accumulations from a large storm, most at Hotham, least at Baw Baw and Thredbo.
Forecasts
Mean forecast event total across the verified resorts against the observed mean of 65 mm, over every initialisation spanning the window. The snow column is indicative: each model's mean total multiplied by its own snow–rain bracket, against the observed 11 mm water equivalent — the rust tick, and the mean snow water actually recorded at the five gauges whose totals can be split, 16.1% of what they caught. This column is arithmetic on the two beside it, not a per-run score. The scored comparisons follow below.
| Model | Inits · runs | Mean total, mm | Bias, dry ◂ ▸ wet | MAE, mm | Implied snow vs observed 11 mm water |
|---|---|---|---|---|---|
| ECMWF IFS | 11 · 66 | 71 | +5.7 | 12.9 20% | 22–45 |
| GFS | 12 · 72 | 75 | +9.8 | 15.8 26% | 30–62 |
| ICON | 11 · 66 | 69 | +3.7 | 17.4 28% | 23–34 |
| ECMWF AIFS | 11 · 66 | 46 | −19.4 | 20.7 30% | 6–22 |
| GDPS 2–0 d only | 3 · 18 | 49 | −16.2 | 22.5 33% | 18–28 |
| ARPEGE 1–0 d only | 5 · 30 | 52 | −13.0 | 24.1 36% | 10–23 |
| UKMO | 10 · 60 | 74 | +8.9 | 24.1 38% | 20–42 |
| JMA | 11 · 66 | 24 | −41.0 | 45.4 69% | 6–11 |
Ordered by rainfall MAE. On water the field splits both ways: four models ran wet (ECMWF IFS, GFS, ICON, UKMO) and four ran dry (ECMWF AIFS, ARPEGE, GDPS, JMA), with JMA the outlier at 42 mm dry. On the indicative snow column almost every bracket sits above what fell. Where one straddles the observed value it can do so for opposite reasons: ECMWF AIFS through the best-placed phase bracket in the field, JMA through a large dry rainfall error offsetting a warm-biased phase call — a cancellation rather than a skill.
Scored comparisons
Phase, rainfall and snow-level are scored separately because they rank the field differently. Phase: each run's snow–rain bracket (the model derivation marks some hours uncertain, so a run gives a range, counted as rain at the low end and snow at the high end) against the observed 16.1% — the rust tick on each bar; green brackets contain it. Rainfall: MAE against event totals, including runs too dry to score for phase. Level: precipitation-weighted mean bias and hourly MAE against the observed curve.
| Model | Snow–rain bracket vs observed 16% | Contains | Miss | Rain MAE, mm | Level bias, m | Level MAE, m | Gust MAE | |||
|---|---|---|---|---|---|---|---|---|---|---|
| ECMWF AIFS | 14–48% | 64% | 3.5% | 1 | 20.7 | 4 | -9 | 304 | 4= | — |
| ARPEGE 1–0 d only | 20–44% | 32% | 7.7% | 2 | 24.1 | 6= | -46 | 193 | 2 | 40 |
| JMA | 26–47% | 32% | 11.3% | 3 | 45.4 | 8 | -148 | 361 | 7 | — |
| UKMO | 26–57% | 18% | 12.1% | 4 | 24.1 | 6= | -149 | 146 | 1 | 33 |
| ECMWF IFS | 32–64% | 16% | 16.4% | 5 | 12.9 | 1 | -188 | 340 | 6 | 25 |
| ICON | 33–49% | 11% | 17.4% | 6 | 17.4 | 3 | -152 | 304 | 4= | 29 |
| GDPS 2–0 d only | 36–56% | 7% | 20.2% | 7 | 22.5 | 5 | -179 | 229 | 3 | 38 |
| GFS | 40–82% | 5% | 23.4% | 8 | 15.8 | 2 | -333 | 438 | 8 | 33 |
Ordered by phase miss — the distance of the observed fraction outside each model's bracket, averaged over runs; circled figures are each model's rank on that skill. ECMWF AIFS, JMA publish no gust field and are marked — rather than scored as failing.
Two models are scored on a fraction of the archive. A run is one model at one site, so the run counts read higher than the number of forecasts behind them; the models with full coverage rest on at least 10 initialisations each.
Lead time
Phase miss against lead time: how far the observed 16.1% snow fraction sat outside each model's bracket, averaged over the runs initialised that many days before the event. Lower is better; zero would mean the bracket contained what fell. ECMWF AIFS, ECMWF IFS and GFS are drawn out for reference and the run-weighted field mean is the heavy line; the rest are in grey. Every series is named at its right-hand end.
The field mean falls from 22.4% four days out to 9.2% on the day — the expected convergence, and evidence the models were responding to real information rather than drifting. But it converges to the wrong answer: even the runs initialised on the day still placed the observed snow fraction 9 points outside their bracket, and only 19% of them contained it. The error that mattered — a snow level a kilometre above the resorts at the height of the rain — was still being under-called hours out, not just days out.
| Model | 4 days out | 3 days | 2 days | 1 day | Same day |
|---|---|---|---|---|---|
| ECMWF AIFS | 1.3% 90% hit | 6.0% 50% hit | 4.5% 60% hit | 4.1% 60% hit | 1.1% 60% hit |
| ARPEGE 1–0 d only | — | — | — | 7.2% 47% hit | 8.5% 10% hit |
| JMA | 17.3% 15% hit | 4.7% n=3 67% hit | 8.9% 38% hit | 10.1% n=6 33% hit | 6.8% n=3 33% hit |
| UKMO | 0.0% n=5 100% hit | 18.1% 7% hit | 14.0% 0% hit | 11.9% 10% hit | 7.3% 20% hit |
| ECMWF IFS | 20.7% 0% hit | 17.4% 20% hit | 13.4% 30% hit | 20.6% 10% hit | 9.3% 20% hit |
| ICON | 28.9% 0% hit | 18.9% 10% hit | 12.7% 13% hit | 15.2% 20% hit | 11.9% n=5 0% hit |
| GDPS 2–0 d only | — | — | 27.0% n=5 0% hit | 14.3% n=5 20% hit | 19.3% n=5 0% hit |
| GFS | 56.7% 0% hit | 10.3% 20% hit | 28.9% 0% hit | 16.6% 0% hit | 14.2% 0% hit |
The same figures per model, ordered by mean miss across the leads, with the share of runs whose bracket contained the observed fraction beneath each. ECMWF AIFS did not converge late onto the right answer — it is the tightest row at every lead from four days out, and no other model that spans the full range stays close to it. That consistency, not a good same-day run, is what puts it top of the phase ranking.
Two cautions on reading individual rows. Coverage is uneven — ARPEGE and GDPS reach only part of the range in this archive, and cells marked n= rest on a handful of runs, so single-model kinks carry little weight. And the lead axis is bounded by the archive rather than the models: 4.8 days is as far back as this window reaches. Scored against the Sunday alone the archive reaches six days out, and at that range every model but JMA had its bracket centred on snow.
Attribution
The three skills above can be tied together on the quantity that matters — snow water — because the model's snow water is its rainfall multiplied by its phase. That makes the over-forecast separable: hold one term at what was observed, let the other keep its error, and see which one carries the mistake. This is the precipitation-weighted logic used for the observed snow level, applied to the models, so a phase error during heavy rain costs more than the same error in drizzle. Totals are pooled over the scored rows; the range is the snow–rain bracket, low end counting uncertain hours as rain.
| Error term | Snow-water error, mm | Share of total | |
|---|---|---|---|
| Rainfall (QPF) error alone model rainfall, observed phase | +321 | 3–11% | |
| Phase error alone observed rainfall, model phase | +2,882 to +9,971 | 96–100% | |
| Interaction the two errors acting together | −228 to −291 | small, negative | |
| Total over-forecast model rainfall, model phase | +2,995 to +9,975 | 100% |
The phase error accounts for essentially all of it. Given the observed phase, the models' own rainfall would have produced only +321 mm of error — a few per cent of the total. Given the observed rainfall, their own phase still produces +2,882 to +9,971 mm. The interaction term is small and negative, meaning the two errors slightly offset rather than compound: pooled across the field the models ran net dry on water while running far too snowy on phase, and the dry bias masked a little of the phase error. This is the sense in which "the models missed the snow" is imprecise — they did not miss the storm, or badly misjudge its water. They put the snow level in the wrong place.
Provenance and limits: the decomposition runs on
17,760 scored rows (17,390 with complete observed
thermometry). It runs on the same resort-point forecasts as the tables above
(legacy-resort-query; not exact gauge-cell forecasts) rather than exact native gauge-cell extractions, so
it inherits any grid-elevation mismatch. Baw Baw is excluded from this section for want of a matching Part-One gauge. The
scorecard records the conservative-cancellation test as
not satisfied, so the low and high ends of
the bracket are both reported rather than a single figure.
Synthesis
The attribution says which error mattered; it does not rank the models. No principled single score exists for that here — weighting phase against rainfall against level needs a stated use-case, and the archived phase derivation carries no probabilities to score properly. Two transparent syntheses are shown instead. Mean of ranks treats the three skills equally. Snow-water miss is the same physical quantity as the attribution, per model: how far each model's implied snow water landed from the observed 11 mm.
| Model | Phase | Rain | Level | Mean rank | Snow-water miss |
|---|---|---|---|---|---|
| ECMWF AIFS | 1 | 4 | 4= | 3.17 | in bracket |
| ARPEGE 1–0 d only | 2 | 6= | 2 | 3.50 | in bracket |
| UKMO | 4 | 6= | 1 | 3.83 | 8 mm |
| ECMWF IFS | 5 | 1 | 6 | 4.00 | 11 mm |
| ICON | 6 | 3 | 4= | 4.50 | 11 mm |
| GDPS 2–0 d only | 7 | 5 | 3 | 5.00 | 7 mm |
| JMA | 3 | 8 | 7 | 6.00 | in bracket † |
| GFS | 8 | 2 | 8 | 6.00 | 18 mm |
ECMWF AIFS leads on mean rank and ties on the other — the only model in the top half of all three skills, and one of two whose implied snow water bracket contained what fell, alongside ARPEGE. But the two columns disagree in instructive ways, and the disagreements are the point. JMA † comes within half a millimetre of the observed snow water while sitting second-last on mean rank: its 42 mm dry rainfall error and its warm-biased phase bracket very nearly cancel, and a synthesis that rewarded that would be measuring luck. GFS is the mirror case — second on rainfall, yet the largest snow-water miss in the field, because the synthesis makes its phase error pay in proportion to all the water it correctly forecast. Mean-of-ranks assumes the three skills matter equally and that rank gaps are even; neither is established, which is why the per-skill tables remain the primary result. It also gives every model one column regardless of what stands behind it, so GDPS and ARPEGE — 3 initialisations and 5 initialisations against 10 or more for the rest — weigh here exactly as much as models with the full archive behind them. That is a weakness of this table rather than of those models, and the strongest reason to read it after the per-skill ones rather than instead of them.
Sensitivity
Three choices sit under the tables above: the wet-bulb threshold for calling observed phase, whether a wide bracket is charged for its width, and which lead times are compared. Each is swept rather than asserted. The snow-level basis is not swept — it must be 0 °C to match what the models publish.
| Phase threshold | Observed fraction, 5 stations | All 8 usable | Top four, raw miss | Top four, width-penalised |
|---|---|---|---|---|
| 0.0 °C | 11.9% | 10.7% | AIFS · ARPEGE · JMA · UKMO | ARPEGE · AIFS · JMA · ICON |
| 0.5 °C | 12.6% | 11.6% | AIFS · ARPEGE · JMA · UKMO | ARPEGE · AIFS · JMA · ICON |
| 1.0 °C primary | 16.1% | 14.9% | AIFS · ARPEGE · JMA · UKMO | ARPEGE · AIFS · JMA · ICON |
| 1.5 °C | 25.4% | 23.5% | AIFS · ARPEGE · UKMO · JMA | ARPEGE · JMA · AIFS · ICON |
Rust marks a position that differs from the primary row. The observed fraction more than doubles across the sweep (11.9% → 25.4%) while the top of the ranking barely moves: ECMWF AIFS leads the raw-miss ordering at every threshold, and below it only JMA and UKMO change places at all. The width penalty is the choice that matters — charging half a bracket's width puts ARPEGE ahead of ECMWF AIFS at every threshold, since AIFS wins on raw miss partly by carrying a wider bracket. The penalty is a heuristic, not a proper interval score: there is no calibration behind the one-half, and a real interval score would need the model's own probabilities, which the archived derivation does not carry.
Gauge undercatch is a scenario, not a correction. Heated gauges lose snow in wind, and gusts reached 130 km/h, so 16.1% is a lower bound. Under assumed catch efficiencies: 100% catch → 16.1% · 80% catch → 19.3% · 70% catch → 21.5% · 50% catch → 27.7%. There is no local calibration for these instruments, so these are assumptions, kept separate from the threshold sweep — they are different kinds of uncertainty and pooling them into one range would misrepresent both.
The wet-bulb method is worth 18 m and changes no ranking. Solving the psychrometric equation at station pressure puts the observed level at 2068 m; Stull's sea-level fit gives 2050 m. Every model's bias shifts by the same amount, so the ordering is untouched — but an absolute bias is only meaningful with the method stated beside it.
Comparing like leads is two separate tests, and one of them bites. Restricting every model to the lead days all of them cover (192 runs) is a common lead-window sensitivity, not matched samples: the models still run at different cycle frequencies inside it. On that basis the phase order is ECMWF AIFS, ARPEGE, JMA, UKMO, ICON, ECMWF IFS, GFS and GDPS. Averaging within each of the 6 site × lead-day cells first and then weighting the cells equally — a stricter test, and the one that removes the cycle-rate advantage — keeps ECMWF AIFS top on phase, but it also moves the rainfall lead from ECMWF IFS to ECMWF AIFS, which is fourth on rainfall over the full archive and drops ECMWF AIFS from 4th to 7th on hourly level tracking. Genuinely matched forecasts barely exist here: 1 initialisation time (2026-08-07) is present for all eight models, at 6 site cases — too sparse to rank on.
Which archived copy of a run is scored changes nothing. Scoring the first archive of each initialisation — the forecast the application actually received — against the last before the event gives identical phase, rainfall and hourly-level rankings.
Footnote
Six ensemble systems cover these resorts. One run per system survives for this event — 7 August 00Z, 38 hours before it started — because ensemble output was not yet being archived; the run persisted in a short-lived cache. Rainfall only is scored, and the last column is the reason.
| System | Members | Observed inside 10–90% spread | Median member error, mm | Grid below resort top |
|---|---|---|---|---|
| ECMWF IFS ENS | 51 | 5 of 6 | +7.1 | 608 m |
| ECMWF AIFS ENS | 51 | 3 of 6 | -18.5 | 803 m |
| GEFS | 31 | 3 of 6 | -12.2 | 684 m |
| ICON-EPS | 40 | 2 of 6 | -7.0 | 861 m |
| GEPS | 21 | 1 of 6 | -29.0 | 1061 m |
| MOGREPS-G | 18 | 1 of 6 | -17.5 | 930 m |
These points are served at native grid elevation, 608–1061 m below the resorts' own top stations. The deterministic models above are placed at station height by the library's phase derivation; no equivalent exists for these members, so their snow fraction is reported in the retained record but not scored. A member reporting no snow at grid elevation is not an error: under a 2068 m snow level, a point that far down really did see rain. ECMWF IFS ENS was the best of the six, consistent with its deterministic sibling topping the rainfall table. One run is an anecdote, not a verification.
Provenance
var/forecast-history/ —
74 initialisations across eight models, scored at the verified resorts for
444 rows in all — one row per model, site and initialisation. Observations come from the Alpine Weather Dashboard at
data/observations/, snapshotted before scoring so the numbers reproduce.
Resort-reported new snow comes from the same dashboard's official-report store; it is context
only and enters no score.Mount Baw Baw is one instrument with two feeds, not two gauges. BoM serves it directly, and a Weather Chaser mirror of the same station stands in when that feed thins: through July the BoM feed ran 48 readings a day and the mirror was absent entirely; on 10 August BoM fell to eight readings and the mirror appeared for the first time, and on the 11th BoM managed two against the mirror's fifty-two. While it was failing the BoM feed also lagged — at 02:30 the mirror read 8.6 mm, BoM read 6.0 mm at 03:00, and the mirror read 10.0 mm at 03:30. Adding the two together, which an earlier version of this page did, produced 45.8 mm: a number that measured nothing. Scoring believes one feed per rain day, chosen on how many readings it filed — a question about coverage and not about the answer, since the counter is cumulative, so a denser feed buys resolution and never a larger total.
Verification window 2026-08-08T14:00:00Z → 2026-08-10T14:00:00Z · level basis wet-bulb 0 °C, pressure-aware · phase threshold +1.0 °C, swept 0–1.5 · run selection earliest archive per model/site/initialisation · 74 initialisations · 444 scored runs · observations snapshot 2026-08-14T10:48:07Z · source var/verification/events/part2-final-verification/scorecard.json