← Model performance explorer

Model performance check: 23–25 August 2026

The window is closed

The scoring window runs to local midnight on 26 August. Every usable gauge is now observed to within ten minutes of that edge, and the 25 August rain day has since closed with at most 0.4 mm falling anywhere in the network between the window edge and 09:00 on the 26th. Nothing in this analysis is provisional, and no model is scored against a gap.

An earlier draft of this page was published while the period was still open, with observations ending at 20:00 AEST on 25 August. Closing it added 4.4 mm across the five verification gauges — nothing at Thredbo Top or Falls Creek, 0.2 mm at Perisher, 1.4 mm at Hotham and 2.8 mm at Buller — because rain went on dribbling at the Victorian gauges until close to midnight. Every model’s error grew slightly as a result, which is the direction the universal dry bias predicts. Rankings are otherwise unchanged, with one exception noted under Ensembles.

Executive summary

A warm, windy rain event with no snow at any station. The observed snow fraction is exactly 0.0 at all six usable gauges and at every wet-bulb threshold in the 0.0–1.5 °C sweep, so this period is a volume and false-phase check, not a phase-skill test.

Every model was much too dry. The best of them, UKMO, still missed by 27.2 mm — a mean relative error of 53% — and the worst, JMA, by 44.8 mm. The dry bias survives at short range: restricted to the lead-day bucket shared by every model, the errors only fall to 22.6–40.2 mm. This was a badly underforecast event across the board rather than a spread of good and bad calls.

The one clear discriminator is phase. GFS forecast a 39–57% snow fraction against an observed zero, and put its snow level 1,515 m below the solved observed height. Every other model kept its bracket in single or low double figures. ARPEGE called all rain and was exactly right, on fifteen runs at lead day 0 only.

Figures come from var/verification/events/period-2026-08-23/scorecard.json and var/verification/events/event-2026-08-24/scorecard.json.

What the gauges recorded

Rain reached the Victorian gauges around 22:50 AEST on 24 August and the New South Wales gauges an hour or so later, ran hard overnight, and eased through the 25th. The last measurable increment came at 18:20–20:10 in the west and north-east and just before midnight at Hotham and Buller. The window totals, on the ten-minute feed:

StationElevationPeriod total24 Aug rain day25 Aug rain day
Thredbo Top Station1957 m80.8 mm64.216.6
Perisher Valley1738 m74.2 mm52.621.6
Falls Creek1765 m50.2 mm45.44.8
Mount Hotham1849 m35.6 mm26.69.0
Mount Hotham Airport1295 m19.2 mm15.63.6
Mount Buller1707 m18.8 mm13.25.6

The five verification gauges total 259.6 mm, a mean of 51.9 mm. An independent recomputation straight from the snapshotted ten-minute NDJSON — rain-day maxima, not going through the harness — reproduces all six figures exactly. 23 August contributed nothing anywhere.

It was also a significant wind event: peak gusts of 104 km/h at Thredbo Top, 98 at Hotham, 82 at Falls Creek and 80 at Buller.

The event was strongly weighted to New South Wales. The two NSW verification gauges took the largest totals, and Buller — the most westerly Victorian site — took the least.

Precipitation ranking

Ranked by mean absolute error against the observed station total. Negative bias means too dry. The last column is the harness’s mean per-run relative error, which is not the same as MAE divided by the network mean.

RankModelRunsMAE (mm)Bias (mm)Mean relative error
1UKMO4527.2−18.252.9%
—ARPEGE1529.4−20.650.8%
2ICON4030.9−28.559.9%
3GDPS5534.8−31.067.0%
4ECMWF IFS11036.0−34.365.7%
5ECMWF AIFS11038.4−36.869.5%
6GFS11038.5−37.269.1%
7JMA8044.8−43.481.9%

ARPEGE is shown but not ranked: fifteen runs, all at lead day 0, is not a comparable sample against models carrying seven lead-day buckets.

Every bias is negative and every model missed by more than half the observed total. That is the headline result of this period and it does not depend on any of the caveats below.

It is not a long-lead artefact

The only lead-day bucket every model reaches in this period is day 0. Restricted to it — 125 runs in all — the order shuffles but the magnitude does not:

ModelRunsMAE (mm)Bias (mm)
ICON1022.6−18.5
ECMWF IFS2027.4−22.2
GDPS1027.9−17.4
UKMO2028.0−15.9
ARPEGE1529.4−20.6
ECMWF AIFS2029.5−25.7
GFS2032.5−31.1
JMA1040.2−38.8

Same-day forecast errors were still 44–77% of the observed network mean, and every bias is still negative. ICON leads here and UKMO drops to fourth, but with ten to twenty runs per model this is a sensitivity, not a second ranking.

Phase and snow level are not ranked

Nothing fell as snow. The observed snow fraction is 0.0 at every usable station across the whole PHASE_SENSITIVITY_C sweep (0.0, 0.5, 1.0, 1.5 °C wet-bulb), and stays 0.0 under every gauge-undercatch scenario down to 50% assumed snow catch efficiency. A model whose phase bracket “contains the observed fraction” here is simply a model that forecast rain.

The observed snow level is an extrapolation, not a measurement. The Victorian wet-bulb-zero regression solves to a precipitation-weighted 3,403 m — about 1,550 m above Mount Hotham, the highest contributing gauge — with a mean R² of 0.713 and a mean lapse rate of −0.49 °C/100 m. Individual wet hours reach 6,829 m where the profile flattens, and none falls below 2,324 m. Only 59% of solved hours clear R² 0.70, against 74% in the 16–22 August week. Every station was well above wet-bulb zero for the entire event, so solving for that height means projecting far outside the data the network can supply. Closing the window improved the fit slightly — the target moved 38 m and mean R² rose from 0.684 — but the objection was never the fit. It is the extrapolation distance, and that has not changed.

Both measures are therefore reported but not ranked, and this period carries no weight in the whole-period snow-level view on the performance explorer. That is a deliberate gap, not an oversight: folding a 3,403 m target into the rolling thermal combination would move every model’s published number by a height no station observed.

What the thermal numbers do support is a directional statement, because the sign is unambiguous even where the magnitude is not:

ModelPhase runsMean forecast levelLevel bias (m)Phase bracketBracket contains 0%
ARPEGE152,706 m−6970.0–0.0%100%
UKMO452,467 m−9361.1–12.8%76%
ICON332,402 m−1,0010.4–20.4%88%
GDPS392,383 m−1,0204.0–7.5%74%
ECMWF IFS762,378 m−1,0258.3–24.8%86%
ECMWF AIFS732,151 m−1,25212.4–31.8%73%
JMA372,110 m−1,29312.0–25.4%76%
GFS801,888 m−1,51539.1–56.6%44%

Every model placed its snow level below the solved observed height, and the ordering is consistent across the two wet-bulb calculations (the Stull sensitivity moves the observed level by 8 m and each model bias by 8–9 m). Of the 398 runs that defined a bracket at all, 288 contained the observed zero and 110 sat entirely above it; none sat below.

GFS is the outlier that matters operationally, and it is the one thermal result that survives without trusting the 3,403 m target at all. It assigned 39–57% of its precipitation to snow against an observed zero, and its bracket failed to contain that zero on 56% of its phase-capable runs — the only model that missed more often than it hit. Its mean forecast level of 1,888 m is 222–818 m below every other model’s, low enough to put snow in Thredbo’s 2,037 m forecast band and “uncertain” in several of the 1,780–1,845 m bands beneath it, through an event that was rain at every gauge from 1,295 m up.

Ensembles — precipitation only

SystemRunsMembersObserved total inside p10–p90Median error (mm)Grid below station
MOGREPS-G701837.1%−34.2920 m
GEPS802133.8%−39.51,020 m
ECMWF IFS ENS905133.3%−35.1609 m
ICON-EPS904026.7%−43.6870 m
ECMWF AIFS ENS1005126.0%−41.7818 m
GEFS1053118.1%−44.2693 m

The same story as the deterministic side, and worse. Every median is 34–44 mm too dry, and the best system’s 10th–90th percentile band failed to contain the observed total on nearly two runs in three. The event fell outside ensemble spread, not merely outside the ensemble mean.

This is the one place where closing the window changed an ordering. On the open window ECMWF IFS ENS led containment at 41.1% and MOGREPS-G sat third at 32.9%. Seven ECMWF IFS ENS runs flipped from contained to missed — five at Hotham and two at Buller, where the observed total climbed past p90s of 34.8 and 18.7 mm — while MOGREPS-G gained three at Buller, where a p10 of 16.1 mm had sat a tenth of a millimetre above the old observed total and now sits below the new one. A ranking that a 4 mm revision can invert is a ranking held very loosely, and the substantive finding — that no system contained the observed total on even two runs in five — is unaffected.

Ensembles are not ranked on phase: the harness reports native grid elevations 609–1,020 m below the stations and marks snow fraction non-comparable.

Nested event lens — 24–25 August

The 23rd was dry, so an event-only window that drops it changes the observed totals not at all and only admits more runs (670 against 565) by shortening the span a run must cover.

RankModelRunsMAE (mm)Bias (mm)
1ICON5026.7−23.5
2UKMO6027.7−16.4
3ARPEGE3029.2−23.3
4GDPS6034.0−30.1
5ECMWF IFS12535.6−33.4
6GFS12537.2−36.3
7ECMWF AIFS12538.0−36.2
8JMA9544.5−42.6

ICON and UKMO swap, ARPEGE reaches lead day 1 and becomes rankable, and GFS moves ahead of AIFS on 0.8 mm. Nothing else moves and no bias changes sign. The lead-day-0-and-1 sensitivity is kinder to ICON still — 16.2 mm, comfortably the best figure any model records anywhere in this event — but on twenty runs.

This is a nested lens on the period above, not a fifth consecutive window, and it is excluded from the rolling continuous record so that no date is counted twice.

How this fits the rolling record

The performance explorer now carries four adjacent primary periods spanning 9–25 August:

  1. 9–10 August: a warm rain-on-snow event.
  2. 11–15 August: a near-dry bridge, useful for false alarms.
  3. 16–22 August: a whole-week check containing a three-day event.
  4. 23–25 August: this period — a heavy, entirely warm rain event.

The whole-period precipitation view extends from fourteen days to seventeen and now reads 4.95 mm/day for ECMWF IFS, ahead of GDPS at 5.66, ECMWF AIFS at 6.01, GFS at 6.11 and JMA at 7.78. ECMWF IFS keeps the lead it held over the fourteen-day span. In the matched short-range view, which holds all seven models to a shared daily lead-day bucket, ECMWF IFS takes first at 4.49 mm/day and UKMO falls from first to third at 5.04: it led the fourteen-day span, and the three days added here were not days on which it was closest on volume.

The whole-period snow-level views are unchanged, to the metre, and still cover 9–22 August across three of the four periods, for the reason set out above. The page says so beside both thermal charts rather than quietly averaging seventeen days of precipitation against fourteen days of snow level.

Confidence and gaps

  • Thredbo Village (95908) is excluded entirely: five readings across the three days is far too sparse to difference a counter. Its counter did read 70.1 mm across those readings, which corroborates Thredbo Top’s 80.8 mm rather than contradicting it.
  • Mount Baw Baw (95901) is excluded from precipitation by the caller, matching the two preceding periods, and remains in the thermal regression. Its own ten-minute counter reports 20.0 mm on a feed that stops at 13:00 AEST on 25 August; that figure is not scored.
  • Lead-day coverage differs by model and is reported as produced. Period: ECMWF IFS and GDPS 0–6; ECMWF AIFS 0–3 and 6; GFS and JMA 0–4 and 6; ICON and UKMO 0–3; ARPEGE 0 only. MAX_LEAD_DAY = 6 folds day 6 and beyond into bucket 6 rather than dropping those runs.
  • Ten initialisations were archived more than once. The default --select earliest scores the first archive received; the variant spread is a median of 0.27 m and a maximum of 59 m on mean snow level, and zero on precipitation.
  • The snow-level regression is Victorian, and this event was heaviest in New South Wales. NSW has two usable stations and cannot support a regression, so there is no NSW thermal figure here — the two largest gauge totals in the event contribute to rainfall scoring and to nothing else.
  • One warm event settles nothing about cold-season phase skill. What it does establish is that on 24–25 August 2026 every available model, deterministic and ensemble, was substantially too dry, and that GFS alone forecast a resort-level snow event that did not happen.

Reproducibility

npm run verify:event -- --start 2026-08-23 --end 2026-08-26 --slug period-2026-08-23 --exclude 95908 --exclude-precip 95901 --precip-source-contains bom_10m
npm run verify:event -- --start 2026-08-24 --end 2026-08-26 --slug event-2026-08-24 --exclude 95908 --exclude-precip 95901 --precip-source-contains bom_10m

Both runs used the default earliest-archive selection and snapshotted the observations first. Both were rerun on 26 August against a closed window; the figures on this page are from those runs. python3 reports/model_performance_data.py rebuilds the explorer payload from the retained scorecards. Relative percentage precipitation error is undefined at a dry gauge; no gauge was dry in this window, so every station contributes to the relative-error field as well as to millimetre MAE and bias.