Model performance check: 16–22 August 2026
Executive summary
ECMWF IFS produced the best all-round forecast across the full week. It had the lowest precipitation MAE and the smallest snow-level bias, although GDPS tracked the snow level more closely in absolute-error terms. UKMO and ICON performed best during the three-day event itself but did not have complete forecasts spanning the full seven-day window. GFS ranked last on the raw weekly precipitation and snow-level measures, yet was competitive at short lead times; that apparent contradiction is examined separately below.
The primary assessment covers 16–22 August and scores the earliest archived forecast received by the application. A second window isolates the precipitation event of 19–21 August. The full-week result is the headline because the dry days test false-alarm rainfall as well as the main event. Figures come from var/verification/events/week-2026-08-16/scorecard.json and var/verification/events/event-2026-08-19/scorecard.json; dry-day findings use the adjacent attribution.json files.
Precipitation ranking
The ten-minute-feed totals passed the independent check exactly: Hotham 69.8 mm, Falls Creek 64.4, Thredbo Top 58.2, Buller 40.8, Perisher 35.2 and Hotham Airport 13.8. Against those totals, the full-week deterministic ranking is:
| Rank | Model | Runs | MAE (mm) | Bias (mm) |
|---|---|---|---|---|
| 1 | ECMWF IFS | 70 | 27.3 | -25.7 |
| 2 | GDPS | 25 | 28.0 | -26.1 |
| 3 | JMA | 25 | 34.9 | -32.5 |
| 4 | ECMWF AIFS | 70 | 37.1 | -37.1 |
| 5 | GFS | 70 | 38.4 | -36.1 |
ECMWF IFS therefore leads the whole-week volume ranking, narrowly ahead of GDPS. Every bias is negative: the dominant failure was underforecast volume. All five models nevertheless produced non-zero QPF on at least one of the dry local dates (16–18 or 22 August), so none called the week’s dry margins perfectly.
On the event alone, UKMO was a clear volume winner and ICON second. Among models present in both windows, GDPS overtook ECMWF IFS, while ECMWF AIFS and GFS both overtook JMA. UKMO and ICON have no complete full-week score and therefore cannot displace the weekly winner; ARPEGE has only five event runs and is insufficient to rank. Complete deterministic results for both windows are in Appendix A.
Why the GFS sensitivity result reverses the order
The headline ranking and the sensitivity answer different questions. The headline averages every complete run the application received, so models are represented at the lead times and cycle frequencies actually present in the archive. GFS contributes 70 runs to that result and ranks fifth at 38.4 mm MAE.
The sensitivity then makes two changes:
- It retains only lead days 0–2, the short-range buckets shared by every model. This removes 40 GFS runs from lead days 4–6+, leaving 30 runs. GFS MAE falls to 27.7 mm and it moves from fifth to third.
- It gives each of the 15 resort × lead-day cells equal weight, so a model with more forecast cycles inside a cell does not count more heavily. The equal-cell GFS MAE is 24.7 mm, narrowly first.
The equal-cell calculation is built from runs that define a usable phase bracket. All 30 short-lead GFS runs qualify, while the corresponding counts are 29 for ECMWF AIFS, 30 for ECMWF IFS, 20 for GDPS and 20 for JMA. It is therefore not an exactly matched precipitation comparison, and only five site-initialisation cases are shared exactly across every model.
The reversal is still informative: GFS’s poor weekly score was concentrated in the longer-lead and operational sample mix, whereas its short-range, equally weighted resort performance was competitive. It does not overturn ECMWF IFS as the best model on the forecasts the application actually received. Appendix D shows each stage of the comparison.
Temperature, snow level and phase
Cold-tail caveat: almost all precipitation was rain; the pooled observed snow fraction was only 6.4%, and phase skill rests mainly on the small, drying cold tail on 21 August. Snow level and phase are inferred from wet-bulb temperature, not measured. The full-week Victorian observed precipitation-weighted snow level was 2,381 m (179 precipitating timestamps; mean regression R² 0.856).
| Rank by level MAE | Model | Runs / phase runs | Level MAE (m) | Level bias (m) | Mean phase bracket | Bracket hit |
|---|---|---|---|---|---|---|
| 1 | GDPS | 25 / 25 | 229 | -271 | 15.5–31.4% | 20.0% |
| 2 | JMA | 25 / 20 | 270 | -341 | 14.8–38.7% | 45.0% |
| 3 | ECMWF IFS | 70 / 70 | 272 | -131 | 2.8–21.0% | 61.4% |
| 4 | ECMWF AIFS | 70 / 59 | 381 | -312 | 2.9–40.7% | 86.4% |
| 5 | GFS | 70 / 42 | 389 | -381 | 15.9–43.2% | 23.8% |
GDPS has the lowest snow-level MAE, but ECMWF IFS is the more balanced thermal call: only 43 m worse in MAE, with by far the smallest magnitude bias and a phase bracket much closer to the observed 6.4%. GDPS, JMA and GFS were systematically too snowy. AIFS’s high bracket-hit rate is qualified by its very wide average bracket. On the event alone, UKMO led level MAE at 167 m, followed by ICON at 184 m; neither has a full-week score. ARPEGE’s four phase-capable runs are insufficient. ECMWF AIFS and JMA gusts are unavailable by design, not failures. Complete thermal, phase and gust results are in Appendix B.
The network-wide ten-minute outage on 19 August (10:40–14:30 and 14:50–19:10 AEST) limits intra-day phase timing: accumulation was landed on isolated readings, including 14.6 mm at Falls Creek. Totals remain valid, and the concealed intervals were unambiguously rain, so the weekly volume ranking is unaffected. Temperature, humidity, wet-bulb and the snow-level regression use every reading and are unaffected by precipitation feed pinning.
Ensembles — precipitation only
ECMWF IFS ENS had the strongest usable full-week spread coverage. MOGREPS-G matched its 40.0% containment rate but has only five full-week runs and is insufficient to rank. Event-only coverage was weak and reordered every system: MOGREPS-G led, narrowly ahead of ICON-EPS, while ECMWF IFS ENS fell to fifth. Complete ensemble results are in Appendix C.
All ensemble median errors are negative. Ensembles are not ranked on phase: the harness reports native grid elevations 609–1,020 m below the stations and marks snow fraction non-comparable.
Overall assessment
No single model won every measure, but ECMWF IFS was the best all-round model for the week. It won the headline full-week precipitation MAE, had the smallest snow-level bias by a wide margin, defined a usable phase bracket on all 70 runs, and its ensemble counterpart had the best adequately sampled full-week containment. Its weaknesses were a still-large 25.7 mm dry bias and only middling snow-level MAE.
GDPS was the strongest challenger and the best pure snow-level tracker. It missed ECMWF IFS by only 0.7 mm on weekly precipitation MAE, beat every other model on snow-level MAE, and became the best of the common models during the event-only window. The reason it is not the holistic winner is calibration: its mean snow level was 271 m too low and its phase bracket was entirely too snowy often enough to contain the observed fraction on only 20% of phase-capable runs.
UKMO and ICON were the event specialists. UKMO led event-only precipitation and snow-level MAE; ICON was second on both and had the lowest available event gust MAE. Their shorter full-window availability prevents a weekly verdict, rather than counting against their forecast quality. ECMWF AIFS produced the highest full-week phase hit rate, but did so with the widest bracket and materially underforecast volume. JMA was competitive on weekly snow-level MAE but strongly low-biased and fell to last on event-only precipitation among adequately sampled models.
GFS was the weakest raw full-week deterministic result, finishing last on both precipitation and snow-level MAE and forecasting a phase mix that was much too snowy. Its short-range result was substantially better, however: it rose to third after restricting the comparison to lead days 0–2 and to first only after resort × lead-day cells were weighted equally. The week therefore supports “weakest on the operational weekly sample,” not a general claim that GFS is intrinsically worst. Across ensembles, none was convincing: all median forecasts were too dry, and even the best adequately sampled full-week system contained the observed total only 40% of the time.
Confidence and gaps
- Thredbo Village (95908) is excluded entirely because its record is too sparse. Baw Baw (95901) is excluded from precipitation because counter gaps span daily resets, but remains in the thermal regression. Its scorecard fields are
precipUsable: false,precipTotalUsable: falseandprecipNote: "excluded by the caller". Every other gauge is usable for both total and phase; the harness attached no further disqualification. - The observed precipitation side uses the ten-minute feed alone. This week is therefore not constructed identically to earlier August event scorecards that used feed election across the reconciled archive.
- Lead-day phase buckets are reported as produced. Full week: AIFS/IFS/GFS 0,1,2,4,5,6; GDPS 0–3; JMA 0–2. Event: AIFS/IFS/GDPS/GFS/JMA 0–6; UKMO/ICON 0–3; ARPEGE 0. The archive does contain longer leads (including 8 August initialisations), but
MAX_LEAD_DAY = 6folds day 6 and beyond into bucket 6; it does not drop those runs. - One warm event with a narrow cold tail cannot settle general phase skill, long-lead skill, or NSW regional snow level. Missing complete full-week samples are left missing; nothing was substituted, interpolated or widened.
Appendix A — deterministic precipitation tables
Ranked by MAE against the observed event total. Negative bias means too dry.
Full week, 16–22 August
| Rank | Model | Runs | MAE (mm) | Bias (mm) | MAE / observed |
|---|---|---|---|---|---|
| 1 | ECMWF IFS | 70 | 27.3 | -25.7 | 50.3% |
| 2 | GDPS | 25 | 28.0 | -26.1 | 49.9% |
| 3 | JMA | 25 | 34.9 | -32.5 | 62.1% |
| 4 | ECMWF AIFS | 70 | 37.1 | -37.1 | 67.9% |
| 5 | GFS | 70 | 38.4 | -36.1 | 70.9% |
Event only, 19–21 August
| Rank | Model | Runs | MAE (mm) | Bias (mm) | MAE / observed | Standing |
|---|---|---|---|---|---|---|
| 1 | UKMO | 30 | 15.4 | -7.9 | 29.2% | Ranked |
| 2 | ICON | 35 | 21.1 | -16.1 | 39.0% | Ranked |
| 3 | GDPS | 45 | 32.3 | -30.9 | 58.7% | Ranked |
| 4 | ECMWF IFS | 95 | 33.9 | -33.4 | 64.0% | Ranked |
| — | ARPEGE | 5 | 34.3 | -34.3 | 65.1% | Insufficient sample |
| 5 | ECMWF AIFS | 100 | 38.5 | -38.5 | 72.6% | Ranked |
| 6 | GFS | 100 | 41.2 | -40.6 | 77.2% | Ranked |
| 7 | JMA | 55 | 42.0 | -41.1 | 76.3% | Ranked |
Appendix B — deterministic thermal and phase tables
Ranked by snow-level MAE. Bias is forecast minus the observed precipitation-weighted level. “Phase runs” are runs wet enough to define a bracket; bracket hit is the share whose forecast snow-fraction bracket contained the observed fraction. Gust MAE is atmospheric context, not part of the snow-level rank.
Full week, 16–22 August
| Rank | Model | Runs | Phase runs | Level MAE (m) | Level bias (m) | Mean phase bracket | Bracket hit | Gust MAE (km/h) |
|---|---|---|---|---|---|---|---|---|
| 1 | GDPS | 25 | 25 | 229 | -271 | 15.5–31.4% | 20.0% | 63 |
| 2 | JMA | 25 | 20 | 270 | -341 | 14.8–38.7% | 45.0% | Unavailable |
| 3 | ECMWF IFS | 70 | 70 | 272 | -131 | 2.8–21.0% | 61.4% | 49 |
| 4 | ECMWF AIFS | 70 | 59 | 381 | -312 | 2.9–40.7% | 86.4% | Unavailable |
| 5 | GFS | 70 | 42 | 389 | -381 | 15.9–43.2% | 23.8% | 66 |
Event only, 19–21 August
| Rank | Model | Runs | Phase runs | Level MAE (m) | Level bias (m) | Mean phase bracket | Bracket hit | Gust MAE (km/h) | Standing |
|---|---|---|---|---|---|---|---|---|---|
| 1 | UKMO | 30 | 30 | 167 | -167 | 1.0–12.9% | 66.7% | 46 | Ranked |
| 2 | ICON | 35 | 34 | 184 | -98 | 1.3–10.4% | 50.0% | 18 | Ranked |
| — | ARPEGE | 5 | 4 | 200 | +41 | 0.0–4.1% | 25.0% | 53 | Insufficient sample |
| 3 | GDPS | 45 | 43 | 205 | -142 | 1.9–11.6% | 48.8% | 64 | Ranked |
| 4 | ECMWF IFS | 95 | 88 | 261 | -54 | 0.9–9.5% | 63.6% | 50 | Ranked |
| 5 | JMA | 55 | 33 | 283 | -151 | 0.3–13.3% | 39.4% | Unavailable | Ranked |
| 6 | ECMWF AIFS | 100 | 60 | 328 | -243 | 0.6–14.5% | 73.3% | Unavailable | Ranked |
| 7 | GFS | 100 | 62 | 346 | -242 | 0.9–15.5% | 56.5% | 71 | Ranked |
Appendix C — ensemble precipitation tables
Ranked by the share of runs whose 10th–90th percentile member range contained the observed total. Median error is ensemble median minus observed total. Phase is deliberately excluded.
Full week, 16–22 August
| Rank | System | Runs | Members | Observed in spread | Median error (mm) | Grid below station (m) | Standing |
|---|---|---|---|---|---|---|---|
| 1 | ECMWF IFS ENS | 30 | 51 | 40.0% | -26.1 | 609 | Ranked |
| — | MOGREPS-G | 5 | 18 | 40.0% | -24.4 | 920 | Insufficient sample |
| 2 | ECMWF AIFS ENS | 30 | 51 | 23.3% | -29.6 | 818 | Ranked |
| 3 | GEPS | 20 | 21 | 20.0% | -37.3 | 1,020 | Ranked |
| 4 | ICON-EPS | 25 | 40 | 16.0% | -39.4 | 870 | Ranked |
| 5 | GEFS | 30 | 31 | 13.3% | -34.0 | 693 | Ranked |
Event only, 19–21 August
| Rank | System | Runs | Members | Observed in spread | Median error (mm) | Grid below station (m) |
|---|---|---|---|---|---|---|
| 1 | MOGREPS-G | 30 | 18 | 16.7% | -37.1 | 920 |
| 2 | ICON-EPS | 50 | 40 | 16.0% | -38.4 | 870 |
| 3 | GEPS | 45 | 21 | 13.3% | -39.9 | 1,020 |
| 4 | ECMWF AIFS ENS | 55 | 51 | 10.9% | -36.0 | 818 |
| 5 | ECMWF IFS ENS | 55 | 51 | 9.1% | -33.2 | 609 |
| 6 | GEFS | 65 | 31 | 6.2% | -39.7 | 693 |
Appendix D — lead and sample-mix sensitivity
The table follows the ranking through three stages. “All runs” is the operational weekly result. “Lead 0–2” removes the differing longer-lead samples but still counts every retained cycle. “Equal cells” averages within each resort × lead-day cell before giving all 15 cells the same weight; its run count is phase-capable runs contributing to those cells.
| Model | All runs: rank · MAE · n | Lead 0–2: rank · MAE · n | Equal cells: rank · MAE · phase n |
|---|---|---|---|
| ECMWF IFS | 1 · 27.3 · 70 | 4 · 29.0 · 30 | 4 · 28.6 · 30 |
| GDPS | 2 · 28.0 · 25 | 1 · 25.2 · 20 | 3 · 25.4 · 20 |
| JMA | 3 · 34.9 · 25 | 5 · 34.9 · 25 | 5 · 32.4 · 20 |
| ECMWF AIFS | 4 · 37.1 · 70 | 2 · 25.3 · 30 | 2 · 25.2 · 29 |
| GFS | 5 · 38.4 · 70 | 3 · 27.7 · 30 | 1 · 24.7 · 30 |
MAE is in millimetres. The equal-cell snow-level results remain GDPS 186 m, ECMWF IFS 239 m, ECMWF AIFS 243 m, GFS 267 m and JMA 269 m; GFS’s precipitation reversal does not extend to its thermal performance.
Reproducibility
npm run test:event-verification
npm run verify:event -- --start 2026-08-16 --end 2026-08-23 --slug week-2026-08-16 --precip-source-contains 10-minute --exclude 95908 --exclude-precip 95901
npm run verify:event -- --start 2026-08-19 --end 2026-08-22 --slug event-2026-08-19 --precip-source-contains 10-minute --exclude 95908 --exclude-precip 95901
The verification harness passed all 91 tests before scoring. Both runs used the default --select earliest and wrote observation snapshots; neither used --no-snapshot.