The honesty board
Model vs. raw forecast
Every forecast we publish gets logged, then graded against what Mumbai's airport actually observed — no cherry-picking. Once the calibrated model is live, this page shows it beating the raw forecast on these same hours. If it can't, it doesn't ship.
Calibrated model — live
3628 / 200 labelled
Enough graded data to trust a model. The board below now grades the calibrated model against the raw forecast on a holdout it never trained on.
What we've collected
- Snapshots logged
- 3,672
- Graded rows
- 3,628
- Observed rain hours
- 1,915
- Awaiting a grade
- 44
across 153 hourly runs
observed by METAR
of 3,628 graded
future hours, not yet observed
The raw forecast's record
On the 3,628 graded hours, counting a rain call whenever the raw forecast read ≥ 0.3 mm·h⁻¹— the same line the Now page draws.
| Observed rain | Observed dry | |
|---|---|---|
| Forecast rain≥ 0.3 mm | 1792hits | 753false alarms |
| Forecast dry< 0.3 mm | 123misses | 960correct dry |
- Base rate52.8%of graded hours actually had rain
- Detection · POD93.6%of 1915 rain hours, 1792 were forecast
- False-alarm ratio29.6%of 2545 rain calls, 753 were wrong
- Raw agreement75.9%forecast matched observation
The raw forecast cries wolf: 2545 rain calls, only 1792 right — and it still missed 123 of 1915 real rain hours. Cutting those false alarms is the whole job of the calibrated model.
What goes live at 200
When the board turns on, it adds the model's column: a Brier score measured on a time-ordered holdout the model never trained on. The model only ships when it beats both baselines — and a worse model can never replace a better one.
- 1 Beat the raw forecast on the holdout.
- 2 Beat the current champion model.
- 3 Otherwise — rejected. The champion stays.