पाऊस

The honesty board

Model vs. raw forecast

Every forecast we publish gets logged, then graded against what Mumbai's airport actually observed — no cherry-picking. Once the calibrated model is live, this page shows it beating the raw forecast on these same hours. If it can't, it doesn't ship.

Calibrated model — live

3628 / 200 labelled

Enough graded data to trust a model. The board below now grades the calibrated model against the raw forecast on a holdout it never trained on.

What we've collected

Snapshots logged
3,672

across 153 hourly runs

Graded rows
3,628

observed by METAR

Observed rain hours
1,915

of 3,628 graded

Awaiting a grade
44

future hours, not yet observed

The raw forecast's record

On the 3,628 graded hours, counting a rain call whenever the raw forecast read ≥ 0.3 mm·h⁻¹— the same line the Now page draws.

Raw forecast outcomes versus observed rain, on 3628 graded hours.
Observed
rain
Observed
dry
Forecast rain≥ 0.3 mm1792hits753false alarms
Forecast dry< 0.3 mm123misses960correct dry

The raw forecast cries wolf: 2545 rain calls, only 1792 right — and it still missed 123 of 1915 real rain hours. Cutting those false alarms is the whole job of the calibrated model.

What goes live at 200

When the board turns on, it adds the model's column: a Brier score measured on a time-ordered holdout the model never trained on. The model only ships when it beats both baselines — and a worse model can never replace a better one.

  1. 1 Beat the raw forecast on the holdout.
  2. 2 Beat the current champion model.
  3. 3 Otherwise — rejected. The champion stays.