Calibration
When the model says 60%, it should happen about 60% of the time. Each dot is a bucket of predictions; the diagonal is perfect calibration.
Rolling RPS
Ranked Probability Score over a trailing 30-day window. Lower is better; the dashed line is the baseline window.
Season by season
Every figure here is produced by a scheduled job from the archived prediction log and the raw match results. Alerts are advisory: a retrained model is only allowed to go live after it passes the same out-of-sample test published above.