Football Model Calibration
Do predicted probabilities match observed frequencies?
Do predicted probabilities match observed frequencies?
A reliability question
Calibration asks whether probabilities correspond to long-run event frequencies. It complements the ability to distinguish stronger and weaker teams. A model can separate teams well while expressing too much certainty.
Reading a reliability diagram
Group similar forecasts into bins. Plot the mean prediction in each bin against the observed event frequency. Publish the sample count alongside each group, because a small bin can fluctuate substantially. For three football outcomes, inspect each class and declare how aggregate diagnostics are formed.
An evaluation plan
NinetyQuant plans to compare untouched pre-match snapshots with final results on a forward holdout period. Training, calibration and final evaluation should use separate time windows. Choices such as bin edges and minimum sample sizes need to be fixed and documented.
No curve without evidence
The demo does not show a synthetic reliability curve as if it were observed performance. Until a sufficiently large audited forecast record exists, calibration is listed as awaiting validation.
Sources & further reading
Prepared by NinetyQuant. Numerical examples are illustrative. References explain general concepts and do not validate NQ’s demo forecasts.