EVALUATION · 3 MIN READ · UPDATED 2026-09-21

Why Prediction Accuracy Is Misleading

A correct label can hide an unhelpful probability.

A correct label can hide an unhelpful probability.

What accuracy measures

Accuracy is the fraction of matches where the model’s highest-probability outcome occurs. It is easy to understand, but discards the strength of the forecast and the remaining probability distribution.

The same pick, different forecasts

Consider two forecasts for the same fixture: 40% / 35% / 25% and 90% / 5% / 5%. Both choose the home team. Accuracy treats them identically whether the home team wins or loses, even though they express very different degrees of certainty.

A more useful scorecard

NinetyQuant’s proposed evaluation combines accuracy with multiclass Brier score, log loss and calibration diagnostics. Scores must use the same match cohort and forecast horizon before they can be compared. A time-separated evaluation protects against selecting parameters using the results being scored.

Current status

The v0.1 Model Arena contains no verified historical performance. Blank scores are deliberate. Real evaluation requires archived predictions generated before kickoff and independently recorded results.

Sources & further reading

Prepared by NinetyQuant. Numerical examples are illustrative. References explain general concepts and do not validate NQ’s demo forecasts.

Explore the NQ methodology