Model Arena
Different models see football differently. We are building a transparent record of how each one performs.
Evaluation Board
Evaluation protocol| MODEL | STATUS | MATCHES EVALUATED | BRIER ↓ | LOG LOSS ↓ | ACCURACY | CALIBRATION |
|---|---|---|---|---|---|---|
| Elo | Illustrative | — | — | — | — | Not evaluated |
| Poisson | Illustrative | — | — | — | — | Not evaluated |
| Expected Goals | Illustrative | — | — | — | — | Not evaluated |
| Recent Form | Illustrative | — | — | — | — | Not evaluated |
| Bayesian | Planned | — | — | — | — | Not evaluated |
| Machine Learning | Planned | — | — | — | — | Not evaluated |
| NQ Ensemble | Illustrative | — | — | — | — | Not evaluated |
Performance will only be reported for timestamped pre-match predictions evaluated against verified results.
Different lenses on the same game.
Elo
Relative team strength, updated after each result and adjusted for home advantage.
Demo v0.1 · Illustrative outputs only.
Poisson
A score distribution built from expected scoring rates, assuming independent goal counts.
Demo v0.1 · Illustrative outputs only.
Expected Goals
Underlying attacking and defensive performance through the quality of chances.
Demo v0.1 · Illustrative outputs only.
Recent Form
Recent performance with more weight placed on newer matches.
Demo v0.1 · Illustrative outputs only.
Bayesian
A planned framework for updating team-strength estimates and representing uncertainty.
Planned · No implementation or forecasts published.
Machine Learning
Planned nonlinear models, evaluated on time-separated holdout data.
Planned · No implementation or forecasts published.
NQ Ensemble
A proposed blend of complementary models. Current outputs are illustrative fixtures.
Demo v0.1 · Illustrative outputs only.
Probability, not certainty.
A 65% win probability still leaves a 35% chance of a different result. Over many comparable matches, a well-calibrated 65% forecast should win about 65% of the time.
Learn to read probabilities