The proof
Is 75% more reliable than 70%?
Pick a market and see what happened every time we stated a given probability. If the model is well tuned, in every band the observed frequency lands close to the stated one. Where it does not, you see by how much.
The whole archive: for each period we refit the model using only the data that existed before it. These are the raw probabilities, without the final correction, because that correction was worked out on these very matches and using it here would be circular reasoning. It is also why the gap shows up clearly.
- matches assessed
- 36,712
- average stated probability
- 50.1%
- observed frequency
- 52.4%
- gap
- −2.27%
perfect calibration
Band by band
| when we said | times | it happened | 95% interval |
|---|---|---|---|
| 10% – 20% | 32 | 50.0% | too few matches |
| 20% – 30% | 979 | 40.1% | 37.1% – 43.2% |
| 30% – 40% | 5,758 | 45.1% | 43.8% – 46.3% |
| 40% – 50% | 11,633 | 49.0% | 48.1% – 49.9% |
| 50% – 60% | 11,506 | 54.1% | 53.2% – 55.1% |
| 60% – 70% | 5,544 | 62.4% | 61.1% – 63.6% |
| 70% – 80% | 1,131 | 65.4% | 62.6% – 68.1% |
| 80% – 90% | 126 | 68.3% | 59.7% – 75.7% |
| 90% – 100% | 3 | 100.0% | too few matches |
Below fifty matches a band says nothing solid. The interval shows you that.
This count covers the leagues. The European cups use a different model and stay out of it. This page is the aggregate. The individual matches of the last few days, one by one, are in Outcomes