The proof
Is 75% more reliable than 70%?
Pick a market and see what actually happened every time the system published a given probability. If the model is calibrated, in each band the observed frequency lands close to the stated probability — and where it does not, you can see by how much.
The whole archive, refitting the model for each window on the data preceding it alone. These are the raw probabilities, before the calibrators: applying them here would be circular, since they are estimated on this very data. It is also why the miscalibration is visible.
No matches assessed with these filters.
The validation covers the leagues, not the European cups: those are forecast by the hierarchical model, which is not in this archive. This page is the aggregate. The individual matches of the last few days, one by one, are in Outcomes