ProofOdds

The live record · updated after every matchday

How we are actually doing

One question, asked the same way every week: over the matches we published before kickoff, is our log loss below the closing line or above it? Lower is better. This page shows the answer whatever it is.

ProofOdds
1.0161
log loss, 599 matches
Closing line
0.9788
same matches, same scoring
Gap per match
+0.0374
±0.0151 at 95% · 599 matches
Top pick correct
49.7%
market 51.6%
How to read the gap Predicting 1/3-1/3-1/3 every week scores 1.0986. The closing line scores 0.9788. So everything anyone knows about football is worth 0.1198 nats, and this model captures 69% of it — though 599 matches only pin that figure between 56% and 81%.
How much of this is signal Not much of it yet, and the honest thing is to say so before somebody else does. The gap is +0.0374 ± 0.0151 per match at 95% confidence — an interval of +0.0223 to +0.0524 over 599 graded matches (standard error 0.0077, t = 4.86). That interval excludes zero, so this sample does now tell the model and the closing line apart. Because both are scored on the same matches, the difference is paired and its interval is far tighter than either log loss alone; it still needs a few thousand matches to close on a gap this size. The number with the sample behind it is the walk-forward backtest: 2660 matches, and there the model's loss is comfortably separated from the market's. Recompute it: diff = model_loss - market_loss per match, then mean ± 1.96 × sd / √n — or run scripts/rescore.py --csv audit.csv from a clone and average the columns yourself. That script recomputes every figure on this page from the sealed files by a second implementation and fails loudly if the two disagree.

Where the ground is lost

-4.7+10.5+25.7 2026-08-28 2026-10-09
Cumulative log loss against the market-average closing line. Above the rule the model is behind; below it, ahead.

Is a stated 30% really 30%?

said 8%, happened 4% (26 outcomes)said 16%, happened 16% (187 outcomes)said 26%, happened 26% (763 outcomes)said 35%, happened 35% (324 outcomes)said 45%, happened 42% (262 outcomes)said 54%, happened 51% (131 outcomes)said 64%, happened 72% (68 outcomes)said 74%, happened 81% (27 outcomes)said 83%, happened 67% (9 outcomes) 0% 100% 100%
Where the dots sit on the diagonal, a stated probability means what it says. Dot size is the number of outcomes in the bin.

Calibration and sharpness are different things, and only one of them is our problem. If the dots sit on the diagonal, a stated probability means what it says — the model is not overconfident, it is simply not seeing as much as the market sees.

A model can be perfectly calibrated and still lose on log loss, by hedging towards the base rates when the market is right to commit. That is what the gap above is mostly made of.

Over/under 2.5 goals

Fewer matches than the table above, and a reader is owed the reason. Every fixture is graded from the first entry that published it — one fixture, one file, so anyone checking by hand knows exactly where to look — and entries sealed before this market was added carry no probability for it. The Premier League therefore joins the totals record in September rather than August. That is a choice about how auditable we want the rule to be, not a law of nature. The reference points differ too — guessing 1/3-1/3-1/3 on a result scores 1.0986, guessing a coin flip on a half-goal line scores 0.6931 — so averaging the two would produce a number that means nothing.

ProofOdds
0.6830
log loss, 555 matches
Closing line
0.6588
market-average closing total
Gap per match
+0.0242
±0.0152 at 95%
Went over
59.8%
of the graded matches
Do not compare these two gaps directly Everything anyone knows about a football result is worth about 0.135 nats on the 1X2 market. On total goals it is worth about 0.020 — there is roughly seven times less to know, because a closing total barely beats a coin flip. A model therefore sits closer to the market on goals almost regardless of how good it is, and reading a smaller gap as "better at goals" is backwards. The comparison that means something is the share of what was there to win: 69% on the result, 30% on the total — both still wide enough at this sample (56–81% and -15–74%) that the difference between them is not yet a finding.

Two groups, counted apart

The headline above pools every division. That is the honest total and it stays, but a pooled log loss moves when the mix of divisions moves, whether or not anything about the model has — so the pooled line on its own would let a composition change look like a result. On 17 September 2026 twelve thinner divisions were added to the eleven this site had been pricing since 28 August. These rows keep them apart, each with its own interval.

This split was decided and published before a single match in the new group had been graded — before one had been sealed. The proof is in the ledger, which nobody can rewrite: the last entry sealed before the decision is 2026-09-17, and no entry on or before that date names any of the twelve. Group membership is frozen. A division added later joins neither group and gets its own line, so neither of these can ever be reshaped into the one that reads better.

Group Divisions Graded ProofOdds Closing line Gap Pending
The original eleven
Priced since 28 August 2026. Top divisions plus the Championship — the most heavily traded football in the world.
11 425 1.0073 0.9732 +0.0341 ±0.018 208
The twelve added on 17 September 2026
Second tiers, the English and Scottish lower divisions and two more top flights. Thinner markets, which is why they were added, and why they are counted separately.
12 174 1.0377 0.9924 +0.0452 ±0.027 97

Which division sits in which group is listed on the method page and set in config.py, in the repository, where the change is dated like everything else here. A group with no graded matches says so rather than borrowing the other group's number.

Division by division

The same question asked 23 times. A model that is genuinely close to the market should be close in most of them; an edge that lives in one division and nowhere else is usually a lucky autumn, not an edge. A division's log loss appears here only once it has 20 graded matches. Below that the gap's own margin of error is several times the gap, so the number would say nothing except which way a coin landed — and the division where it landed our way is the one that would get quoted.

Division Graded ProofOdds Closing line Gap Pending
E0 Premier League · original eleven 40 1.0434 1.0765 -0.0331 ±0.069 20
E1 Championship · original eleven 71 1.1146 1.0148 +0.0998 ±0.053 37
SP1 La Liga · original eleven 49 0.9406 0.9198 +0.0207 ±0.049 21
I1 Serie A · original eleven 40 0.9550 0.9393 +0.0156 ±0.047 20
D1 Bundesliga · original eleven 36 0.9853 0.9562 +0.0292 ±0.063 18
F1 Ligue 1 · original eleven 36 1.0685 1.0385 +0.0300 ±0.051 18
P1 Primeira Liga · original eleven 38 0.9545 0.9355 +0.0190 ±0.048 10
N1 Eredivisie · original eleven 38 1.0241 0.9839 +0.0401 ±0.077 19
B1 Jupiler Pro League · original eleven 26 0.9203 0.8624 +0.0579 ±0.069 8
SC0 Scottish Premiership · original eleven 18 too few to score — needs 20 6
BRA Brasileirao Serie A · original eleven 33 0.9542 0.9368 +0.0173 ±0.054 31
E2 League One · added 17 Sep 2026 18 too few to score — needs 20 12
E3 League Two · added 17 Sep 2026 30 1.1113 1.0555 +0.0558 ±0.055 12
EC National League · added 17 Sep 2026 36 0.9668 0.9557 +0.0111 ±0.048 11
SC1 Scottish Championship · added 17 Sep 2026 7 too few to score — needs 20 5
SC2 Scottish League One · added 17 Sep 2026 10 too few to score — needs 20 5
SC3 Scottish League Two · added 17 Sep 2026 10 too few to score — needs 20 5
D2 2. Bundesliga · added 17 Sep 2026 7 too few to score — needs 20 7
I2 Serie B · added 17 Sep 2026 9 too few to score — needs 20 9
SP2 Segunda Division · added 17 Sep 2026 29 1.0060 0.9985 +0.0075 ±0.049 12
F2 Ligue 2 · added 17 Sep 2026 3 too few to score — needs 20 4
T1 Super Lig · added 17 Sep 2026 8 too few to score — needs 20 8
G1 Super League Greece · added 17 Sep 2026 7 too few to score — needs 20 7

Per-division rows are small samples for a long time, and even the ones shown carry an interval wider than the gap they report. The headline at the top of this page is the one with the most matches behind it, and it needs a season too. "Pending" counts sealed predictions still waiting on a result or a closing price.

Week by week

Week of Matches ProofOdds Closing line Gap ± 95%
2026-10-05 8 0.9026 0.9256 -0.0230 ±0.055
2026-09-28 47 0.9746 0.9374 +0.0372 ±0.045
2026-09-21 35 1.1056 1.0819 +0.0237 ±0.046
2026-09-14 212 1.0302 0.9899 +0.0403 ±0.025
2026-09-07 103 1.0313 0.9688 +0.0625 ±0.043
2026-08-31 123 0.9991 0.9717 +0.0274 ±0.031
2026-08-24 71 0.9778 0.9546 +0.0232 ±0.046

Coverage

What is missing from the table above

Every other table on this page starts from the ledger and asks how the predictions did. This one starts from the matches that were actually played and asks whether the ledger has them at all — the only direction that can see a match we never sealed. A prediction that does not exist cannot appear in a table of predictions, so without this section a division whose fixture feed goes dark simply contributes fewer rows and the headline never flinches. Counted from the same results files the grading uses, against the same ledger, joined the same way.

599 of 649 matches played since each division entered the chain were sealed before kickoff — 92.3%. 38 were never sealed at all, and never can be: the ledger is append-only and a prediction published after kickoff is worse than no prediction.A further 12 were sealed but carry a club name that does not yet join to the results file — those are scored retroactively the moment the spelling is added, without anything being rewritten.

Division Fixtures from In the chain since Played Sealed Never sealed Name unjoined Coverage
E0 Premier League football-data.org 2026-08-27 40 40 — — 100%
E1 Championship football-data.org 2026-08-28 71 71 — — 100%
E2 League One football-data.co.uk 2026-09-18 18 18 — — 100%
E3 League Two football-data.co.uk 2026-09-18 30 30 — — 100%
EC National League football-data.co.uk 2026-09-18 47 36 11 — 77%
SP1 La Liga football-data.org 2026-08-28 49 49 — — 100%
SP2 Segunda Division football-data.co.uk 2026-09-18 33 29 3 1 88%
I1 Serie A football-data.org 2026-08-28 40 40 — — 100%
I2 Serie B football-data.co.uk 2026-09-18 10 9 1 — 90%
D1 Bundesliga football-data.org 2026-08-28 36 36 — — 100%
D2 2. Bundesliga football-data.co.uk 2026-09-18 9 7 2 — 78%
F1 Ligue 1 football-data.org 2026-08-28 36 36 — — 100%
F2 Ligue 2 football-data.co.uk 2026-09-18 9 3 6 — 33%
P1 Primeira Liga football-data.org 2026-08-28 38 38 — — 100%
N1 Eredivisie football-data.org 2026-08-28 39 38 — 1 97%
B1 Jupiler Pro League football-data.co.uk 2026-09-03 29 26 3 — 90%
SC0 Scottish Premiership football-data.co.uk 2026-09-02 23 18 5 — 78%
SC1 Scottish Championship football-data.co.uk 2026-09-18 12 7 5 — 58%
SC2 Scottish League One football-data.co.uk 2026-09-18 10 10 — — 100%
SC3 Scottish League Two football-data.co.uk 2026-09-18 10 10 — — 100%
T1 Super Lig football-data.co.uk 2026-09-18 9 8 1 — 89%
G1 Super League Greece football-data.co.uk 2026-09-18 7 7 — — 100%
BRA Brasileirao Serie A football-data.org 2026-09-02 44 33 1 10 75%

“In the chain since” is the first date the ledger can prove the division was being asked about. Before 6 September an entry recorded only the divisions that sealed, not the ones that were asked and answered nothing, so for the older divisions that date is a lower bound — these percentages can therefore only flatter us, never the reverse.

Three counts, and they are not the same number

23 divisions are configured. 23 have had a match to seal since entering the chain. 23 have a sealed prediction in the ledger, and 23 appear on the scorecard above. A single number covering all four would be the easiest thing on this page to get quietly wrong.

Why a missing match is worse than a wrong one

A wrong prediction is on the scorecard, costing us. A match we never sealed is on nobody’s scorecard — it neither helps nor hurts the headline, which is exactly what makes it dangerous. If the matches we miss are not a random sample of the matches played, the gap above is measured on a subset somebody else selected. That is why this table exists and why it is on this page rather than a status page nobody reads.

New markets

BTTS and Asian handicap

BTTS forecast only

0.6891 log loss over 419 matches; 0.6931 is a coin flip. There is no free closing BTTS benchmark in this source, so this is forecast validation—not an edge claim. At 419 matches the distance from the coin flip is -0.0041 ± 0.0165, which does not clear zero.

Asian handicap scored vs close

Model 0.7195 vs close 0.6940, 384 matches. Quarter-line half-wins and half-losses are stake-weighted. The gap is +0.0256 ± 0.0240; at 384 matches that is a separation.

What this page will never do

No ROI headline

Return on investment over a few hundred bets is mostly noise, and it is the number every tipster picks because it is the easiest to flatter. Log loss cannot be cherry-picked.

No hidden window

Every graded match since the first publication is in the numbers above. There is no "since we improved the model" start date.

No quiet edits

The ledger is a hash chain in a public repository. If a past prediction changed, this site would say so itself. What a chain cannot prove on its own is that nothing was left out; the ledger page says so rather than letting you assume otherwise.