ProofOdds

Method

How this is built and how it is measured

Short version: a Dixon-Coles goals model, refitted every day on matches played before that day, published before kickoff, and scored with log loss against the market-average closing line. The code is public.

Two kinds of number, and the difference matters

This site publishes more than it can grade. That is a deliberate change from how it started, and the honest way to do it is to label every figure with which kind it is — on the card itself, not in small print underneath.

Scored against the close

A public closing price exists for this exact selection, so the forecast is compared with the market by log loss and can be shown to lose. This is the scorecard, and it is the only place an edge claim could ever live. Today: the result in every division, plus over/under 2.5 and the Asian handicap in the divisions whose source publishes them. Those three are not three independent measurements, and the scorecard should not be read as if they were — why, and by how much.

Scored against guessing

No free closing benchmark exists anywhere for these, so there is nothing to be measured against except guessing — log loss against the coin flip, which says whether a forecast is calibrated, not whether it beats anyone. Sealed and scored on that basis, never presented as an edge. Today that is both teams to score, and only that: its number is on the scorecard.

Sealed, not yet scored

Published on the match cards, sealed into the ledger before kickoff, and carrying no score anywhere on this site: the goal-total lines other than 2.5, the correct-score view, corners in the divisions that qualify, and the markets derived from the same scoreline distribution — clean sheet, win to nil, odd/even, exact total goals, goals by team and the winning margin. They are shown as context, they are falsifiable later because they are sealed now, and until a number exists they are evidence of nothing.

Why there are three boxes and not two Until September this page had two, and it listed the goal-total ladder, the correct-score view and a corners model in the second one under the words "sealed and scored on that basis". Only BTTS was scored. Nothing else on the site produced a number for the other three, and a site whose whole argument is that it keeps the score it promises to keep had promised a score it was not keeping. The third box is that correction. The corners model came out altogether while that was being sorted out; it is back on the match cards as of 18 September 2026, in the third box where it belongs, with no separate page and no navigation link of its own — see corners below for which divisions show it and why the others do not.

The earlier version of this page said the site would publish nothing without a closing line at all. We changed our mind on that and we would rather say so than quietly widen the definition: a forecast measured against guessing is still falsifiable, still sealed before kickoff, and still ours to be wrong about in public — it is simply a weaker claim than beating the market, and it is labelled as one everywhere it appears. What has not changed is the rule underneath: nothing is published that cannot be checked against something.

Which bucket a market falls in depends on the division, not just the market. The season-by-season European files carry a closing 1X2, over/under 2.5 and one main Asian-handicap line. The "new leagues" files — the Brasileirão among them — carry a closing 1X2 and nothing else. So for Brazil only the result is scored against the close; its over/under and handicap sit in the third box, tagged sealed, not scored on every match card.

A correction, 18 September 2026. Those two markets used to be tagged forecast only on Brazilian cards. That tag means measured against guessing, and it was not true: grade.py needs a closing price to mark an over/under or a handicap as graded at all, so a Brazilian over/under reached no scorecard on this site and was measured against nothing. The second box is for markets something actually scores, and only both-teams-to-score qualifies. The tag now says sealed, not scored, which is what it always was.

The model

Every match gets the result, BTTS, a 0.5–5.5 goal-total ladder, a quarter-line Asian-handicap grid, the correct-score view, clean sheets, win to nil, odd/even, exact total goals, goals by team and the winning margin. All of them are different sums of the same Dixon–Coles scoreline distribution — not separate fitted models, and not a second set of assumptions. Corners are the one exception on the card and are a genuinely separate model; they are marked as such below.

Three scorecard lines, and not three measurements

The paragraph above is the pleasant half of the sentence. This is the other half, and it changes how the scorecard should be read.

If every market is a different sum of one fitted distribution, then the result, over/under 2.5 and Asian handicap lines on the scorecard are not three independent verdicts on the model. They are three partitions of the same distribution, graded on overlapping sets of the same matches. Every one of the 384 matches graded on the handicap, and every one of the 555 graded on the total, is a match the result market also grades. A reader who sees the model behind on all three — and as this page built, it is behind on all three will naturally count three pieces of evidence. It is closer to one piece of evidence seen three ways. That reader could easily be us, which is the reason this is on the page.

Correlated is not the same as identical, and the overstatement runs both ways. The markets do ask different questions of the same distribution. Over/under asks about the total, the result and the handicap about the difference, and a large handicap line asks about the tail of that difference — a margin the result market never has to resolve, because two goals and five goals are the same home win to it.

How much overlap, measured rather than asserted. Each scorecard headline is a mean of one number per match: the model's log loss minus the closing line's. Correlate that per-match number between two markets, over the matches both of them grade, and you have how far the two headlines are one number. At build time, on the live graded sample:

  • Result and Asian handicap: r = 0.70 over 353 matches. Much of the same fact twice.
  • And it splits by line size, which is the point. On handicaps of 0.25 or less — where the handicap is very nearly draw-no-bet, or "home win or not" — r = 0.89 over 146 matches: all but the same measurement. On handicaps of 1.00 or more, r = 0.45 over 78 matches. That gap between the two is the margin tail being genuinely tested, and it is the honest case for keeping the handicap line on the page at all.
  • Result and over/under 2.5: r = 0.15 over 555 matches. Close to independent, which is what you would expect: knowing the total tells you little about who won.

So the scorecard's three scored lines are worth somewhere between one measurement and three, nearer two, and the handicap is the line that adds least of its own. Read the result line as the headline it is, read the total as a second and largely separate question, and do not read the handicap as a third opinion on the first — at small lines it mostly is the first. The coefficients above are Pearson on a paired difference, computed at build time from the same graded frame the scorecard uses, and they will move as the sample grows.

Not every entry carries every market, and none was backfilled. The markets arrived on different days and first publication wins, so a fixture sealed before a market existed carries only what was sealed that day and its card says so. Expected goals go back to 26 August 2026, over/under 2.5 to 28 August, and the full goal-total ladder, the Asian-handicap grid, both-teams-to-score and corners all start on 2 September. The 630 predictions sealed between 28 August and 1 September show the 2.5 line alone, because that is the only goals line they sealed — filling in the other five now would mean rewriting a sealed entry, which this site does not do for any reason, including to make a page look tidier.

Sealed as numbers, and derived from sealed numbers

Both kinds of figure are on every match card and a reader has to be able to tell them apart, so the card says which is which and so does this page.

  • Sealed directly. Written into the ledger entry before kickoff and never rewritten: the result, all 6 goal-total lines, both teams to score, all 25 Asian handicap lines, xg_home, xg_away, and the corner distribution where a division has one. Open the entry and they are there.
  • Derived from sealed inputs. Recomputed when the page is built and stored nowhere: the correct-score grid, clean sheet, win to nil, odd/even, exact total goals, goals by team and the winning margin. Each is a different sum of one Dixon–Coles distribution rebuilt from three numbers that are sealed — xg_home, xg_away and the division’s rho — at 10 goals a side.

This is not a loophole in "sealed before kickoff". A deterministic function of three sealed numbers is fixed the moment those numbers are: there is no freedom left in it, nothing is fetched at build time, and anyone with the entry can recompute every one of these figures and get the same answer.

The check is easy to run and worth running. Across all 413 sealed predictions that carry the inputs, the derived figures reproduce the sealed ones — the result, the goal-total ladder, both teams to score — to within 3 parts in 100,000. That residual is not slack in the method: the entry stores xg_home, xg_away and rho rounded to four decimal places, and rebuilding a grid from a value that may be 5×10⁻⁵ off moves a probability by about that much. Truncating the grid at 10 goals a side costs 5×10⁻⁷ by comparison. If you recompute and land further away than that, something is wrong and we would like to hear about it.

They are derived rather than sealed for one practical reason: the entry for a single day is already 625 KB for 88 predictions, and sealing a dozen more partitions per fixture would multiply that for numbers that carry no information the three sealed inputs do not already carry. The correct-score view has worked this way since it went up; these follow the same precedent rather than inventing a third pattern.

Every one of the derived markets is printed as a partition: a set of outcomes that is exhaustive and mutually exclusive, adding to one by construction, with nothing left in an "any other result" line. They are rounded in a single pass in tenths of a percent, so the printed figures add to exactly 100.0 — and where one outcome falls below a tenth of a percent it is printed as the floor <0.1% rather than as 0.0%, because the model did not rule it out, in which case the printed figures no longer add to 100.0 and the card says so under that market instead of claiming otherwise. The cut points are chosen the same way: exact total goals prints 0 to 7 and one "8 or more" bucket, goals by team prints 0 to 5 to match the axes of the correct-score grid, and the buckets mean nothing is dropped either way.

One model is fitted per division, on that division's matches only. Ratings are not comparable across countries — a Championship attack rating of 1.1 says nothing about a Bundesliga one — which is also why this site does not price European competitions: there is no common scale on which to compare the two teams.

Every team carries an attack rating and a defence rating, and there is one league-wide home advantage. Expected goals for a fixture are

λ_home = α_home × β_away × γ × league_mean
λ_away = α_away × β_home × league_mean

Goals are then modelled as Poisson counts. Two Poissons multiply into a grid of scorelines; below the diagonal is a home win, the diagonal is a draw, above it an away win. Nothing in the model knows anything about football — the ratings are chosen by maximum penalised likelihood, which is to say the numbers are pushed around until the observed results stop being surprising.

Dixon and Coles (1997) add two things to that, and both matter:

  • A low-score correction. Independent Poissons say the two teams' goals are unrelated. They are not: 0-0 and 1-1 happen more often than the maths predicts and 1-0 and 0-1 less often, because a level game is played differently from a decided one. Four cells of the grid get a multiplier.
  • Exponential time decay. A match from 2016 says nothing about this season. Each match enters with weight exp(−ξ · days_ago). We use ξ = 0.002 per day, a half-life of 347 days.

The rule that makes the numbers mean anything

To price a match on date D, the model may only see matches played strictly before D. No exceptions. Fit on everything and then "predict" the past and the results come out spectacular and worthless — that is lookahead bias, and it invalidates most amateur backtests you will ever be shown.

Here it is enforced structurally rather than by care: matches are sorted by date once and the training slice is taken by binary search, so a future match physically cannot enter a fit. The same discipline applies to the two hyperparameters — the decay rate and the prior width were chosen on 2017/18–2020/21 and then frozen, because picking them on the data you report is a slower form of the same bias.

Which matches make it in — and which do not

The scorecard measures the predictions we published. It is silent, by construction, about matches we never published a prediction for: a forecast that does not exist cannot appear in a table of forecasts. So a division whose fixture feed goes quiet does not show up as a problem — it just contributes fewer rows, and the headline gap does not move. For a site whose entire claim is that the bad numbers stay up, that is the most dangerous blind spot available, because it is the one that flatters by omission.

So the coverage is measured and published. For every division, from the results files themselves: how many matches have been played since that division entered the ledger, and how many of them we sealed a prediction for before kickoff. Currently 599 of 649, or 92.3%. The per-division breakdown is on the scorecard, beside the numbers it qualifies.

Two different things can put a played match outside the ledger, and they are counted separately because only one of them is permanent:

  • Never sealed. The fixture feed did not offer the match in time. This cannot be repaired — the ledger is append-only, and a prediction added after kickoff would be worthless and dishonest. The match is gone from the record for good.
  • Sealed, but the club name does not join. The prediction exists and was published on time; it simply cannot yet be matched to the result, because the fixture feed and the results file spell a club differently. Adding the spelling scores it retroactively, without any entry being altered.

Fourteen of the 23 configured divisions take their fixtures from football-data.co.uk’s shared fixtures.csv, a file published when its author publishes it. When it goes stale, those divisions seal only the matches it happened to carry. That is a real limitation on the scorecard and not a temporary one, so it is stated here rather than discovered: for those divisions the score is computed over the matches the feed offered, and the coverage column says how many that was.

Being explicit about the direction of the risk: if the matches we miss are not a random sample of the matches played — and a feed that posts in batches gives no reason to think they are — then the gap on the scorecard is measured on a subset selected by somebody else. We cannot correct for that from inside the feed. We can report it, which is what the coverage column does.

How the score is computed

Log loss: take the probability we gave to the outcome that actually happened, apply −log(p), average over matches. Say 90% and be right, you pay 0.105. Say 10% and be right, you pay 2.303. Lower is better, and unlike ROI it cannot be flattered by picking which bets to count.

Closing odds are turned into probabilities by inverting the three prices and dividing each by their total, which removes the bookmaker's margin proportionally. Two reference points frame every number on this site: predicting 1/3 each time scores 1.0986, and the closing line scores 0.9639 over the Premier League backtest below. The whole of football knowledge is the gap between those two. That second figure moves with the division and the sample: across the 23 divisions with graded matches on the live scorecard the closing line is currently scoring something different again, which is why every page here quotes the number belonging to the sample it is talking about rather than one house figure.

The benchmark — and why it changed once

The line we grade against is the market-average closing price as published by football-data.co.uk (the AvgC columns): the mean across the books they survey, taken at kickoff, de-vigged proportionally like everything else here.

Until January 2026 the site graded against Pinnacle's closing price, the usual sharp-book benchmark. football-data.co.uk stopped carrying Pinnacle's columns mid-January 2026, in every division at once, which left a choice: blend benchmarks match by match, or switch to one line that exists for every match. Blending is how numbers stop meaning anything, so we switched — everywhere, including the backtest below, which was recomputed against the average from scratch.

Before switching, we measured what the change does. On every match since 2019/20 that carries both prices, in the eight original divisions in this benchmark comparison, the two de-vigged benchmarks were scored against each other:

  • Premier League (2,490 matches): Pinnacle 0.9584 vs average 0.9582 on the result market — a difference of 0.0001 in log loss. On the totals market, 0.6729 vs 0.6732. The two probability series correlate at 0.9995.
  • All eight original divisions: the largest difference anywhere is the Portuguese Primeira Liga at +0.0020 (the average is a slightly softer benchmark there); the Eredivisie reads +0.0013 and every other division ±0.0007 or less. On totals, no division differs by more than 0.0004. No single season in any division differs by more than 0.005.

So the switch moves the goalposts by at most two thousandths of a nat, in a game where the model's distance to the market is about twenty. Where the bias exists it makes the benchmark slightly easier, and this paragraph is the disclosure. The measurement is reproducible: scripts/check_benchmark.py in the site repository reruns it from the raw CSVs. Pinnacle's prices are still recorded in the data where they exist, as a cross-check, and are never graded against.

Two groups on the scorecard, and when that was decided

The site began on 28 August 2026 pricing eleven divisions. On 17 September 2026 twelve more were added, all from the same source. They were added because they are thinner markets — a closing line ought to be easier to match where fewer people are pricing it, and that is a claim worth testing rather than asserting.

It is also a claim that would quietly corrupt the headline. A pooled log loss moves when the mix of divisions moves, so from the first new-division match onwards the pooled figure is a composition change and a performance change added together. The scorecard therefore keeps the pooled number — it is the honest total, and quietly dropping it would be its own kind of selection — and shows these two groups beside it, each with its own interval.

The timing is part of the claim. A split introduced after results arrive cannot be told apart from a split chosen because of them. This one was written, committed and published before a single match in the new group had been graded, and before one had been sealed. The evidence is not our word for it: the ledger is append-only and public, the last entry sealed before the decision is 2026-09-17, and no entry dated on or before it names any of the twelve. Anyone can check that with proofodds/verify.py and a clone.

Membership is frozen, and that matters more than the split itself. These are not buckets for “big” and “small”; they record what was already being measured on 17 September 2026. A division added in future joins neither group and gets its own line, so neither of these can be reshaped later into whichever reads better.

GroupDivisions
The original eleven
Priced since 28 August 2026. Top divisions plus the Championship — the most heavily traded football in the world.
E0 Premier League, E1 Championship, SP1 La Liga, I1 Serie A, D1 Bundesliga, F1 Ligue 1, P1 Primeira Liga, N1 Eredivisie, B1 Jupiler Pro League, SC0 Scottish Premiership, BRA Brasileirao Serie A
The twelve added on 17 September 2026
Second tiers, the English and Scottish lower divisions and two more top flights. Thinner markets, which is why they were added, and why they are counted separately.
E2 League One, E3 League Two, EC National League, SC1 Scottish Championship, SC2 Scottish League One, SC3 Scottish League Two, D2 2. Bundesliga, I2 Serie B, SP2 Segunda Division, F2 Ligue 2, T1 Super Lig, G1 Super League Greece

The backtest — and why it is on this page, not the scorecard

Before a single live prediction existed, this exact model was run walk-forward over nine seasons of the Premier League. The average closing price is published from 2019/20, so the scored window is the last seven of them — the first two train the model without being graded. That result is a useful prior and it is not a track record, so it lives here and never appears next to the live numbers.

What this backtest does not cover It is the Premier League only, and so are the two hyperparameters: the decay rate and the prior width were chosen on E0 seasons 2017/18–2020/21 and then frozen. The other 22 divisions that have published a forecast run the same model with those settings transferred, and have no walk-forward history of their own. Most of the live scorecard is therefore made of divisions this backtest says nothing about — the Championship alone supplies more of the graded sample than the Premier League does. Treat the figures below as a prior for the model's shape, not as a forecast of what it will do in Belgium or Brazil. Fitting and validating a walk-forward per division is the obvious next piece of work and it has not been done.
Backtest, model
0.9827
2660 matches, 2019/20 – 2025/26
Backtest, closing line
0.9639
same matches
Gap
+0.0188
the model loses, per match
Held-out slice
+0.0227
1900 matches that informed no decision
We do not beat the closing line, and we say so on the front page A goals-only model loses to a market that also knows about injuries, lineups, rest days and dead rubbers. Publishing that is the point. Anyone selling you certainty is either not measuring, or is measuring and not showing you.

Corners — a separate model, and a rule about where it appears

Corners are the one market on a match card that is not a sum of the goals distribution. They have their own model: a team-strength negative binomial fitted only on this division’s final home and away corner counts (HC, AC) from football-data.co.uk. It ingests no odds, because there is no free closing corner price anywhere, in any division. Corners are therefore sealed, not scored everywhere they appear and are never an edge claim.

What was wrong with it until 18 September 2026 The corner model applied no time weighting at all, in any division, while the goals model decayed its history at ξ = 0.002 per day. A corner count from 2015/16 counted exactly as much as one from last week. It now decays at the same ξ, in the fit, in the starting level and in the dispersion estimate. Measured before and after across every division with corner data, expected total corners move by 0.3 to 0.8 per fixture on average. That is why corners stayed off the cards until the weighting was in.

The ledger is append-only, so the corner blocks sealed between 2 September and 18 September 2026 are unweighted and stay exactly as they were sealed. A match card shows a corner section only when the block it is reading came from the weighted fit; older fixtures show no corner section and say why. Nothing was rewritten to make this page true.

The rule

A corner model is fitted, sealed and shown in a division only when

the median club in the current season carries at least (number of clubs − 1) time-weighted corner appearances, decayed at the same ξ = 0.002 per day.

A club plays twice that many matches in a season, so the bar is half a season of recent corner history for the typical club, and it scales with the division because a season does. It replaced a flat threshold of 100 raw rows, which could not see either thing that matters. It could not see a hole in the coverage: the National League’s file carries corner counts in 2015/16 and again in 2026/27 and in none of the ten seasons between, so 648 rows sailed past a bar of 100 while 552 of them described a division that no longer exists. And it could not see division size: 100 rows is a fifth of a Championship season and most of a Scottish League Two one.

It is deliberately a rule and not a list of leagues. It is measured the same way everywhere, it is re-measured on every run, and a division whose source stops publishing corner counts drops out on its own — no one has to notice.

What it passes today

Measured at build time from the same data the ledger gates on, so this table is the rule’s own output rather than a list anyone typed.

Division Rows with corners Effective Newest season Median per club Bar Corners shown
Premier League 4230 472.5 2026/27 47.1 19 yes
Championship 6166 693.2 2026/27 49.8 23 yes
League One 6010 689.1 2026/27 42.7 23 yes
League Two 6062 699.2 2026/27 47.3 23 yes
National League 695 133.9 2026/27 11.2 23 no
La Liga 4249 493.2 2026/27 49.3 19 yes
Segunda Division 4225 602.3 2026/27 39.8 21 yes
Serie A 4230 475.9 2026/27 47.6 19 yes
Serie B 3512 469.5 2026/27 27.6 19 yes
Bundesliga 3401 378.1 2026/27 42.0 17 yes
2. Bundesliga 2808 388.5 2026/27 36.2 17 yes
Ligue 1 3901 392.2 2026/27 42.7 17 yes
Ligue 2 3230 410.7 2026/27 33.9 17 yes
Primeira Liga 2816 399.6 2026/27 43.4 17 yes
Eredivisie 2743 396.2 2026/27 43.5 17 yes
Jupiler Pro League 2622 398.5 2026/27 49.2 17 yes
Scottish Premiership 2501 291.0 2026/27 48.1 11 yes
Scottish Championship 1574 234.0 2026/27 44.0 9 yes
Scottish League One 1552 235.2 2026/27 32.5 9 yes
Scottish League Two 1551 237.5 2026/27 43.1 9 yes
Super Lig 3111 416.3 2026/27 43.2 17 yes
Super League Greece 2180 298.5 2026/27 41.4 13 yes
Brasileirao Serie A 0 0.0 none 0.0 19 no

The two that fail fail for different reasons. The Brasileirão file carries no corner columns at all, so there is nothing to fit. The National League has the columns but not recently enough: its median club carries about 8 effective appearances against a bar of 23, because almost all of its corner history is a decade old. A division that fails stops sealing a corner block rather than sealing one nobody should trust — and the blocks it sealed in the past stay in the ledger untouched, because nothing here is ever rewritten.

What this site will not do

  • Sell picks, "sure things" or accumulator tips.
  • Put an "AI" badge on a number without the score beside it.
  • Quote a return on investment as evidence of skill.
  • Start the record over after a bad run.

Limitations, plainly

  • Goals only. No shot data, no expected goals, no lineups. The model has never heard of a suspension.
  • Promoted teams start at league average and shrink towards their real level as matches arrive. Their first few weeks are the model's worst.
  • De-vigging is proportional, the simplest method. It slightly overstates longshot probabilities, which if anything flatters us.
  • Closing odds are not available for every match. Matches without them stay out of the score rather than being quietly counted.

Where these numbers come from. The backtest above is not a claim you have to take on faith: the research repository holds the model, the walk-forward harness, the hyperparameter search and the tests, along with the raw CSVs, so python run.py reproduces every figure on this page — https://github.com/GitSimaao/pl-dixon-coles. The code running this site is at https://github.com/GitSimaao/proofodds. Results and closing prices come from football-data.co.uk.