Track record

How good a prediction model is isn't claimed, it's measured. This page shows the model's past performance unfiltered — unfiltered.

These figures come from backtesting, not from live tracking. For each past match the model was trained only on matches played before it, then tested. Once live tracking is running this page will update with real-time data. Scope: 16 leagues, 795 matches, 12,344 predictions. These numbers are computed from the database every time the page loads; there is no hand-updated copy. (July 2026September 2026)

Last updated: 16 September 2026 at 19:10 · Methodology version: 10

The model knows when it knows

The real test of a prediction model is not that it answers every match, but that it is right on the ones it is sure about. Below are the predictions where the model said “I am at least this confident”, and how many of them actually turned out correct. The thresholds were not cherry-picked; every tier is here.

What the model saidPredictionsCorrect
Match result “at least 50% confident”37058.9%
Match result “at least 60% confident”18073.9%
Match result “at least 70% confident”7482.4%
Double chance “at least 75% confident”74980.1%
Double chance “at least 80% confident”35286.6%
Double chance “at least 85% confident”17791.5%

The thing to notice: the stated percentage and the actual percentage track each other, and the actual is slightly higher. The model does not overstate. If you see “95% accuracy” somewhere, the question to ask is: over how many predictions, in which market, and were the incorrect ones counted.

If we had to answer every match

The harshest measure: the case where the model is forced to tick one of three boxes on every single match. Nobody plays that way, but we have nothing to hide — what it is compared against and where the ceiling sits are written next to it.

74.2%

The ceiling of this game. 25.8% of matches end in a draw, and a draw is almost never on its own the single most likely outcome — nobody forced to tick one of three boxes, us or anyone else, gets meaningfully above this.

80.1%

of the 749 double-chance predictions we called “at least 75%” were correct (600 of them).

The double-chance box is shown on the home page too. If it had come out the other way — delivering less than we claimed — it would sit in the same place at the same size. The showcase exists to show what was measured, not to pick the good news.

Does it work in other leagues too

A model doing well in a single league can also be explained by it having fitted that league's quirks. The real test is producing the same result in other leagues with no tuning at all. The parameters were chosen on the Süper Lig; the other leagues below took no part in that choice.

LeagueMatchesOur accuracyvs random
EredivisieNetherlands5455.6%+22.2 pts
Serie AItaly4060.0%+26.7 pts
Ligue 1France3641.7%+8.3 pts
UEFA Europa Conference League7052.9%+19.5 pts
Premier LeagueEngland4045.0%+11.7 pts
PremiershipScotland3652.8%+19.4 pts
Ligue 2France5446.3%+13.0 pts
La LigaSpain5653.6%+20.2 pts
1. LigTurkey7045.7%+12.4 pts
2. BundesligaGermany4546.7%+13.3 pts
ChampionshipEngland8345.8%+12.4 pts
UEFA Champions League4353.5%+20.2 pts
BundesligaGermany2755.6%+22.2 pts
Segunda DivisiónSpain5545.5%+12.1 pts
Süper LigTurkey4546.7%+13.3 pts
UEFA Europa League4153.7%+20.3 pts

In 15 of 16 leagues the model beats random ticking by 8.3 to 26.7 points — 795 matches in total. In every one of those leagues it also beats that league's own historical average. In one or more leagues we came in below the base rate; that row is not hidden.

What this table proves, and what it does not

The accuracy figures here are the numbers under a rule of “tick one of three boxes on every single match”. Nobody has to play that way; the confidence table above shows what the product actually does. This table's job is not to persuade — it is to show that the METHOD WORKS IN EVERY LEAGUE.

The reference point is ticking one of three options at random: 33.3%. Not a chosen number but the definition itself — three options, one is right. The difference is what the model genuinely adds, and it is written in points: the difference between two percentages is points, not a percentage.

A harder comparison is measured too: each league's own historical average — a forecaster that already knows home advantage. The model beats that in every league as well; the audit and the track record on this page rest on that measure.

The numbers in context: ticking at random gives 33%. The upper bound is 74.2% — because 25.8% of matches are draws and nobody can foresee a draw. The room to move is narrower than it looks, and that is a fact about football.

So wherever you see a high accuracy figure, ask the same question: on which matches, in which market, and were the incorrect ones counted too. On this page they all are.

This table covers only the leagues published on the site; leagues ingested for measurement but not put on screen do not enter it either — the track record's cohort and the site's cohort are the same one. The second-tier leagues sit clearly below the first-tier ones; that row is not hidden, it is in the table.

Calibration curve

The answer to “of the matches we called 60%, did 60% actually happen?” In a well-calibrated model the two columns are close together.

Confidence bandModel saidActually happenedMatches
0%10%5.8%0.0%11
10%20%15.8%21.2%52
20%30%25.7%21.1%90
30%40%35.5%28.6%161
40%50%44.7%42.7%206
50%60%54.0%41.4%140
60%70%64.4%64.9%74
70%80%75.3%82.1%39
80%90%85.1%84.2%19

Being watched: under-rested home teams

When the home team has had 4 days or less since its last match, the model used to systematically underrate it. In the latest measurement that bias is gone: we said 44.4%, the actual figure was 44.6% (3,129 matches).

The flaw was neither invented nor buried: it was genuinely measured on 8 August 2026 across 719 matches in 6 leagues (z=+3.14) and has been stated on this page ever since. As coverage grew and the sample multiplied, the gap dissolved statistically — meaning the original finding was largely small-sample noise. This box updates itself after every audit run; if the bias returns it will show up here as a warning again.

Largest deviation: in the 50-60% band the model said “54.0%” on average, and 41.4% actually happened. So in that band it trusts itself too much — the error is on the optimistic side, which is more serious. It stands here as a shortcoming to be fixed; it isn't hidden.

Other markets

Every market is measured separately against its own baseline. Instead of saying “the model is good already, so this market must be good too”, each one was tested. The n column shows how many matches each row rests on — lower Brier is better.

This table does not by itself decide which market gets published. That decision came from a separate walk-forward test: the coefficient is chosen on the first 75% and tested on the final 25% holdout. A result that looks good on all the data but collapses on the holdout is data dredging, not a real edge. The first-half goal markets failed that test and are not published.

PazarnModel BrierTaban oranEdge
Double chance 127950.38340.3882+1.2%
Double chance 1X7950.39050.4458+12.4%
Double chance X27950.42490.4836+12.2%
Expected fouls — away6300.45830.4866+5.8%
Expected fouls — home6300.46170.4806+3.9%
First half — level6410.47230.4734+0.2%
First-half double chance 126410.47230.4734+0.2%
First-half double chance 1X6410.40240.4126+2.5%
First-half double chance X26410.42270.4318+2.1%
First half — away ahead6410.40240.4126+2.5%
First half — home ahead6410.42270.4318+2.1%
Both teams to score6410.48560.4915+1.2%
Expected corners — away5930.41690.4835+13.8%
Expected corners — home5930.45130.4925+8.4%
Match result — draw7950.38340.3882+1.2%
Match result — away win7950.39050.4458+12.4%
Match result — home win7950.42490.4836+12.2%
3+ goals6410.46230.4828+4.3%

Why double chance adds up to 200%

1X, 12 and X2 always sum to exactly 200%. This is not an error, it follows from the definition: the options are not mutually exclusive, they overlap — every outcome is counted in two of the three, so the 100% of 1X2 is doubled. For the same reason 1X plus X2 exceeds 100% by exactly the probability of a draw. This is an invariant, verified by a unit test on every build.

An honest note about the first half: the first-half model is trained separately (it is not a fraction of the full-match prediction), but its edge is markedly smaller than the full-match model's — the measured figure is in the table above. First-half goal markets are not published at all because they failed to beat the base rate.

Goals market (3+ goals)

In this market the model is not symmetric. The measurement showed:

So the goals market is only shown for matches where the model gives a clear signal in the “3+ goals” direction. For other matches, rather than inventing a number, it says “no clear signal”. Staying silent when you have nothing to say is more honest than filling the space.

Why the incorrect ones were wrong

Counting the correct predictions is easy. What actually teaches you something is what happened when we were wrong. All 400 incorrect predictions were examined one by one.

83.8%

Share of losses decided by a single goal.

52.5%

Share of losses caught out by a draw.

6.0%

Reverse result by 3+ goals — a genuine surprise.

The main culprit behind our losses: the draw

The model said “draw” in only 0.0% of 795 matches. Yet 26.4% of matches ended level.

This is not an error, it's what the maths gives: a draw is almost never the single most likely outcome — with probability spread across three results, one side usually stands out. But when draws do happen they are recorded as losses, and they make up half of ours. That is exactly why double chance performs so much better than 1X2: it covers the draw instead of losing to it.

Is there a parameter to be extracted from the losses

Factors knowable BEFORE the match were scanned: days of rest, fixture congestion, missing players, stage of the season, expected goals. Result: no usable parameter emerged.

One factor looked significant in the raw scan (“accuracy is 8.6 points higher in matches expected to be high-scoring”). Controlled for within the same confidence band, the sign reversed (−6.9 points) — high-scoring matches were already concentrated in the high-confidence bands, and those bands are naturally more accurate. The raw finding was an artefact of the distribution, not of the factor. Without that control we would have drawn a wrong parameter out of it and damaged the model.

Our ceiling: 84% of losses are decided by a single goal. That is football's natural variance; chasing it makes the model fit noise. The share attributable to red cards and in-match events will be measured separately once event data is ingested — that data can't be known before kick-off so it can't be a parameter, but it shows our ceiling.

Tried it, didn't add it

An idea sounding right is not enough. Approaches that were tried and found not to work are listed here too — because what we DIDN'T add tells you as much as what we did.

Head-to-head history (the “bogey team” effect)

The intuition that “team X has a bogey team in Y” was tested. Result: no signal.

  • The correlation between the bias in past meetings and the bias in the current match is r = −0.05 — indistinguishable from chance.
  • Of the 165 pairings with at least 4 matches, 52 look “dominant”. But even in a world with no pairing effect at all, simulation produces 46.3 of them. The excess of +5.7 sits inside the noise band.
  • When a correction was applied to expected goals, the best coefficient came out as zero — that is, applying no correction. Every positive coefficient made the prediction worse.

The honest limit: across the 3 seasons we hold, there are at most 6 matches per pairing. That is enough to say “this data can't show it”, not enough to say “there is no effect”. The test will be repeated when more seasons are available.

An attempt to correct the known bias (sharpening)

To fix the “not confident enough” flaw described above, the probability distribution was sharpened with a correction that preserves the ranking. Result: it didn't work and wasn't applied.

  • With the coefficient chosen on the first 75%, the holdout Brier went 0.6020 → 0.6031 (worse).log loss de bozuldu.
  • It improved only three of the six leagues and hurt three. A correction that doesn't work everywhere isn't a correction, it's overfitting.
  • The calibration error (ECE) got worse at the very thing we were trying to fix.

The bias itself turned out to be fragile: the +3.2 point gap seen for very strong away teams across all the data becomes −0.8 in the most recent 25%. So it isn't a stable enough flaw to correct. Another method will be tried; if that fails too, the flaw will keep standing here as it is.

Pulling the draw probability toward the base rate

Since half our losses are caught out by draws, this was tried: pull each match's draw probability toward the league's long-run draw rate. Result: no gain, it makes things worse.

The measurement showed why: the model's average draw probability is 25.6%, while the actual draw rate is 23.3%. So the model doesn't under-call draws — it slightly OVER-calls them. The impression that “the model doesn't see draws” is wrong — it gives each match the right probability, but among three outcomes a draw is almost never the highest (in 3 of 943 matches; the highest value seen was 37.8%). That is what the maths gives, not a flaw to be corrected. The right answer was not to force the model but to bring double chance forward.

Number of missing players (injuries)

Tested on 859 matches against real per-match injury records. Result: no measurable contribution.

  • The correlation between the difference in missing players and the bias in the result is r = −0.04. The direction is right — the depleted team scores less — but the magnitude doesn't separate from noise (p = 0.28).
  • Added to the model as a correction, the best coefficient again came out as zero.

Why this isn't surprising: we have no player-quality data, so a missing star and a missing substitute count the same. On top of that, the effect of ongoing injuries is already reflected in the team's recent results and therefore in its rating. This is not a “never”, it's a “not like this”.

Team-specific home advantage

League-wide home advantage is already in the model (home teams score 1.63 goals per match, away teams 1.28). What was tested is whether the advantage varies from team to team. Result: no contribution.

The team factors looked meaningful (highest 1.25, lowest 0.79) but adding them to the model made the prediction worse at every coefficient. There are only ~17 home matches per team per season; most of the “fortress at home” label that comes out of that little data is noise.

An independent strength measure (ClubElo)

Because squad-value data could not be obtained from a source with clean licensing, an independent team-strength measure was tried. Result: the effect is real but small (+0.46%).

What shows this isn't noise: when the coefficient was chosen on the first 75%, the best value came out as 0.12 — and the holdout's own optimum was also at 0.12. Under overfitting those two points diverge. Even so, for a gain of that size we don't take on a dependency with unclear licensing; if the terms become clear and it's confirmed across more seasons, it will be added.

Technical detail: Brier score and method

The measure behind the comparisons above. The Brier score is the mean squared error between the probability stated and the outcome that occurred; lower is better. Accuracy asks “did you tick the right box”, the Brier score asks “how right was the percentage you stated” — the second is stricter, and it is the one we use to correct the model.

MethodBrierDifference
Model0.5994
Baseline: historical base rate0.6588+9.0%

Each match's prediction was computed only from matches played BEFORE it (walk-forward validation), and the baseline used the same window — so the comparison is fair. The first 120 matches of each league are a warm-up period and are not predicted; the 795 matches above are the ones the model actually produced a prediction for.

How we measure

For every prediction the model is trained only on matches played BEFORE that match (walk-forward validation). Matches played on the same day are excluded too — since we can't know which finished first, the safe side is taken. The model can never see the result of the match it is predicting; otherwise the numbers on this page would come out falsely good.