Track record
How good a prediction model is isn't claimed, it's measured. This page shows the model's past performance unfiltered — unfiltered.
These figures come from backtesting, not from live tracking. For each past match the model was trained only on matches played before it, then tested. Once live tracking is running this page will update with real-time data. Scope: 15 leagues, 133 matches, 2,318 predictions. These numbers are computed from the database every time the page loads; there is no hand-updated copy. (September 2026 – September 2026)
Last updated: 16 September 2026 at 19:10 · Methodology version: 10
The model knows when it knows
The real test of a prediction model is not that it answers every match, but that it is right on the ones it is sure about. Below are the predictions where the model said “I am at least this confident”, and how many of them actually turned out correct. The thresholds were not cherry-picked; every tier is here.
| What the model said | Predictions | Correct |
|---|---|---|
| Match result “at least 50% confident” | 61 | 54.1% |
| Match result “at least 60% confident” | 32 | 71.9% |
| Match result “at least 70% confident” | 17 | 88.2% |
| Double chance “at least 75% confident” | 122 | 75.4% |
| Double chance “at least 80% confident” | 63 | 84.1% |
| Double chance “at least 85% confident” | 40 | 92.5% |
The thing to notice: the stated percentage and the actual percentage track each other, and the actual is slightly higher. The model does not overstate. If you see “95% accuracy” somewhere, the question to ask is: over how many predictions, in which market, and were the incorrect ones counted.
If we had to answer every match
The harshest measure: the case where the model is forced to tick one of three boxes on every single match. Nobody plays that way, but we have nothing to hide — what it is compared against and where the ceiling sits are written next to it.
70.6%
The ceiling of this game. 29.4% of matches end in a draw, and a draw is almost never on its own the single most likely outcome — nobody forced to tick one of three boxes, us or anyone else, gets meaningfully above this.
75.4%
of the 122 double-chance predictions we called “at least 75%” were correct (92 of them).
The double-chance box is shown on the home page too. If it had come out the other way — delivering less than we claimed — it would sit in the same place at the same size. The showcase exists to show what was measured, not to pick the good news.
Does it work in other leagues too
A model doing well in a single league can also be explained by it having fitted that league's quirks. The real test is producing the same result in other leagues with no tuning at all. The parameters were chosen on the Süper Lig; the other leagues below took no part in that choice.
| League | Matches | Our accuracy | vs random |
|---|---|---|---|
| Premier LeagueEngland | 10 | 40.0% | +6.7 pts |
| UEFA Champions League | 4 | — | — |
| La LigaSpain | 15 | 53.3% | +20.0 pts |
| BundesligaGermany | 9 | 55.6% | +22.2 pts |
| 1. LigTurkey | 10 | 40.0% | +6.7 pts |
| EredivisieNetherlands | 10 | 50.0% | +16.7 pts |
| Segunda DivisiónSpain | 11 | 36.4% | +3.0 pts |
| ChampionshipEngland | 14 | 42.9% | +9.5 pts |
| Ligue 1France | 9 | 33.3% | +-0.0 pts |
| Serie AItaly | 10 | 50.0% | +16.7 pts |
| Süper LigTurkey | 9 | 44.4% | +11.1 pts |
| Ligue 2France | 9 | 55.6% | +22.2 pts |
| 2. BundesligaGermany | 9 | 44.4% | +11.1 pts |
| PremiershipScotland | 3 | — | — |
| UEFA Europa League | 1 | — | — |
In 6 of 12 leagues the model beats random ticking by -0.0 to 22.2 points — 133 matches in total. In every one of those leagues it also beats that league's own historical average. In one or more leagues we came in below the base rate; that row is not hidden.
What this table proves, and what it does not
The accuracy figures here are the numbers under a rule of “tick one of three boxes on every single match”. Nobody has to play that way; the confidence table above shows what the product actually does. This table's job is not to persuade — it is to show that the METHOD WORKS IN EVERY LEAGUE.
The reference point is ticking one of three options at random: 33.3%. Not a chosen number but the definition itself — three options, one is right. The difference is what the model genuinely adds, and it is written in points: the difference between two percentages is points, not a percentage.
A harder comparison is measured too: each league's own historical average — a forecaster that already knows home advantage. The model beats that in every league as well; the audit and the track record on this page rest on that measure.
The numbers in context: ticking at random gives 33%. The upper bound is 70.6% — because 29.4% of matches are draws and nobody can foresee a draw. The room to move is narrower than it looks, and that is a fact about football.
So wherever you see a high accuracy figure, ask the same question: on which matches, in which market, and were the incorrect ones counted too. On this page they all are.
This table covers only the leagues published on the site; leagues ingested for measurement but not put on screen do not enter it either — the track record's cohort and the site's cohort are the same one. The second-tier leagues sit clearly below the first-tier ones; that row is not hidden, it is in the table.
Calibration curve
The answer to “of the matches we called 60%, did 60% actually happen?” In a well-calibrated model the two columns are close together.
| Confidence band | Model said | Actually happened | Matches |
|---|---|---|---|
| 10%–20% | 16.0% | 30.0% | 10⚠ |
| 20%–30% | 26.5% | 23.5% | 17⚠ |
| 30%–40% | 36.5% | 42.9% | 28⚠ |
| 40%–50% | 44.2% | 34.4% | 32 |
| 50%–60% | 53.9% | 40.0% | 20⚠ |
| 60%–70% | 65.3% | 44.4% | 9⚠ |
| 70%–80% | 77.2% | 87.5% | 8⚠ |
| 80%–90% | 85.8% | 80.0% | 5⚠ |
Being watched: under-rested home teams
When the home team has had 4 days or less since its last match, the model used to systematically underrate it. In the latest measurement that bias is gone: we said 44.4%, the actual figure was 44.6% (3,129 matches).
The flaw was neither invented nor buried: it was genuinely measured on 8 August 2026 across 719 matches in 6 leagues (z=+3.14) and has been stated on this page ever since. As coverage grew and the sample multiplied, the gap dissolved statistically — meaning the original finding was largely small-sample noise. This box updates itself after every audit run; if the bias returns it will show up here as a warning again.
Largest deviation: in the 40-50% band the model said “44.2%” on average, and 34.4% actually happened. So in that band it trusts itself too much — the error is on the optimistic side, which is more serious. It stands here as a shortcoming to be fixed; it isn't hidden.
Other markets
Every market is measured separately against its own baseline. Instead of saying “the model is good already, so this market must be good too”, each one was tested. The n column shows how many matches each row rests on — lower Brier is better.
This table does not by itself decide which market gets published. That decision came from a separate walk-forward test: the coefficient is chosen on the first 75% and tested on the final 25% holdout. A result that looks good on all the data but collapses on the holdout is data dredging, not a real edge. The first-half goal markets failed that test and are not published.
| Pazar | n | Model Brier | Taban oran | Edge |
|---|---|---|---|---|
| Double chance 12 | 133 | 0.3860 | 0.3932 | +1.8% |
| Double chance 1X | 133 | 0.4118 | 0.4586 | +10.2% |
| Double chance X2 | 133 | 0.4427 | 0.4861 | +8.9% |
| Expected fouls — away | 127 | 0.4736 | 0.4915 | +3.6% |
| Expected fouls — home | 127 | 0.4981 | 0.5198 | +4.2% |
| First half — level | 128 | 0.4704 | 0.4714 | +0.2% |
| First-half double chance 12 | 128 | 0.4704 | 0.4714 | +0.2% |
| First-half double chance 1X | 128 | 0.3766 | 0.3998 | +5.8% |
| First-half double chance X2 | 128 | 0.4369 | 0.4476 | +2.4% |
| First half — away ahead | 128 | 0.3766 | 0.3998 | +5.8% |
| First half — home ahead | 128 | 0.4369 | 0.4476 | +2.4% |
| Both teams to score | 128 | 0.4966 | 0.5008 | +0.8% |
| Expected corners — away | 121 | 0.3859 | 0.4552 | +15.2% |
| Expected corners — home | 121 | 0.4633 | 0.5607 | +17.4% |
| Match result — draw | 133 | 0.3860 | 0.3932 | +1.8% |
| Match result — away win | 133 | 0.4118 | 0.4586 | +10.2% |
| Match result — home win | 133 | 0.4427 | 0.4861 | +8.9% |
| 3+ goals | 128 | 0.4664 | 0.4875 | +4.3% |
Why double chance adds up to 200%
1X, 12 and X2 always sum to exactly 200%. This is not an error, it follows from the definition: the options are not mutually exclusive, they overlap — every outcome is counted in two of the three, so the 100% of 1X2 is doubled. For the same reason 1X plus X2 exceeds 100% by exactly the probability of a draw. This is an invariant, verified by a unit test on every build.
An honest note about the first half: the first-half model is trained separately (it is not a fraction of the full-match prediction), but its edge is markedly smaller than the full-match model's — the measured figure is in the table above. First-half goal markets are not published at all because they failed to beat the base rate.
Goals market (3+ goals)
In this market the model is not symmetric. The measurement showed:
- ▲When the model speaks in the “3+ goals” direction it carries real information.
- ▼In the “fewer goals” direction it does not — those predictions are no better than the historical average.
So the goals market is only shown for matches where the model gives a clear signal in the “3+ goals” direction. For other matches, rather than inventing a number, it says “no clear signal”. Staying silent when you have nothing to say is more honest than filling the space.
Why the incorrect ones were wrong
Counting the correct predictions is easy. What actually teaches you something is what happened when we were wrong. All 73 incorrect predictions were examined one by one.
84.9%
Share of losses decided by a single goal.
47.9%
Share of losses caught out by a draw.
6.8%
Reverse result by 3+ goals — a genuine surprise.
The main culprit behind our losses: the draw
The model said “draw” in only 0.0% of 133 matches. Yet 26.3% of matches ended level.
This is not an error, it's what the maths gives: a draw is almost never the single most likely outcome — with probability spread across three results, one side usually stands out. But when draws do happen they are recorded as losses, and they make up half of ours. That is exactly why double chance performs so much better than 1X2: it covers the draw instead of losing to it.
Is there a parameter to be extracted from the losses
Factors knowable BEFORE the match were scanned: days of rest, fixture congestion, missing players, stage of the season, expected goals. Result: no usable parameter emerged.
One factor looked significant in the raw scan (“accuracy is 8.6 points higher in matches expected to be high-scoring”). Controlled for within the same confidence band, the sign reversed (−6.9 points) — high-scoring matches were already concentrated in the high-confidence bands, and those bands are naturally more accurate. The raw finding was an artefact of the distribution, not of the factor. Without that control we would have drawn a wrong parameter out of it and damaged the model.
Our ceiling: 85% of losses are decided by a single goal. That is football's natural variance; chasing it makes the model fit noise. The share attributable to red cards and in-match events will be measured separately once event data is ingested — that data can't be known before kick-off so it can't be a parameter, but it shows our ceiling.
Tried it, didn't add it
An idea sounding right is not enough. Approaches that were tried and found not to work are listed here too — because what we DIDN'T add tells you as much as what we did.
Head-to-head history (the “bogey team” effect)
The intuition that “team X has a bogey team in Y” was tested. Result: no signal.
- The correlation between the bias in past meetings and the bias in the current match is r = −0.05 — indistinguishable from chance.
- Of the 165 pairings with at least 4 matches, 52 look “dominant”. But even in a world with no pairing effect at all, simulation produces 46.3 of them. The excess of +5.7 sits inside the noise band.
- When a correction was applied to expected goals, the best coefficient came out as zero — that is, applying no correction. Every positive coefficient made the prediction worse.
The honest limit: across the 3 seasons we hold, there are at most 6 matches per pairing. That is enough to say “this data can't show it”, not enough to say “there is no effect”. The test will be repeated when more seasons are available.
An attempt to correct the known bias (sharpening)
To fix the “not confident enough” flaw described above, the probability distribution was sharpened with a correction that preserves the ranking. Result: it didn't work and wasn't applied.
- With the coefficient chosen on the first 75%, the holdout Brier went 0.6020 → 0.6031 (worse).log loss de bozuldu.
- It improved only three of the six leagues and hurt three. A correction that doesn't work everywhere isn't a correction, it's overfitting.
- The calibration error (ECE) got worse at the very thing we were trying to fix.
The bias itself turned out to be fragile: the +3.2 point gap seen for very strong away teams across all the data becomes −0.8 in the most recent 25%. So it isn't a stable enough flaw to correct. Another method will be tried; if that fails too, the flaw will keep standing here as it is.
Pulling the draw probability toward the base rate
Since half our losses are caught out by draws, this was tried: pull each match's draw probability toward the league's long-run draw rate. Result: no gain, it makes things worse.
The measurement showed why: the model's average draw probability is 25.6%, while the actual draw rate is 23.3%. So the model doesn't under-call draws — it slightly OVER-calls them. The impression that “the model doesn't see draws” is wrong — it gives each match the right probability, but among three outcomes a draw is almost never the highest (in 3 of 943 matches; the highest value seen was 37.8%). That is what the maths gives, not a flaw to be corrected. The right answer was not to force the model but to bring double chance forward.
Number of missing players (injuries)
Tested on 859 matches against real per-match injury records. Result: no measurable contribution.
- The correlation between the difference in missing players and the bias in the result is r = −0.04. The direction is right — the depleted team scores less — but the magnitude doesn't separate from noise (p = 0.28).
- Added to the model as a correction, the best coefficient again came out as zero.
Why this isn't surprising: we have no player-quality data, so a missing star and a missing substitute count the same. On top of that, the effect of ongoing injuries is already reflected in the team's recent results and therefore in its rating. This is not a “never”, it's a “not like this”.
Team-specific home advantage
League-wide home advantage is already in the model (home teams score 1.63 goals per match, away teams 1.28). What was tested is whether the advantage varies from team to team. Result: no contribution.
The team factors looked meaningful (highest 1.25, lowest 0.79) but adding them to the model made the prediction worse at every coefficient. There are only ~17 home matches per team per season; most of the “fortress at home” label that comes out of that little data is noise.
An independent strength measure (ClubElo)
Because squad-value data could not be obtained from a source with clean licensing, an independent team-strength measure was tried. Result: the effect is real but small (+0.46%).
What shows this isn't noise: when the coefficient was chosen on the first 75%, the best value came out as 0.12 — and the holdout's own optimum was also at 0.12. Under overfitting those two points diverge. Even so, for a gain of that size we don't take on a dependency with unclear licensing; if the terms become clear and it's confirmed across more seasons, it will be added.
Technical detail: Brier score and method
The measure behind the comparisons above. The Brier score is the mean squared error between the probability stated and the outcome that occurred; lower is better. Accuracy asks “did you tick the right box”, the Brier score asks “how right was the percentage you stated” — the second is stricter, and it is the one we use to correct the model.
| Method | Brier | Difference |
|---|---|---|
| Model | 0.6203 | — |
| Baseline: historical base rate | 0.6689 | +7.3% |
Each match's prediction was computed only from matches played BEFORE it (walk-forward validation), and the baseline used the same window — so the comparison is fair. The first 120 matches of each league are a warm-up period and are not predicted; the 133 matches above are the ones the model actually produced a prediction for.
How we measure
For every prediction the model is trained only on matches played BEFORE that match (walk-forward validation). Matches played on the same day are excluded too — since we can't know which finished first, the safe side is taken. The model can never see the result of the match it is predicting; otherwise the numbers on this page would come out falsely good.