Goal Model
How likely is today's fixture to produce three goals — and will both sides score? Every number comes with the working behind it, and every past board is published whether it was right or wrong.
Today's matches
How it works
The full method, in plain language — including what the model cannot see. Expand any question to read it.
What exactly is being predicted?
Two questions per fixture, and only these two. Over 2.5 goals asks whether the match will finish with three or more goals. BTTS — both teams to score — asks whether each side will score at least once.
Both come from one underlying estimate: how many goals each side is expected to score, written λ (lambda). Everything else on this page is the working behind those two numbers. Regulation time only — extra time and penalties are excluded.
Where does the data come from?
Fixture lists, standings tables and recent results come from a football data provider, refreshed on a schedule and cached so that a single day's run stays inside a fixed request budget. Historical results are kept in our own archive and are what the over-2.5 and BTTS base rates are computed from — a league table records goals, but not how those goals were distributed across matches.
Every published number names its own source. Open any match and the “League baseline” block tells you whether it came from a live standings table, from our archive, or from a global fallback used when neither was available.
How the model turns data into a probability
Start with the competition average: how many goals a home side scores in a typical match of this league, and how many an away side scores. Then measure each team against that average — an attack strength of 1.20 means a side scores twenty per cent more than a typical team would in the same fixtures; a defence strength below 1.00 means it concedes fewer.
Multiplying those strengths by the league average gives the expected goals for each side. From a pair of expected-goal numbers you can compute the probability of every scoreline, and from the full grid of scorelines you can read off both answers: the chance of three or more goals, and the chance that neither side is kept out.
The scoreline grid uses a low-score correction. Treating the two sides' goals as completely independent is known to under-count 0-0 and 1-1 and over-count 1-0 and 0-1 — real matches tighten up when they stay level. The correction nudges those four cells back towards what actually happens.
Why small samples get pulled towards the average
A team that has scored six goals in two matches has not proven it is a three-goal-a-game side; it has proven very little. Taking that 3.0 at face value would push its fixtures to the top of the list on the strength of two results.
So every strength estimate is pulled back towards the league average by an amount that depends on how many matches stand behind it — hard for two matches, barely at all for twenty. Early in a season the pull is stronger still. This is why the page shows more tier B and C confidence in August than in March: the model is reporting that it does not yet know much, which is the honest answer.
How recent form is weighted
Three views of a team are blended rather than one: its season record at this venue (home form for the home side, away form for the away side), its last six matches across all venues, and its last five at this venue. The season record carries the most weight because it has the largest sample; recent form carries less, but enough to register a side that has genuinely changed.
When one of the three has too few matches to mean anything, it is dropped and the remaining weights are redistributed — rather than filling the gap with a guess.
Blending back to the league base rate
The final published probability is not the raw model output. It is blended with the competition's own base rate — how often matches in this league actually go over 2.5 — in a proportion set by how much data the estimate rests on. A fixture between two sides with twenty matches each leans mostly on the model; one built from four matches leans mostly on the league. This is what stops a thin sample producing a confident-looking number.
What A, B and C confidence mean
Confidence is about the evidence behind a number, not the size of the number. An 80% probability built on two matches is less trustworthy than a 62% built on twenty, and the tier is what tells them apart.
- A — high. Full standings, healthy venue and form samples on both sides, an ordinary league or cup fixture.
- B — moderate. Something is thin: a short form window, an older standings snapshot, an early cup round.
- C — low. Several caveats at once, or a fixture type the model handles poorly — friendlies above all.
Every match card carries its tier, and opening a match lists exactly which caveats produced it.
What the model does not know
This is the part worth reading before trusting any number on this page.
- Injuries and suspensions. A side missing its first-choice striker looks identical to a full-strength one.
- Team selection. A rotated XI in a dead-rubber fixture is invisible to the model.
- Weather and pitch. No weather data is used at all.
- Motivation and context. A side already safe, a manager sacked yesterday, a derby's temperature — none of it is modelled.
- Player-level information. The model works entirely at team level, from goals scored and conceded.
Those gaps are why published probabilities are clamped at both ends and never exceed the low nineties. A model with no knowledge of who is actually playing has not earned the right to say 97%. When you see a number near the top of the range, read it as “as strong as this method gets”, not as near-certainty.
How we grade ourselves
Two boards of ten matches — one for Over 2.5, one for BTTS — are chosen and frozen before kick-off, with the probabilities exactly as published. Once results are in, those same picks are scored. They are never re-selected afterwards, which is the only thing that makes a claimed hit rate mean anything.
Yesterday's boards, and the rolling accuracy over the last 7 and 30 days, are published below. Alongside the hit rate you will find a Brier score: the average squared error of the probabilities themselves, where lower is better and 0.25 is what you would get by saying “50%” to everything. It punishes confident wrong calls in a way a hit rate does not.
Postponed and abandoned matches are voided rather than counted as either result.