How the model gets checked.
Not football tactics — this is where we're transparent about our own methods. How the numbers are built, how we validate them, and where they've gotten it wrong.
Does a 70% consensus actually win seven times in ten over a real sample? The methodology behind checking that, and what we've found. Full note below.
The tradeoffs we weighed before deciding a single number wasn't enough, and why we still publish one anyway on Club Rankings.
How we tested the 1.70–2.50 odds band rule against several seasons of historical data before ever publishing a live pick.
We tested Form ratings at several different lookback windows before settling on five matches. Here's what changed at each one.
A look back at every fixture where the consensus missed, searching for a shared pattern rather than treating each one as unrelated.
The technical detail behind the calibration note above — bucket sizes, sample thresholds, and how we handle small-sample leagues.
Mapping the specific markets and leagues where FDVS's read and the betting market's price diverge more often than chance would predict.
How we validate consensus accuracy — and what "calibration" actually means
A consensus percentage is only as trustworthy as its calibration — a technical word for a simple question: when we say a fixture is 70% likely to go a certain way, does that outcome actually happen close to seven times in ten, across enough matches for the pattern to be real rather than luck? This note is about how we actually check that, not just claim it.
What "calibration" means, concretely
Take every fixture where our consensus sat between 65% and 75% for a given outcome — call it the "70% band." If our model is well-calibrated, that outcome should have actually happened somewhere close to 70% of the time across all fixtures that landed in that band, over a large enough sample. If it happened 90% of the time, our model is underconfident in that range — we could be more assertive. If it happened 50% of the time, we're overconfident — the number is promising more certainty than the underlying data supports. Calibration isn't about whether any single call was right or wrong; it's about whether the percentage itself means what it claims to mean, averaged across many fixtures.
Why we bucket into bands rather than checking single percentages
No single fixture can validate a percentage — a 70% call that loses tells you almost nothing on its own, since a genuinely well-calibrated 70% is still expected to lose three times in ten. The only way to check calibration honestly is to group many fixtures with similar consensus readings into a band and look at the hit rate across the whole group. We currently check five bands — 50-59%, 60-69%, 70-79%, 80-89%, and 90%+ — and expect each band's actual hit rate to land reasonably close to its midpoint over a full season of tracked fixtures.
What happens when a band comes back miscalibrated
When a band's real hit rate drifts meaningfully from what it should be — say, our 80-89% band winning only 65% of the time over a large enough sample to matter — that's treated as a genuine finding, not noise to explain away. It usually means one of the inputs feeding that confidence level needs re-weighting, or that a specific league or market type is systematically pulling that band off target. This is exactly the kind of check that led to some of the model adjustments referenced in our other research notes, including how we settled on a five-match lookback window for Form.
Why the results page is part of this validation, not separate from it
Every result graded as a "miss" on the results page is a real data point feeding this same calibration check — publishing our misses isn't just a transparency gesture, it's the actual mechanism that makes calibration checking possible. A results archive that only kept the wins would make genuine calibration auditing impossible, because the hit rate for any confidence band can only be computed honestly if every outcome in that band is counted, not just the favourable ones.
What good calibration looks like versus what confidence feels like
A well-calibrated 55% and a well-calibrated 90% are both doing their job correctly even though they feel completely different to read — the 55% is telling you honestly that this one is close to a coin flip, and the 90% is telling you honestly that this one is about as close to settled as football gets. The goal was never to make every number feel confident; it was to make every number mean what it says, whatever that number happens to be.
Sample size and why some leagues take longer to calibrate
A newly added league doesn't get treated as calibrated until it has enough tracked fixtures across enough confidence bands to check honestly — a handful of matches isn't enough to know whether a band is accurate or just got lucky. This is the same minimum-sample standard applied to league averages, player statistics, and club records elsewhere on the platform, and it's part of why a newer competition's consensus numbers should be read with slightly more caution until that history builds up.
This is an ongoing process, not a one-time check
Calibration isn't something we verified once and moved on from — it's checked continuously as results come in, because a model that was well-calibrated at the start of a season can drift as team form, league tempo, and squad turnover shift the underlying patterns it was originally tuned against. The research notes linked above go deeper into specific pieces of this process; this page is the overview of why the process exists at all. If a band ever drifts and stays drifted, the honest response is adjusting the model and saying so — not quietly retiring the check.