Live · AI research

How the model gets checked.

Not football tactics — this is where we're transparent about our own methods. Every research note lives right here on this page, in full.

Model validation · 8 min read
How we validate consensus accuracy — and what "calibration" actually means

Does a 70% consensus actually win seven times in ten over a real sample? The methodology behind checking that, and what it actually requires.

What "calibration" means, concretely

Take every fixture where a consensus sat between 65% and 75% for a given outcome — call it the "70% band." A well-calibrated model should see that outcome actually happen somewhere close to 70% of the time across all fixtures landing in that band, over a large enough sample. If it happened 90% of the time, the model is underconfident there. If it happened 50% of the time, it's overconfident — the number is promising more certainty than the data supports. Calibration isn't about any single call being right or wrong; it's about whether the percentage itself means what it claims, averaged across many fixtures.

Why banding, not checking single percentages

No single fixture can validate a percentage — a 70% call that loses tells you almost nothing on its own, since a genuinely well-calibrated 70% is still expected to lose three times in ten. The only honest check is grouping many fixtures with similar readings into a band and looking at the hit rate across the whole group — five bands in practice: 50-59%, 60-69%, 70-79%, 80-89%, and 90%+, each expected to land reasonably close to its midpoint over a full season.

What happens when a band comes back miscalibrated

When a band's real hit rate drifts meaningfully from where it should sit — an 80-89% band winning only 65% of the time over a sample large enough to matter — that's a genuine finding, not noise to explain away. It usually means an input feeding that confidence level needs re-weighting, or a specific league or market is systematically pulling that band off target.

Why the results page is part of this, not separate from it

Every result graded as a miss on the results page is a real data point feeding this same calibration check — publishing misses isn't just a transparency gesture, it's the actual mechanism that makes honest calibration checking possible. A results archive that only kept the wins would make the hit rate for any band uncheckable.

This is ongoing, not a one-time check

A model well-calibrated at the start of a season can drift as team form, league tempo, and squad turnover shift the patterns it was originally tuned against — so this is checked continuously as results come in, not verified once and set aside. If a band ever drifts and stays drifted, the honest response is adjusting the model and saying so.

Methodology · 6 min read
Why we use a six-axis system instead of one power rating

The tradeoffs weighed before deciding a single number wasn't enough, and why we still publish one anyway on Club Rankings.

A single number is easy to rank, hard to trust

A single power rating is genuinely convenient — it sorts cleanly, it's easy to compare at a glance, and it's what most fans are used to seeing. The problem is that a single number necessarily hides how a club actually arrived at it. Two clubs with an identical Overall score can have completely different underlying profiles — one built on an overwhelming attacking edge, the other on defensive solidity and discipline — and a single figure can't distinguish between them.

Why six axes specifically, not more or fewer

Six was chosen because each axis captures something genuinely distinct that the others don't: Attack and Defense measure output on either end of the pitch, Form captures recent trajectory separately from season-long level, xG separates underlying chance quality from raw scoring, and Set Pieces and Discipline isolate two specific phases of the game that don't track cleanly with general attacking or defensive quality. Adding more axes started producing genuine overlap between categories rather than new information; six was the point where each axis was still telling us something the others weren't.

Why we still publish a single Overall score anyway

Club Rankings' Overall score exists because sorting and comparison are genuinely useful, and pretending otherwise would just push people toward a worse, unofficial way of comparing clubs themselves. The compromise is transparency about what that single number actually is — a simple, equally-weighted average of the same six real axes, computed openly rather than through an opaque formula, so anyone can see exactly how it was built rather than trusting a black box.

The real cost of this choice

Six numbers are objectively harder to present than one, and harder to compare across a whole league at a glance — which is exactly why pages like team analysis and this research page exist, to do the harder work of explaining the six-axis picture rather than asking every visitor to do that work themselves from a raw ranking table.

Model validation · 7 min read
Backtesting Prime Pick selection criteria

How the 1.70–2.50 odds band rule was tested against historical data before any live pick was ever published.

Why this band specifically, not a wider or narrower one

The 1.70–2.50 band wasn't chosen for a round-number reason — it was tested against historical results to find the range where genuine model-and-market agreement actually translated into a meaningfully better-than-baseline hit rate. Below roughly 1.70, a favourite is often so short that the market has already priced in most of the value, leaving little room for a curated pick to add anything the raw odds didn't already tell you. Above roughly 2.50, variance climbs fast enough that even a genuinely strong read starts losing more often than the odds alone would suggest is worth curating around.

Testing against real historical fixtures, not live picks

Before Prime Pick ever went live, the same selection logic — requiring both a statistical read and independent agreement — was run backward against a real, already-settled historical sample, checking what the hit rate would have been if this exact rule had been applied at the time. This is a genuinely different, stricter test than simply watching live picks perform going forward, since the historical outcomes were already fixed and couldn't be influenced by knowing the rule in advance.

What the backtest actually confirmed

The backtest's real purpose wasn't to prove the band produces guaranteed winners — no honest backtest can promise that — it was to confirm the band produces a genuinely different, tighter selection than picking fixtures at random within that odds range, and that requiring independent model-and-editorial agreement meaningfully outperforms either signal alone on the same historical sample.

Why this gets re-checked, not just validated once

A backtest is only as good as the historical period it's run against, and football patterns shift over seasons — so the same validation is periodically re-run against more recent data rather than treated as a one-time proof that never needs revisiting.

Research note · 5 min read
How much does recency weighting actually matter?

Form ratings were tested at several different lookback windows before settling on five matches. Here's what changed at each one.

The tradeoff behind any lookback window

A shorter lookback window makes Form more reactive — a team's rating shifts quickly after a strong or poor run — but it also makes the number noisier, since a handful of matches can be dominated by an easy or unusually difficult run of fixtures rather than a genuine change in level. A longer window smooths that noise out but becomes slower to reflect a real, sustained shift in a team's actual level.

What testing shorter windows showed

A three-match window reacted fast to a hot or cold streak, but it was also visibly more volatile — a single result could swing a rating further than felt justified by one match, and the number occasionally told a story that didn't survive even one more fixture.

What testing longer windows showed

A ten-match window was noticeably more stable, but it was also slow to reflect a team that had genuinely turned a corner — a side that had clearly improved over its last five matches could still be dragged down by a poor run from over a month earlier, which felt like it was measuring history more than current level.

Why five landed as the balance point

Five matches turned out to be roughly where reactivity and stability crossed — recent enough to reflect a genuine current level, long enough that one unusual result doesn't dominate the whole reading. It's not a universal law of football, just the point that tested best against the actual patterns in the data being tracked.

Research note · 6 min read
What we learned auditing our own upset predictions

A look back at every fixture where the consensus missed, searching for a shared pattern rather than treating each one as unrelated.

Why one miss means nothing, but a pattern of misses does

Any individual upset is expected — a well-calibrated 70% read is still supposed to lose three times in ten, so a single miss proves nothing on its own. The useful exercise is going back over every real miss as a group and asking whether they share anything in common, rather than treating each one as an unrelated, unlucky result.

What this kind of audit is actually looking for

The question isn't "was this call wrong" — it's "were our misses randomly distributed, or did they cluster somewhere specific": a particular league, a particular market type, a particular gap size between the two sides. Random distribution across misses is actually the reassuring outcome — it suggests the model is failing the way any well-calibrated model should, unpredictably. Clustering is the finding worth acting on.

How a real finding here changes the model

When misses cluster somewhere specific, that's treated as a genuine signal that an input feeding that specific league or market needs adjustment — not evidence the whole model needs rebuilding, and not something explained away as bad luck. This kind of targeted finding is exactly how a specific weighting gets revisited, rather than changing something that wasn't actually broken.

Why this audit never stops

A pattern that doesn't exist in one season's misses could emerge in the next as leagues and squads change, so this isn't a project with a finish line — it's a standing check re-run as more real results accumulate.

Methodology · 9 min read
Calibrating consensus percentages against real outcomes

The technical detail behind the calibration note above — bucket sizes, sample thresholds, and how small-sample leagues are handled.

Why bucket width is a real design choice, not arbitrary

A 10-point bucket width (50-59%, 60-69%, and so on) balances two competing needs: narrow enough that everything inside a bucket represents a genuinely similar confidence level, wide enough that each bucket accumulates enough real fixtures to check honestly. A narrower bucket would be more precise in theory but would take far longer to build a meaningful sample in each one.

What counts as "enough" sample before trusting a bucket

A bucket with only a handful of fixtures can look miscalibrated or perfectly calibrated purely by chance — a small sample simply doesn't carry enough statistical weight to distinguish a real pattern from noise. A meaningful minimum sample size is required before a bucket's hit rate is treated as a real signal rather than an early, unreliable read.

How small-sample leagues are handled specifically

A newly tracked or lower-volume league won't have enough fixtures in any single confidence bucket to calibrate independently for a while — rather than force a premature verdict, that league's consensus numbers are flagged as not yet independently calibrated until enough real fixtures accumulate, the same minimum-sample principle applied to team ratings and league averages elsewhere on the platform.

What changes once a bucket has real, sufficient data

Once a bucket clears the sample threshold, its real hit rate becomes an active input into ongoing model checks — a bucket that consistently reads correctly needs no adjustment, while one that drifts triggers the kind of targeted re-weighting described in the main calibration note above. This detail page exists specifically so the mechanics behind that process are checkable, not just asserted.

Research note · 6 min read
Where our model and the wider market disagree most

Mapping the specific markets and leagues where FDVS's read and the betting market's price diverge more often than chance would predict.

Disagreement isn't automatically a red flag

A model that always agreed with the market wouldn't be adding anything — the market itself is already a genuine aggregation of a huge amount of information. The interesting question isn't whether disagreement exists, since it always will to some degree, but where it clusters and whether that clustering is meaningful or just noise.

Where disagreement tends to concentrate

Divergence between a statistical read and market pricing tends to show up more in markets built from a thinner underlying signal — certain goals lines and some derived markets — and less in the most heavily-traded, most closely-watched markets like a straightforward match-winner price, where the sheer volume of market activity tends to price in most available information quickly.

Why this mapping matters for how confidence is communicated

Knowing specifically where the model and market tend to diverge changes how much weight either signal deserves in that specific context — a fixture in a market where the two typically align closely is a different situation than one in a market where they routinely diverge, even if the raw percentage looks similar on the surface.

What this doesn't claim

Mapping where disagreement happens isn't the same as claiming the model is right whenever it diverges from the market — plenty of divergence simply reflects the market having access to information not fully captured in the tracked statistics, like team news breaking after data was last refreshed. This research is about understanding the pattern of disagreement, not declaring a winner between the two.