Week One Told You Nothing

sports
football
premier-league
statistics
betting
data-viz
I ran an 18-season LOSO horse race — prior, rational updater, market posterior — to ask whether the Week-2 odds are actually better forecasts than the preseason odds were. They are not.
Author

Kivan Polimis

Published

September 26, 2026

The 2026/27 Premier League season opened last month, and the bookmakers did what they always do after an opening weekend: they moved. Teams that won saw their season odds shorten by Tuesday; teams that lost drifted out. The question I want to ask is not whether the market should have moved — new information arrived, and a price that ignores information is a worse price — but whether it should have moved that much.

The Bayesian framing makes the stakes precise. The preseason odds are the market’s prior, the distillation of everything knowable before a ball is kicked. Week 1 is one observation — 1 of 38 matches, about 2.6% of the season’s evidence, and a noisy 2.6% at that. The Week-2 odds are the market’s posterior. A rational updater confronted with one match barely moves. If the market moves substantially more than the results justify, and seasons still end near the preseason prior anyway, the Week-1 move was overreaction. I have eighteen seasons of data to check.

The prior is hard to beat

Before asking whether the posterior overshoots, it is worth establishing what the prior is worth — because the preseason odds turn out to carry a startling amount of information. For every team in every season from 2008/09 through 2025/26 I built a preseason strength rating: a regression-toward-the-mean seed from last season’s table, refined by the Week-1 closing lines, which are posted before any match is played. That prior strength correlates with final league points at r = 0.81. For instance, sort the 360 team-seasons into preseason terciles and the top tier goes on to finish in the top four 52% of the time, the middle tier 5% of the time, and the bottom tier never — not once in eighteen seasons. Whatever Week 1 tells the market, it is competing against a prior that already knows most of the answer.

Preseason prior strength against final league points, 360 team-seasons (2008/09–2025/26). Colour is the preseason tercile; the correlation is r = 0.81.

The market moves — about 20% more than the results justify

Now the move itself. Per team, I define the Week-1 shock as points earned minus points expected — where the expectation comes from that very match’s own pre-match odds, so opponent quality, venue, and the market’s prior all net out in one step. Per Week-2 match, I measure the surprise: how far the Week-2 closing line sits from where the prior ratings said it should. Pooling 167 Week-2 matches across 18 seasons and regressing surprise on shock (season-clustered errors) gives the market’s average response, β̂ = 0.047 (SE = 0.011). The rational benchmark — an Elo-style updater with its response fitted historically on the same data, refit onto the same scale — moves K̂ = 0.039 per unit of shock. The ratio is 1.20: the market moves roughly 20% more than a calibrated results-only updater would.

One honest caveat before that number hardens into a headline. The gap β̂ − K̂ is about 0.008, which is less than one clustered standard error (0.011). At full sample the ratio is descriptive — suggestive of overreaction, not statistically distinguishable from a rational update. Which is exactly why the next section exists: rather than argue about the coefficient, I can race the three forecasts and let the season endings adjudicate.

The horse race

Here is the design. For each of the 18 seasons I hold it out entirely and fit everything — the goals model, the Elo K, the seed and shrinkage weights, the market-response β — on the other 17. Then I build three ratings snapshots for the held-out season: the prior (preseason ratings, untouched), the rational posterior (prior + K̂ · shock), and the market posterior (prior + β̂ · shock). Each snapshot simulates the full 380-fixture season 10,000 times, and the resulting top-4 and relegation probabilities are scored against what actually happened, with Brier scores. No checkpoint ever sees the season it is scored on.

Table 1
  Brier (top-4) Brier (relegation)
Prior 0.0743 0.1005
Rational 0.0751 0.0987
Market 0.0758 0.0987

On top-4, the prior wins outright: the market posterior scores worse (mean difference +0.0015, season-block bootstrap 95% CI [−0.0016, +0.0046]) — and so, more quietly, does the rational updater (+0.0007). On relegation, the market is marginally better (−0.0018), but the CI straddles zero ([−0.0062, +0.0024]). The pre-registered primary sign tests agree: market vs prior on Brier top-4, Holm-corrected p = 1.000; on Brier relegation, Holm-corrected p = 0.961. Some caution is owed to the sample — eighteen seasons is eighteen coin flips for a sign test — but the direction of the caution matters: nothing here suggests the Week-2 odds are better season forecasts than the preseason odds were, and on top-4 the point estimate runs the other way. Eighteen seasons is a small sample of heavily overlapping clubs, so these season-level intervals and sign tests are best read as descriptive: the design has limited power to detect an effect of this size, and the club overlap across seasons means the intervals are, if anything, slightly too narrow.

Per-season Brier score difference, market posterior minus prior, LOSO. Positive = the market’s Week-2 read was a worse season forecast. Orange dots are seasons; the black diamond and rule are the bootstrap mean and 95% CI.

The per-season dots are the honest picture: in some seasons the Week-2 read helps (2009/10 saves the market a solid margin on top-4), in others it hurts (2014/15 costs it more than the average gap), and the mean sits within a hair of zero on relegation and on the wrong side of it on top-4. Eighteen seasons, ten thousand simulations per checkpoint, every fit walled off from the season it was scored on. The market moved. It didn’t help.

Does the market take it back?

If the Week-1 move is overreaction, there should be a signature in the weeks that follow: re-run the surprise regression against the Week-k odds for k = 3, 4, 5, 6, and the coefficient should decay toward zero as the market unwinds its overshoot. These are full-sample descriptive fits, not LOSO scores, so read them as texture rather than verdict.

Week-1 shock coefficient β_k against Week-k odds, k = 2…6, with 95% CIs (full-sample, season-clustered; descriptive, not LOSO). Dashed line: zero. Dotted line: the Week-2 anchor, β = 0.047.

The clean story would run: β falls week by week and reaches zero once the market has digested its mistake. The actual curve half-cooperates. β drops from 0.047 at Week 2 to the low 0.03s at Weeks 3 and 4, then collapses to 0.004 at Week 5 — which looks exactly like the reversal signature, except the confidence interval straddles zero ([−0.015, +0.025]) — and then bounces back to 0.044 at Week 6. A true unwind does not bounce. The honest reading is that the decay is noisy: the Week-5 dip may be the market taking the overreaction back, or it may be sampling noise in ~163 matches, and this design cannot say which. I will not oversell a curve that refuses to decay.

What Week 1 is worth

The thesis survives the caveats. The preseason odds already contained the information — r = 0.81 against final points, a 52/5/0 top-4 gradient across the terciles — and Week 1 gave the market an excuse to move on top of it. The market took the excuse, moved about 20% more than a calibrated results-only updater would, and by May had nothing to show for it: the Week-2 posterior beat the preseason prior on neither primary metric, and on top-4 the point estimate says the move made the forecast worse. So when the odds shift dramatically after this or any opening weekend, some caution about reading too much into it is warranted. The season’s evidence arrives 38 matches at a time. Week 1 is one of them.


Match results and 1X2 odds from football-data.co.uk, seasons 2008/09–2025/26 (Bet365 lines; closing where posted from 2019/20, opening before — about 39% of the surprise rows use true closing lines). All numbers in this post are frozen analysis outputs in the post’s data/ directory; the horse race is nested leave-one-season-out with 10,000 simulation replicates per checkpoint, and sign tests are Holm-corrected. Full method spec and code: github.com/kpolimis/kpolimis.github.io.