← Blog
2026-09-12

Rating deflation: today's chess elite are underrated

Are modern chess players stronger than those of the past? It makes sense the answer should be yes. Humans tend to get better at things over time, learning from the past, and for chess there has also been an explosion of databases and powerful engines which could make it easier to learn and excel at the game.

But ratings among top players haven’t caught up. Not so long ago, Magnus Carlsen declared a goal of reaching 2900, but now he sits at 2823. Where there were once as many as six players rated 2800+, now there is only one, and the second-highest-rated player has been largely inactive recently. Meanwhile, a new young generation has risen up, and most peak-age players sit far from their career rating peaks.

Have classical ratings deflated, such that a given Elo rating now corresponds to stronger play? According to XPA (Expectation Points Added), the answer is clearly yes.

The method

Computing expectation points added by play with engines and machine learning is not foolproof, but it offers a compelling way to measure changing play quality across time. XPA grades each move by comparing the player’s choice with other candidate moves, according to their predicted move probabilities. It tries to measure how much a player’s expectation is changed each move, compared with what would be expected of a modern GM.

It can be produced move-by-move, objectively, for any game in any era. And it shows a strong relationship with Elo ratings, stronger than accuracy.

To assess how particular rating levels have changed, I first took all analyzed classical games in the database, filtered to those where Elo ratings are available, and grouped them together in two-year bins. Within those bins, each player’s performance in a game contributes its average XPA per move, paired with that player’s Elo rating at the time. Then I fit a linear model targeting three Elo levels: 2600, 2650, and 2700. Within each fit, data near the target rating is weighted more heavily. From these fits we can determine what average level of XPA per move corresponds to each rating, for each two-year time window. In total, the analysis covers 341,935 games and 610,572 player-games from 1970–2025.

The results

This is a striking figure! There is a sustained rise in playing strength at each level, consistently across the last 20-25 years. We can also notice that this process hasn’t been the story since the dawn of Elo ratings; during the 1990s the level of play at each rating seemed to decrease.

I want to be clear about how I’m interpreting this: the dip in the 1990s does not mean players of that era were weaker than those before or after. But it does suggest their ratings were inflated and that in other eras their ratings would have been lower.

To check whether this result depended on how the archive’s composition changed over time, as a control I tried this focused purely on Olympiads. Doing the same analysis on Olympiad data reveals the same story:

Just what is the “modern GM baseline”?

In many places I have described the 0 XPA line as the “modern GM baseline,” which is a way of waving my hand at this complexity. I derived this initially as the average rating in the training set, which spanned classical games from 1994 through early 2026. This was a bit below 2600, and when the site produces XPA today it does so by producing predictions at this level.

But as we see now, the Elo rating of this benchmark has moved! The following shows the Elo rating corresponding to 0 XPA over time:

It turns out that in 2025 terms, 0 XPA corresponds to the mid 2500s! This is a drop of 100 points in just the last 15 years.

Checking phases of the game

Another nice feature of XPA is that we can ask: where in the game is this happening? We might expect play to be improving most in the opening, where computers and databases reward homework. Play has improved in the opening, but also in other phases:

All phases show improved play at a given rating level. To resolve this better I checked the improvement of each phase for 2600s, for the most recent two-year window and one twenty years before. Openings do improve the most, but barely:

Another interesting view is by move number. Here we see an effect consistent with time pressure, which is often present in the lead-up to move 40, where classical players often receive added time. And it’s interesting how the long-ago era actually fares the best beyond move 40, perhaps due to more generous time controls for endgames:

Potential confounders

As mentioned, there is the potential that these results are affected by selection bias over eras. One way to check for this is to see if the effect is sensitive to selection, by checking across cohorts of data:

All of these variants yield a highly similar story of significant deflation over the last 20 years:

  • Equal player-game weights. Each player’s performance in a game contributes one value of XPA and Elo. This is the method used in the above figures.
  • Move-count weights. Here each player’s performance in a game is weighted by its length – longer games contribute more, weighting the endgame more highly.
  • Equal player weights. Each player receives equal total weight within each window, before Gaussian rating weights are applied.
  • Olympiads excluded. This condition removes Olympiad games from the analysis.
  • Same recorded time control. This restricts to perfectly matched time controls, considering only games played with the official FIDE time control of 90 minutes for the first 40 moves, another 30 minutes for the remaining moves, and a 30-second increment from move one.
  • Players aged 25-34. This helps show that a mix-shift of ages is not driving our conclusions.

Prior work

XPA offers a new perspective on this question, but I am also aware of substantial prior work which helps put it in context:

  • Regan and Haworth, “Intrinsic Chess Ratings” (2011). Here the authors tried a very similar idea, using engine evaluations and a statistical model of player choice to assess play quality at move-level, and looking across historical periods for inflation. They found little evidence of inflation, and some evidence of deflation.
  • Sonas, Compression and Calculation Improvements (2024). For FIDE, Jeff Sonas investigated deflation using existing Elo ratings. He described underperforming favorites, underrated entrants, and pandemic-related disruptions. FIDE did adopt reforms in 2024, although I doubt they are strong enough to reverse the trend seen here.
  • Sebastian and Voigt, The Elo ratings: Inflation or Deflation? In this ChessBase article two IMs used centipawn loss to come to a similar conclusion, also finding some evidence of deflating chess ratings.

Future work

I’ve been looking forward to getting this far! Soon I think we can try inflation-adjusted ratings, which are likely to tell an interesting story. I’m also thinking about ways to approach the natural question that follows: why?

Methodological details

  • Fitting XPA to Elo ratings. Games are grouped into two-year windows. Within these, each player’s performance in a game contributes an average XPA per move, paired with that player’s Elo rating. Within each window, I fit the relationship between XPA and Elo, with more weight given to points near the target rating. This weighting is done by a Gaussian kernel (sd = 125 Elo; this parameter was tested and doesn’t influence conclusions).
  • Fitting Elo ratings to 0 XPA. Within each two-year bin, to find the Elo value associated with 0 XPA, I use the relationship fitted from all data centered on 2600 Elo as above. I then solve for the Elo rating at which the fitted line reaches 0 XPA.
  • Defining game phases. Openings here are positions within the first 15 moves of the game. Endgame borrows the Lichess definition, requiring six or fewer pieces remaining (across both sides combined), excluding kings and pawns. All other positions are middlegame.
  • Confidence intervals. All figures show 95% confidence intervals calculated from 2000 bootstrap resamples. Whole tournaments are resampled with replacement within each time window and the model is refitted on each draw. For individual Olympiad CIs, players are resampled. For comparisons between eras, each time period is independently resampled and differences are calculated once per replicate. Confidence limits show the 2.5th and 97.5th percentiles of these bootstrap estimates.