Back in March 2025, initially most performant model observed was GPT-3.5 Turbo Instruct, playing white in Continuation mode. Here is a spontaneous live video demonstration: YouTube
In full information chess (Reasoning mode), o1-mini showed strength, winning the first tournament (long since decrowned & ˟deprecated).
Elo progression page: Chess Elo Race

Data collection: Matchups and game volume are mostly influenced by model competency/speeds, data-gaps, API constraints and/or budget.
This chess inference ran me ~$5500. Data collected move-by-move — see an example match: Log | Replay.
Any recent (≥ Feb 15, 2026) reasoning replay automatically features full move-by-move model commentary. All replays have custom move-labeling with eval bars.

Elo is always relative to the player pool it's measured in. lichess rating chess.com rating fide rating engine rating llm rating perceived rating
LLMs operate alien in comparison to Humans and traditional Chess engines. They have a completely different error distribution and are thus not directly comparable. Due to frequent inquiries, while I don't mix isolated Elo pools, and bot tuning is inconsistent, from low-scale testing and under these caveats I can disclose that in my ballpark observation ~1200 Continuation-rated models matched ~1400 rated Lichess bots, while ~1850 Elo in this leaderboard (best reasoners) correlated roughly to Lv16 Expert 2000 chess.com bots. Example matches of Gemini-3.1-Pro-Preview (Reasoning ~1837) vs 1900 (W), 2000 (WD), 2100 (L), 2200 (L), 2400 (L)

Like all my benchmarks, unless specifically indicated otherwise, models are tested on their default settings. This prevents exploding cost, test time & flooding the leaderboard with unrealistic benchmark-optimized configurations that aren't representative of the average user experience. This principled method aligns incentives correctly and keeps the evaluation ecosystem feasible, honest and representative.