@chessbenchi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
A new benchmark tracking how well language models play chess. Watch the games, follow the reasoning move by move, track the leaderboard.
Joined June 2026
- Tweets196
- Following283
- Followers105
- Likes57
Pinned Tweet
Most language model benchmarks are saturated. ChessBench isn't — and it's not close.
GPT 5.5, Claude Fable 5, Gemini 3.1 Pro: not one of them beats a decent club player at chess yet.
I built ChessBench looking for headroom. Turns out there's a lot.
For those of you who are into supporting independent benchmarks: chessbench.ai/support
Every dollar goes to API costs. More funding means more models tested, faster, with more games behind each rating.
Muse Spark 1.3 (and 1.2 and 1.1) have been added to ChessBench.
@alexandr_wang, @AIatMeta not bad
Gemini 3.8 Flash is a monster on the chess board!
@GoogleDeepMind @OfficialLoganK nice
Claude Fable 5.1 has been added to ChessBench. Results may be surprising... @AnthropicAI what's up
Replying to @AnthropicAI
@AnthropicAI's Claude Opus 5 has been added to ChessBench! You may be surprised by the results...
Calling GPT-5.6 on the API? Check your bill closely.
In ChessBench, every move is one API call, so the cost of each is easy to see. OpenAI is billing me for far more than the models actually generate. One example of many: on a hard position, GPT-5.6 used about 21k reasoning tokens. I was billed for 466k.
I checked it against my OpenAI dashboard and the inflated numbers match to the cent, so it's what's charged, not a display bug. Another developer reproduced it. It also quietly hits o1, o3-mini, and o4-mini at 2x.
Anyone else seeing this? @OpenAIDevs