How FumbleChess works

Your opponent is a neural network that learned chess by imitation. It was shown tens of thousands of positions from real games and trained to guess which move the human played. On top of that it runs a small Monte-Carlo tree search, but the search is deliberately short and always guided by the human-trained network. The goal is not the strongest possible chess. It is play that feels like a person of a given rating, blunders included.

1 · From board to move

  1. See the boardThe position becomes 64 tokens, one per square. Each token is 22 numbers: 12 on/off flags (white pawn … black king), side to move, 4 castling rights, the en-passant square, two flags for a position that has already occurred once or twice (repetition), and the two move counters (the 50-move clock and the move number).
  2. Square encoderA learned layer turns those 22 numbers into a 256-number vector and adds a learned “I am square e4” embedding, so the network knows where each token sits.
  3. Local mixerTwo ConvMixer blocks slide a 3×3 filter over the 8×8 grid. Neighbouring squares matter in chess (pawn chains, a castled king), so they get mixed cheaply before the expensive part.
  4. TransformerEight layers of self-attention, 8 heads each. Every square can look at every other square, so a bishop can “see” down a diagonal to a far-away king. About 7 million parameters in total.
  5. Two heads Policy: a score for every from-square → to-square pair (64 × 64 = 4,096 possible moves).
    Value: three numbers: win, draw and loss estimates.

2 · How it was trained

Real human games

From the Lichess open database (January 2024). A game is kept only if both players are inside the rating band. Then 10 random positions are sampled from each of 8,000 games, about 80,000 positions per model.

Three models, three bands

Beginner 800–1200, Club 1600–2400, Master 2800+. A 1000-rated and a 2900-rated player differ in which moves they choose, so each model imitates its own group. Master games are so rare that bullet games are included to find enough.

What it learns from

The policy head is trained to put probability on the move the human actually played (illegal moves are masked out). The value head is trained on the final result of the game. Each model trained for 1,500 steps, about 4 minutes on a laptop GPU.

3 · Thinking ahead: tree search

A single glance at the board gives the AI a gut feeling: which moves look natural, and who seems to be winning. The search lets it test that feeling by playing a few lines out in its head. It builds a tree of positions, one simulation at a time:

  1. Select. Starting from the current position, walk down the tree. At each step pick the move with the best mix of how good it has looked so far and how natural the network thought it was, with a bonus for moves that have been tried rarely (the “PUCT” rule).
  2. Expand and evaluate. When the walk reaches a position it has never seen, ask the network about it once. The policy head says how natural each legal move is (the prior). The value head says how good the position is. There are no random playouts to the end of the game.
  3. Back up. Pass that value back up the path, flipping its sign at every ply because what is good for one player is bad for the other.
  4. Repeat, then choose. After the budget is spent the AI samples its move from how often each one was visited. It does not just take the most visited, so it never plays the same game twice.

Why so little search? The budget depends on the level: about 24 simulations for Beginner, 96 for Club and up to 180 for Master, with a time cap. With so few simulations the visit counts stay close to the network's human-like priors, so the AI keeps its human habits. The search mostly nudges it away from moves its own value head dislikes. Blunders still happen, on purpose, for three reasons: it samples instead of always taking the best, its value head is a rough judge, and it only sees a few moves ahead.

The AI thoughts panel is this search, live. Each row is a candidate move: the bar and percentage show its share of the simulations, and the small number on the right is the AI's expected score after that move (0–100). Arrows on the board get thicker for moves searched more. When it decides, the chosen move turns red. The coloured strip is the value head's first impression of the position, before any searching. Hover a row to see the move's prior probability.

4 · The evaluation bar is a different brain

The bar uses Stockfish, a classical engine that calculates: it searches for 300 ms and scores the position in centipawns (100 = one pawn). That score is converted into an expected result for White, which fills the bar; forced mates show as M3. So the AI plays like a human, while the bar judges like an engine. When they disagree, that is usually a human-style mistake.

Show the maths

Priors. The policy head gives each legal move i a score si. They become prior probabilities with a softmax: Pi = exp(si) / Σj exp(sj).

Value. The value head outputs win, draw and loss probabilities for White. For the side to move, v = P(win) − P(loss), a number between −1 and 1. A move's expected score is (v + 1) / 2.

Selection rule (PUCT). At a node with N visits, choose the move that maximises Q + c · P · √N / (1 + n), where Q is the average value of that move so far (from the mover's point of view), P its prior, n its own visit count and c = 1.5. Untried moves start slightly pessimistic (their parent's value minus 0.2).

Final choice. With visit counts ni at the root, the move is drawn with probability ni1/T / Σj nj1/T. T is 1.0 for Beginner, 0.8 for Club and 0.6 for Master: a lower T leans harder on the most-visited move.

Speed. Positions are evaluated 8 at a time (with “virtual loss” so the 8 walks go to different places). Each position costs about 20 ms in WebAssembly on one thread, which is why the budgets are small.

Centipawns to bar fill. E = 1 / (1 + 10−cp / 400): the standard Elo logistic. 0 cp gives 50%; +100 cp gives about 64%; +400 cp gives about 91%.

Policy size. 64 from-squares × 64 to-squares = 4,096. The head builds a query vector and a key vector for every square and scores each from/to pair by their dot product. Promotions are always to a queen.

Training loss. Policy: cross-entropy between the softmax over legal moves and the human move. Value: cross-entropy between the predicted win/draw/loss and the game result.

5 · It all runs in your browser

The models run with ONNX Runtime Web (WebAssembly) and Stockfish runs in a Web Worker. Nothing is sent to a server. Your game, level, colour and theme are kept in your browser's local storage.

Credits

Ideas and research

Data

Engine

  • Stockfish by the Stockfish developers (GPLv3).
  • Stockfish.js by Nathan Rugg and Chess.com (GPLv3), the WebAssembly build used for the evaluation bar.

Software

  • Chessground, the board, by the Lichess team (GPL-3.0-or-later).
  • chess.js by Jeff Hlywa (BSD-2-Clause), game rules.
  • python-chess by Niklas Fiekas (GPL-3.0), data preparation.
  • PyTorch (Paszke et al., NeurIPS 2019), model and training.
  • ONNX Runtime Web by Microsoft (MIT), running the models in the browser.
  • Vite (MIT) and TypeScript (Apache-2.0).

Pieces

  • Piece set “cburnett” by Colin M.L. Burnett (GPL-2.0-or-later), shipped with Chessground; board colours from its brown theme.

FumbleChess is an educational project. It is not affiliated with Lichess, Chess.com or the Stockfish team.