Your opponent is a neural network that learned chess by imitation. It was shown tens of thousands of positions from real games and trained to guess which move the human played. On top of that it runs a small Monte-Carlo tree search, but the search is deliberately short and always guided by the human-trained network. The goal is not the strongest possible chess. It is play that feels like a person of a given rating, blunders included.
From the Lichess open database (January 2024). A game is kept only if both players are inside the rating band. Then 10 random positions are sampled from each of 8,000 games, about 80,000 positions per model.
Beginner 800–1200, Club 1600–2400, Master 2800+. A 1000-rated and a 2900-rated player differ in which moves they choose, so each model imitates its own group. Master games are so rare that bullet games are included to find enough.
The policy head is trained to put probability on the move the human actually played (illegal moves are masked out). The value head is trained on the final result of the game. Each model trained for 1,500 steps, about 4 minutes on a laptop GPU.
A single glance at the board gives the AI a gut feeling: which moves look natural, and who seems to be winning. The search lets it test that feeling by playing a few lines out in its head. It builds a tree of positions, one simulation at a time:
Why so little search? The budget depends on the level: about 24 simulations for Beginner, 96 for Club and up to 180 for Master, with a time cap. With so few simulations the visit counts stay close to the network's human-like priors, so the AI keeps its human habits. The search mostly nudges it away from moves its own value head dislikes. Blunders still happen, on purpose, for three reasons: it samples instead of always taking the best, its value head is a rough judge, and it only sees a few moves ahead.
The AI thoughts panel is this search, live. Each row is a candidate move: the bar and percentage show its share of the simulations, and the small number on the right is the AI's expected score after that move (0–100). Arrows on the board get thicker for moves searched more. When it decides, the chosen move turns red. The coloured strip is the value head's first impression of the position, before any searching. Hover a row to see the move's prior probability.
The bar uses Stockfish, a classical engine that calculates: it
searches for 300 ms and scores the position in centipawns (100 = one pawn). That score is converted into an
expected result for White, which fills the bar; forced mates show as M3. So the AI plays like a human,
while the bar judges like an engine. When they disagree, that is usually a human-style mistake.
Priors. The policy head gives each legal move i a score si. They become prior probabilities with a softmax: Pi = exp(si) / Σj exp(sj).
Value. The value head outputs win, draw and loss probabilities for White. For the side to move, v = P(win) − P(loss), a number between −1 and 1. A move's expected score is (v + 1) / 2.
Selection rule (PUCT). At a node with N visits, choose the move that maximises Q + c · P · √N / (1 + n), where Q is the average value of that move so far (from the mover's point of view), P its prior, n its own visit count and c = 1.5. Untried moves start slightly pessimistic (their parent's value minus 0.2).
Final choice. With visit counts ni at the root, the move is drawn with probability ni1/T / Σj nj1/T. T is 1.0 for Beginner, 0.8 for Club and 0.6 for Master: a lower T leans harder on the most-visited move.
Speed. Positions are evaluated 8 at a time (with “virtual loss” so the 8 walks go to different places). Each position costs about 20 ms in WebAssembly on one thread, which is why the budgets are small.
Centipawns to bar fill. E = 1 / (1 + 10−cp / 400): the standard Elo logistic. 0 cp gives 50%; +100 cp gives about 64%; +400 cp gives about 91%.
Policy size. 64 from-squares × 64 to-squares = 4,096. The head builds a query vector and a key vector for every square and scores each from/to pair by their dot product. Promotions are always to a queen.
Training loss. Policy: cross-entropy between the softmax over legal moves and the human move. Value: cross-entropy between the predicted win/draw/loss and the game result.
The models run with ONNX Runtime Web (WebAssembly) and Stockfish runs in a Web Worker. Nothing is sent to a server. Your game, level, colour and theme are kept in your browser's local storage.
FumbleChess is an educational project. It is not affiliated with Lichess, Chess.com or the Stockfish team.