In brief
Toby is a chess engine that predicts human moves, and the result of each move, for players of a given rating and time control. He uses both Maia (intuition) and Leela (calculation). Toby’s evaluations adapt not only to the position but also to the ratings of both players and to the time control, are shown as WDL (win / draw / loss), and are overall much more accurate than Maia’s. Toby is free and open source (AGPL v3), and works in chess GUIs like ChessBase or Nibbler.
- On 1,500 classical test games, Toby finds the move actually played in 58.7% of positions, against 52.4% for Maia-3 at the player’s real rating.
- Toby blunders as often as real players do (3.5% of moves), where Maia with the rating boost would play more than twice as many blunders (8.9%).
Table of contents
- Author’s note
- Where Maia goes wrong, and what Toby changes
- Does Toby really work?
- Oracle’s little brother
- Thinking fast, thinking slow: Maia and Leela
- How Toby predicts the next move
- How Toby evaluates each move
- Limits
- Using Toby
- Name, licence, cost
- What Toby can be used for
- Download
- Support and donations
Author’s note
I am Yosha Iglesias, a FIDE Master and Woman International Master, and I’m not a developer. I came up with the concept, the logic, and the tests; Claude Sonnet 5.5 turned them into code and helped me with the calculations.
This article has no scientific pretensions. But nothing is hidden: Toby was tested only on games he had never seen, with the rules of the test fixed in advance, and the method page gives every result, including the tests he failed. Toby is a big step forward in predicting human moves, but he has limits, listed at the end.
Where Maia goes wrong, and what Toby changes
Before Toby, Maia was considered the best existing predictor of human moves. Toby starts from her predictions and corrects them with Leela. Yakubboev (2685) against Sarin (2718) was the game that decided the 2026 Olympiad. Two positions from it show what Toby changes. For each move, the figure gives the probability that a player of that level plays it, and the expected result (win / draw / loss, from White’s point of view).
Move 37. Sarin blundered with the natural but mistaken 37…Nf5??. Maia gives this move 91.6%, and she thinks it is winning for Black (65% wins). Toby thinks the likeliest move for a 2718-rated player is the correct Kf6! (51%; the also correct Kh6 is much less likely according to Toby). He still gives 37…Nf5?? 42%, but he understands that it is in fact losing for Black (60% wins for White, 11% for Black). Stockfish gives the usual and useless 0.00 before 37…Nf5?? (and +4.8 after it), where Toby explains that the position is dangerous for Black, with a predicted score of 61% for White.
Move 38. Maia gives the mistaken gxh4+ as the likeliest move (87%) and almost ignores the winning 38.Ne6+!! (0.1%), which she thinks is losing for White. Toby gives 96% that a player of Yakubboev’s level finds Ne6+, and correctly shows that it is winning (86% wins for White).
You probably wonder: « Sure, this looks great, but does Toby really work? »
Does Toby really work?
All the numbers below come from test games: 1,500 classical games, and about 1,400 rapid and 1,400 blitz games (over the board and online).
How often does the engine’s top choice match the move actually played? On the 1,500 classical test games, Toby finds the move actually played in 58.7% of positions, against 52.4% for Maia-3 at the player’s real rating. The gain is large in slow games, smaller in rapid, and small in blitz.
In the charts, « Maia » is Maia with the rating boost that Toby gives her: the best version of Maia, better than at the player’s real rating.
The gain depends on the level. The correction brings more as the players get stronger, in classical and in rapid. In blitz it stays small.
Does he play like real players? The AESL (average expected-score loss) is the number of points of expected score a move gives away, on average, compared with the best move according to Leela. Toby blunders as often as real players do (3.5% of moves in classical), where Maia with the rating boost would play more than twice as many blunders (8.9%).
| AESL (points per move) | Real players | Maia (rating boost) | Toby |
|---|---|---|---|
| Classical | 3.0 | 5.8 | 3.1 |
| Rapid | 3.3 | 5.8 | 3.5 |
| Blitz | 3.7 | 5.3 | 3.9 |
What does this mean in practice? If you drew each move at random from Toby’s probabilities, you would globally play at the level intended: about as many mistakes as real players of that rating. If you drew them from Maia’s probabilities, you would play at a much lower level, because Maia has no calculation and gives too much probability to blunders.
It also works beyond the first choice: in classical games, the move that was actually played gets about 10% more probability from Toby than from Maia with the rating boost.
So, Toby works. But how does he work? To answer that, let’s start with where he comes from.
Oracle’s little brother
In 2024, I created Oracle, a module that combined a language model (GPT), used as a stand-in for chess intuition, with a calculation engine (Stockfish). It predicted human moves with 56–63% accuracy, against 52% for Maia alone. Oracle had limits, which Toby removes:
- it depended on a paid service (GPT-3.5 turbo-instruct, now out of service);
- there was no UCI version, so no direct use in a chess GUI;
- it gave evaluations as expected score only, not as win / draw / loss.
Toby is Oracle’s little brother: the same idea, but free, local, and adjustable.
Thinking fast, thinking slow: Maia and Leela
Garry Kasparov summed up the thinking process of grandmasters in one sentence: « It’s intuition first, then calculation. » In Thinking, Fast and Slow, Daniel Kahneman describes two ways of thinking: System 1 is fast, automatic, and effortless (intuition); System 2 is slow, deliberate, and laborious (calculation). This is an analogy, a theoretical simplification, but a useful one for describing and predicting human behaviour in general, and in chess in particular. As with Oracle, I use the same idea to build an AI predictor of human moves.
- Maia-3 is System 1. She was trained on millions of human games. Think of her as a super super-grandmaster to whom you give a position and the level of the players: her task is to predict the next move and the expected result, but she is not given time to think and can only rely on intuition.
- Leela is System 2. She calculates and finds the best move, but she plays like a 3600-Elo player: far too well to predict a human, especially at low level. She also predicts the result, but for herself against herself.
- Toby has both. Toby combines his System 1 and his System 2, like a human player does. He is a super super-grandmaster asked to predict the next move of a human, and he can take his time to think and calculate (though he does it super fast!).
How Toby predicts the next move
- Maia’s intuition. For each legal move, the probability that a player of the given rating would play it.
- Leela’s calculation. For each move, how much expected score it loses compared with the best move.
- Toby’s correction. One rule: each move loses part of its probability for every point of expected score lost, with a severity that grows with the rating; Leela’s best move gets a small bonus. Then everything is renormalized.
There are three modes (classical, rapid, blitz), each with its own rule. Below 2200 in blitz and below 2100 in rapid, Toby plays exactly like Maia.
The exact rule and its constants are on the method page.
How Toby evaluates each move
Toby gives his evaluations in WDL (win / draw / loss) rather than in centipawns, for a simple reason.
There are only three objective or theoretical evaluations in chess. A position can be winning for White, a draw, or winning for Black. These evaluations are essential properties of the positions they refer to and do not depend on the context in which the position has arisen.
All other evaluations are subjective or practical: they are only valid for some player against another player, in certain conditions. They can be expressed in:
- phrases (« White is probably winning but Black has drawing chances »);
- signs (« +/-« );
- centipawns (« +1.00 »);
- expected score (« 75% »);
- WDL (« 700 100 200 »).
Leela’s WDL are not useful here: they are the evaluation of a player rated about 3600 against herself. Maia’s value head is closer to what we want but it can be very mistaken, especially at high level or in classical chess. Toby gives an evaluation for a human player against a human player, taking into account:
- the rating of the player to move and the time control;
- the rating of the opponent, when you give it;
- whether the games are played over the board (OTB) or online. At equal strength, there are far more draws OTB than online (Toby knows which one from the rating source you choose).
Here is the same position, evaluated for different players and time controls. For each move, the figure gives Maia’s probability, Toby’s probability, and Toby’s win / draw / loss from White’s point of view. The Position line is the evaluation of the position itself: the evaluations of the moves, weighted by the probability of each move.
Can you trust Toby’s evaluations? This chart compares what Toby announces with what really happened in the test games, for each time control (the closer to the dotted line, the better):
Limits
Toby is not perfect, and I want you to know where he is not:
- There is probably room for improvement. Even keeping the idea of a System 1 and a System 2, nothing says that Maia (not trained on slow games) plus Leela is the best possible combination. A better predictor might not be built on this logic at all.
- Blitz: the gain is small (+0.7 point) and was not demonstrated on the pre-registered test (Lichess games only).
- Online rapid and blitz evaluation: Maia alone is almost as well calibrated as Toby. The final test failed on this point, before I separated OTB and online games; the corrected version was validated on a second set of fresh games.
- Puzzle-like positions (when the best move is much better than the second best): Toby underestimates the player to move by 6 to 8 points (still better than Maia, who underestimates by 12 points).
- A short calculation: Leela calculates with 1,024 nodes (two to three seconds). In some positions, more nodes would refine the result, but on the test games a larger search (4,096 nodes) did not predict better, neither the moves nor the evaluations, so Toby keeps 1,024.
- Not fully deterministic (because Leela isn’t): the same position can give numbers that differ by a point or two from one analysis to the next.
Other, less important limits can be found on the method page.
Using Toby
Toby is a single engine, with a display style for ChessBase, one for Nibbler, and one for other GUIs, and it picks the right one by itself. Unfortunately, ChessBase’s interface is not designed for engines like Toby, and Nibbler, while better suited, does not let an engine customize its display. The options are:
- White Rating and Black Rating: the ratings of the two players. If you fill only one, both sides get that rating (Toby then assumes the same level for the opponent). If you fill both, Toby also uses the opponent’s rating.
- Rating Source: where the ratings come from (FIDE, Lichess Classical / Rapid / Blitz, Chess.com Rapid / Blitz).
- Time Control: classical, rapid, or blitz.
In ChessBase, open the engine’s parameters and fill in the two rating fields (once per game), plus the rating source and the time control. The figure below shows how to read the lines.
In Nibbler, I provide small scripts (in the download, folder nibbler_scripts) that you can copy into Nibbler’s scripts folder: Toby 1 sets the time control, Toby 2 White sets White’s rating, Toby 3 Black sets Black’s, so you can set the right ratings and time control in three clicks.
Name, licence, cost

The name is Toby, from The Oracle By Yosha. The licence is AGPL v3 (Maia-3 is AGPL v3; Leela is GPL v3). Toby is free. The default download is a single package that contains everything (Toby, Maia, Leela, and Leela’s network; about 1.4 GB): unzip it and add it to ChessBase or Nibbler. The Leela in the package needs an NVIDIA graphics card (it was tested on one). If you already have Leela and Maia, lighter packages are available: Toby alone (about 0.3 GB), or Toby with Maia’s weights (0.6 GB) or with Leela and its network (1.1 GB). You can then point Toby to your own versions; the options are described in the repository.
What Toby can be used for
Toby could be used for broadcasts, training, opening preparation, anti-cheating, bot creation, etc. Contributions are welcome.
Download
github.com/Yosha87/Toby: code, instructions, and licence (AGPL v3).
Support and donations
Between Oracle and Toby, I have dedicated hundreds of hours of work and had to invest money. While I am happy to offer Toby to the chess community for free, any donation would be greatly appreciated.
If you value my work and wish to support it, please consider making a donation.
Method page: how the numbers were obtained, with every test result, including the failed ones.









