During a game review, you note the move the engine calls best. In your next game, a familiar shape appears, so you play it. Your opponent replies differently from the line you reviewed, and soon you are lost. The move may not be the problem. You remembered the point, but not the continuation it depended on.
Tournament strength is not a move score
The 27th Gomocup took place June 5–7, 2026. In the freestyle division, RAPFI25 finished first with an Elo rating of 2495 and an aggregate score of 349:11. KATAGOMO26 was second at 2290 and 324:36. Their head-to-head score was 19:5. Results like those make it tempting to trust every point the winner recommends.
Under standard rules, RAPFI25’s Elo was 2121 and KATAGOMO26’s was 2112; their head-to-head score was 7:7. The gap between the same two programs looked different under different rules. Change the rules, and the conditions you need to check in a review change too. A program’s name alone cannot tell you how sound a particular move is.
Tournament Elo describes a program’s overall strength across a set of games. It is not a score for one move. The position evaluation on a review screen comes from the engine’s assessment and search. Its meaning depends on the rules, search depth and opponent replies considered. There is a missing test between “this move scores slightly higher” and “I can handle what happens after I play it.”
A recommendation is not an explanation
Suppose the engine ranks a quiet move ahead of a four-in-a-row. It may have found that the four-in-a-row gives your opponent a convenient exchange, while the quiet move preserves your options for the next turn. The interface usually gives you a ranking, a score or a win rate. Without the variations, none of those numbers tells you which line needs defending.
If you play the recommended move and your opponent chooses a branch you never examined, you have to work it out over the board. Under the conditions of its search, the engine may still prefer its original choice. That conclusion cannot do the calculation for you. Find out what reply it expects before you carry the move into another game.
The CHI PLAY 2022 paper “(A)I Will Teach You to Play Gomoku” examines the gap between an AI that plays well and one that helps a person learn. The CHI PLAY 2025 paper “(A)I Can Play Gomoku” goes on to study and evaluate a teaching system that offers hints, feedback and review support. Both point to the same distinction: a playing AI supplies moves; a teaching layer must help you understand, practice and revisit them.
Strong play need not come with a rulebook
Cal Poly’s “Mastering the Game of Gomoku Without Human Knowledge” presents a player trained through self-play without relying on human knowledge. It can find effective moves, but that does not mean it has produced a set of principles for you to memorize. If you copy only coordinates from a review, you take away one search result, not a way to make the judgment yourself.
The Rapfi paper describes another route to strength: with limited computing power and no GPU, it combined a distilled neural network with tuned evaluation and search to surpass Katagomo, then the strongest open-source AlphaZero-style engine. That result shows how much a move can owe to carefully arranged computation. To learn from it, you still have to examine the move in its actual variations.
Turn the recommendation into a question
An engine is more useful as a reference for variations than as a teacher whose top choice you simply obey. When you open the candidate moves list, start with moves two through five. What does each defend, and what does each concede? If you can explain one candidate by pointing to two lines on the board, you have something to compare with the top choice.
Then change the search depth and run the same position again. Keep the moves that rank near the top both times, and examine your opponent’s strongest reply to each. Mark moves whose rankings shift sharply for further study. A depth check cannot prove a move wins, but it can stop you from treating a tiny edge in a short search as a verdict.
As you study a variation, replace “What is its score?” with questions about the board. Which line must your opponent address after this move? How many points can they use to address that line? Those questions make you separate a threat from the points that answer it. They may also reveal a reply you missed.
Count threats, then count the answers
Consider a coordinate example under freestyle rules, with no forbidden moves. Black has stones at H8, I8 and J8, and at K7, K9 and K10. Black now plays K8. H8–K8 horizontally and K7–K10 vertically each become an open four. Under Renju rules, which impose forbidden moves on Black, K8 would be a forbidden double-four and could not be played.
In this freestyle example, White faces two independent threats, one horizontal and one vertical. G8 and L8 are possible blocking points at the ends of the horizontal line; K6 and K11 are possible blocking points at the ends of the vertical line. Four coordinates do not mean four threats. Blocking one end leaves the other end open, let alone the other line. Count this way, and you can see how many problems a move asks your opponent to solve.
Back in your own review, confirm the rules first. Then mark each threat and the points that answer it. If a move creates only one easily neutralized threat, do not rush to memorize its coordinates, even if the interface gives it a slightly higher score. Look at what remains for you after your opponent replies.
Take away one judgment at a time
Near the end of a review, choose two candidates the engine considers acceptable and compare them. Ask: “Which leaves my opponent fewer viable replies?” That question only helps if you have checked the variations for both moves. Fewer replies do not, by themselves, make a move good. If you cannot explain the difference, put one of your opponent’s replies on the board and keep playing until the difference appears.
In your next game, practice only that judgment. When you see a similar shape, identify the independent threats first, then your opponent’s viable replies. Afterward, check what you overlooked. One judgment you understand is easier to use over the board than a string of recommended coordinates.
Compare the replies to two acceptable moves, one judgment at a time