WUZIQI — Gomoku
Beauty is reason enough
← Back to blog

Engines · Evaluation

How Gomoku Engines Judge a Position

Engines still search through possible moves. But an evaluation function decides which positions look promising. Here is what changed as engines moved beyond hand-written pattern counts.

An ink-wash study of a Gomoku board: before scoring a position, an engine must account for shapes that are still forming and already connected.
An ink-wash study of a Gomoku board: before scoring a position, an engine must account for shapes that are still forming and already connected.

Black places a stone near the middle of the board. White faces no four-in-a-row that demands an immediate block, but two diagonals are beginning to converge. The engine favors a defensive move away from their intersection. If you count only the patterns already on the board, you may wonder what it sees.

When search stops, evaluation takes over

An engine searches ahead through candidate moves and possible replies. In alpha-beta search, for example, it prunes branches that no longer need checking. It still cannot follow every branch to the end of the game.

When search stops at a position, an evaluation function scores that leaf node so the engine can compare moves higher up the tree. Black’s diagonal may pose no immediate threat, yet the score might suggest that White will struggle to cover both sides later. Search uses that estimate to decide whether the line is worth pursuing.

The score estimates the current position; the moves to come will decide the game. The earlier a search stops, the more its evaluation must recognize connections that have not yet become threats but will shape later choices.

What a hand-written pattern table sees

A traditional evaluation recognizes shapes such as an open three or a four-in-a-row, then adds up preset weights. The logic is legible: a four-in-a-row usually demands attention before a shape that needs another preparatory move. An engine can use that distinction to check defensive moves first.

Consider a 15×15 freestyle position with White to move. Black has a diagonal shape in the upper left and a horizontal one on the right, both extending toward the empty tengen. A White stone on tengen answers both. If a hand-written table scores the shapes separately and assumes each requires its own defense, it counts the defensive cost twice and overvalues Black’s position. Such a table is also limited to shapes its author anticipated. Search can catch some omissions, but only after expanding the relevant branches. If an early evaluation pushes a crucial move aside, the engine may never spend enough time on that line.

A diagonal shape in the upper left and a horizontal shape on the right extend toward the same empty central intersection. The board has no text.
On an ink-wash board, two quiet shapes point toward one empty intersection. White can answer both by playing there.

From written rules to trained judgment

NNUE stands for “efficiently updatable neural network.” According to public accounts of NNUE’s origins in the shogi engine YaneuraOu, Yu Nasu invented it in 2018 and first used it in YaneuraOu. Its name also echoes the pronunciation of the nue, a creature of Japanese legend. NNUE lets an engine update a trained evaluation relatively quickly after each move—a useful property when search must score positions again and again.

Chess offers a useful comparison. According to public engine records for Stockfish NNUE, by August 2020 it searched at roughly half its former speed yet played at least 80 Elo stronger than it had with traditional evaluation. It was then adopted by the official engine. A better judgment at the leaves can sometimes outweigh a slower search.

Evaluation can change; search still has to test the line

A learned judgment still needs testing

A trained evaluation can learn how patterns work together across many positions. It remains an estimate. If White has a four-in-a-row that demands a reply, Black’s apparent positional advantage may vanish quickly. The engine still has to search through that forcing sequence.

Search methods divide the work differently. PVS fully searches its preferred move first, then tests other moves with a narrow window, searching again when necessary. Monte Carlo tree search devotes more simulations to promising branches. For sequences of continuing threats such as VCT, some engines use proof-number search to check whether a forced attacking line exists.

A move can enter the shortlist on the strength of its evaluation, then fall away when search finds a concrete reply. When an engine changes its mind, look first for the move that changed the threats. Comparing the two scores alone tells you less.

An ink-wash study of a board: stones near the center suggest several uneven lines of play, with open space at the edges and no text or labels.
Stones near the center suggest several lines of play; the empty edges leave room for the position to unfold.

Different combinations at Gomocup 2026

The engine updates on the official Gomocup 2026 results page describe several combinations. Figrid pairs alpha-beta search with NNUE evaluation and uses a threat cache and continuation history to order moves. Chloris trained its own NNUE on about three million Standard-15 positions; alongside search, it has checks for forcing moves and a layer for urgent defense.

The same official Gomocup 2026 page notes that TKGomoku retains a hand-written pattern evaluation while also using a lightweight NNUE, a policy file and an opening book. DDQK-Conquer uses an NNUE-like algorithm. Starpoint pairs Monte Carlo tree search with transformer evaluation and uses proof-number search internally for VCT. Trained judgment appears in several architectures, but not in the same role.

Nor can the results be credited to any one kind of evaluation. According to the official Gomocup 2026 results page, RAPFI retained its 20×20 freestyle title in 2026 and won both 15×15 freestyle and blitz. New entrant KATAGOMO finished second in both freestyle divisions, trailing by just 9 Elo in 15×15. Evaluation, search and time management all shape how an engine plays.

Reading a surprising recommendation

An engine may rank a quiet move ahead of one that immediately forms a three because the quiet move also restricts two future lines. A hand-written pattern table may see no high-scoring shape when the stone lands. A trained evaluation may give more weight to the relationship between the shapes. Whether that judgment holds up depends on the replies search finds.

Read a score in context. The rules, search depth, time available and evaluation method all affect it. Another engine may give the same position a different number and a different move order. Look at how both sides play after the recommendation, then ask whether the score gap holds.

Try this: Choose an engine-recommended move you would have passed over. First, write down which line you think it defends. Play a few moves along the engine’s variation, then return to the original position and look for the intersection your assessment missed.

Bring the score back to the board

When reviewing a game, find the threats that demand an answer before studying the more distant connections. If a move creates no immediate four-in-a-row but leaves the opponent able to defend only one side on the next turn, note when that constraint first arose.

Evaluation changes how an engine selects and judges positions. Search remains responsible for testing the moves that follow. Keep those jobs separate, and an unfamiliar recommendation invites two sharper questions: What does it value, and how far has it checked?

Check the forcing moves before you trust the score

Done reading? Open WUZIQI and play a round.