A PyTorch policy network, cloned then trained by playing
A small convolutional network copies a two-move search, then improves through 16,384 games. It finished about 40% short of its teacher.
rejectedrecordedEach page here is one theory of how to choose a column, grouped by the technique it uses; the engines that play the games and the instruments that measure them live under Engines and Diagnostics.
Train a network that picks columns directly, nudging it toward the choices that led to longer games.
Read the primer →No approach page matches that search.