Interactive · Plain JavaScript

Self-Taught Connect Four

A pocket-sized AlphaZero, written from scratch. Nobody taught it a single strategy: it learns only by playing itself, and you can watch it happen in your browser. Then try to beat it.

01

Play against it

You are black, it is blue. The bars above the board show its thinking: the pale bar is its instinct from the network alone, the solid bar is where its search ended up.

Instinct After search

Its read on the game
You are winningIt is winning

. At zero it plays on instinct alone; each simulation looks one line deeper into the future.

Its brain

02

Watch it learn

Start from random weights and let it play itself. Every 25 games it is tested against three fixed opponents, searching 200 simulations per move. Random picks any legal column. Greedy takes a win when it can, blocks yours when it must, and otherwise favours the centre. Classic MCTS runs the same tree search, also with 200 simulations, but judges positions by playing random games to the end instead of asking the network. Random and Greedy fall within minutes; classic MCTS is the real test. Switch its brain to "Your training run" above to play what it has learned so far.

Self-play games0
Positions seen0
Games / min-
vs Classic MCTS-
Live self-play

Self-play appears here

Score against fixed opponents · smoothed
vs Random vs Greedy vs Classic MCTS
Training loss
Policy Value
03

How it works

AlphaZero combines a neural network with Monte Carlo tree search and trains the network on games it plays against itself. This page implements a small version of the method for Connect Four.

01

Policy and value network

A fully connected network with two hidden layers takes the board as input. It outputs a policy, a probability for each of the seven columns, and a value, an estimate of the outcome for the player to move. The weights are initialized randomly.

board → (p₁…p₇, v) · - weights
02

Tree search

Monte Carlo tree search builds a tree of possible continuations. At each node it selects the move with the highest PUCT score, which weighs the move's average value (Q) against its policy prior (P) and its visit count (n). New positions are evaluated with the value output rather than played out to the end.

pick argmax Q + c·P·√N / (1 + n)
03

Self-play training

Each self-play game produces training examples: the position, the visit distribution of the search (π) and the final result (z). The network is trained to match π with its policy and z with its value, and the updated network plays the next games.

loss = (z − v)² − π · log p