D4 flow audit
Replays the reference on used games and records each board, decision, and search value for behavioral analysis.
Measuring instruments, harnesses and probes; none of these is a way to play.
Tools that take a reading from games already played or from positions: where the points come from, which discs can never clear, how fast discs must clear to keep up, and what a perfect-information solver would have done.
Replays the reference on used games and records each board, decision, and search value for behavioral analysis.
When two columns are worth the same, their order chooses between them. This measures the effect over a whole game.
Runs a player that can see hidden numbers beside public-information players on the same games, then compares their boards every five moves to show what self-sustaining play looks like.
Measures how often low numbered discs become impossible to clear, when death follows, and whether the evaluator notices.
Lets a privileged planner test whether any policy can remove discs as fast as the game adds them.
Attributes every point to its source and shows that Hardcore score mostly measures survival time.
Measures whether the boards a clairvoyant oracle chooses to stand on are still good boards when the future is replaced by ordinary public randomness, by re-valuing them and matched fair-search boards under the same 32 independent public futures.
Show a planner the whole future to measure what the simulator makes possible. It is never a deployable policy.
The fixed scaffolding that runs policies side by side under one set of conditions, and the checks on the benchmark itself.
Runs simple hand-written players on paired games so differences reflect policy rather than luck.
Plays complete games with the shared hand-written policy and reports score, clears, reveals, board height, search work, and early stops.
The shared harness that makes two hand-built policies play the same games, so that a difference between them is about the policies and not about which one got luckier discs.
Ranks nine policies on 128 saved positions, compares that ranking with complete games, and retires the benchmark when they disagree.
A locked set of 477 positions labels every legal column through long play-forwards, creating a reusable move-ranking test bench.
Small experiments that ask what a network of a given size can hold, before a training run is paid for.
Asks whether telling a simple player to care much more about breaking open gray discs makes it live longer, and finds that turning that dial up barely moves anything.
Retrains the small board evaluator at several sizes and settings to test whether more capacity predicts lifetime better.
Trains the smallest network that could run inside a search leaf on the exact values the four-ply search gives every sibling move, and asks whether it ranks those moves the way the search does.