Drop7 Research
← Docs
docs/research/status.md

Research Status

August 2026

The best dependable policy found so far is fair depth-4 expectimax — a four-move look-ahead that treats the game's luck honestly. Even with its best chance model, its measured means sit under half of that target. This page shows where the gap is, what has been measured, and which directions the evidence has closed.

How score responds to search depth, under each of the two chance models:

Mean score against search depth, fair expectimax, by chance-stratum count Mean score against search depth, fair expectimax, by chance-stratum count. All arms play the same 64-seed development cohort 0xa51d1000-0xa51d103f under corrected 17,000-point Hardcore scoring, except the depth-5 seven-stratum arm, which was stopped by decision after the first 32 seeds (0xa51d1000-0xa51d101f) and is drawn as its own one-point series so the two sample sizes are never joined by a line. Depth 2 was run only at seven strata. No confidence band is drawn on means; the paired contrasts with their bounds are in the companion bar figure. Evidence tier: development / public-development. The historical eight-game D3/D4 cohorts used 7,000-point scoring and are not shown. Sources: docs/exploratory/finding-05-chance-strata.md, RS-20260821T181917Z-9a34ba02, RS-20260821T205102Z-d89df4b5. Mean score against search depth, fair expectimax, by chance-stratumcount250,000300,000350,000400,000450,0002345Mean score (points)Search depth (plies)5 strata (biased next-disc estimate), Search depth 3 | depth 3, 5 strata | Mean score: 305,051 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · Cost | table, mean score5 strata (biased next-disc estimate), Search depth 3depth 3, 5 strataMean score: 305,051 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md · Costtable, mean score5 strata (biased next-disc estimate), Search depth 4 | depth 4, 5 strata (frozen reference) | Mean score: 297,327 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · | Confirmation cohort detail, mean score5 strata (biased next-disc estimate), Search depth 4depth 4, 5 strata (frozen reference)Mean score: 297,327 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md ·Confirmation cohort detail, mean score5 strata (biased next-disc estimate), Search depth 5 | depth 5, 5 strata | Mean score: 288,704 points | n = 64 games | source: RS-20260821T181917Z-9a34ba02 · metrics.d5s5MeanScore | (288,703.67)5 strata (biased next-disc estimate), Search depth 5depth 5, 5 strataMean score: 288,704 pointsn = 64 gamessource: RS-20260821T181917Z-9a34ba02 · metrics.d5s5MeanScore(288,703.67)7 strata (exact next-disc estimate), Search depth 2 | depth 2, 7 strata | Mean score: 265,294 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · Cost | table, mean score7 strata (exact next-disc estimate), Search depth 2depth 2, 7 strataMean score: 265,294 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md · Costtable, mean score7 strata (exact next-disc estimate), Search depth 3 | depth 3, 7 strata | Mean score: 312,327 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · Cost | table, mean score7 strata (exact next-disc estimate), Search depth 3depth 3, 7 strataMean score: 312,327 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md · Costtable, mean score7 strata (exact next-disc estimate), Search depth 4 | depth 4, 7 strata | Mean score: 398,498 points | n = 64 games | source: RS-20260821T181917Z-9a34ba02 · | metrics.d4s7ControlMeanScore (398,498.23)7 strata (exact next-disc estimate), Search depth 4depth 4, 7 strataMean score: 398,498 pointsn = 64 gamessource: RS-20260821T181917Z-9a34ba02 ·metrics.d4s7ControlMeanScore (398,498.23)7 strata, first 32 seeds only (arm stopped), Search depth 5 | depth 5, 7 strata (partial: 32 of 64 planned games) | Mean score: 411,874 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · metrics.d5s7MeanScore | (411,873.66)7 strata, first 32 seeds only (arm stopped), Search depth 5depth 5, 7 strata (partial: 32 of 64 planned games)Mean score: 411,874 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 · metrics.d5s7MeanScore(411,873.66)5 strata (biased next-disc estimate)7 strata (exact next-disc estimate)7 strata, first 32 seeds only (arm stopped)
Mean score against search depth under the five-stratum and seven-stratum chance models. Hover or focus a point for its value, bounds, cohort size and source record.
Source data

All arms play the same 64-seed development cohort 0xa51d1000-0xa51d103f under corrected 17,000-point Hardcore scoring, except the depth-5 seven-stratum arm, which was stopped by decision after the first 32 seeds (0xa51d1000-0xa51d101f) and is drawn as its own one-point series so the two sample sizes are never joined by a line. Depth 2 was run only at seven strata. No confidence band is drawn on means; the paired contrasts with their bounds are in the companion bar figure. Evidence tier: development / public-development. The historical eight-game D3/D4 cohorts used 7,000-point scoring and are not shown.

SeriesSearch depthMean scoreBoundsnSource
5 strata (biased next-disc estimate)3 depth 3, 5 strata305,051 points64docs/exploratory/finding-05-chance-strata.md Cost table, mean score
5 strata (biased next-disc estimate)4 depth 4, 5 strata (frozen reference)297,327 points64docs/exploratory/finding-05-chance-strata.md Confirmation cohort detail, mean score
5 strata (biased next-disc estimate)5 depth 5, 5 strata288,704 points64RS-20260821T181917Z-9a34ba02 metrics.d5s5MeanScore (288,703.67)
7 strata (exact next-disc estimate)2 depth 2, 7 strata265,294 points64docs/exploratory/finding-05-chance-strata.md Cost table, mean score
7 strata (exact next-disc estimate)3 depth 3, 7 strata312,327 points64docs/exploratory/finding-05-chance-strata.md Cost table, mean score
7 strata (exact next-disc estimate)4 depth 4, 7 strata398,498 points64RS-20260821T181917Z-9a34ba02 metrics.d4s7ControlMeanScore (398,498.23)
7 strata, first 32 seeds only (arm stopped)5 depth 5, 7 strata (partial: 32 of 64 planned games)411,874 points32RS-20260821T205102Z-d89df4b5 metrics.d5s7MeanScore (411,873.66)

Spec: web/content/figures/score-vs-depth.json · 3 source records

The same factorial as paired contrasts, which is the form the conclusions rest on:

Paired contrasts in the depth x chance-resolution factorial (mean delta, one-sided 95% lower bound) Paired contrasts in the depth x chance-resolution factorial (mean delta, one-sided 95% lower bound). Whisker is the one-sided 95% whole-game bootstrap lower bound; the upper end of the whisker is the point estimate itself (no upper bound is drawn). A contrast is significant when its lower bound clears zero: both stratum contrasts (7 minus 5 strata) do; neither depth contrast (5 minus 4 plies) does, and each sits below its own detection floor (47,052 at n=64, 107,988 at n=32). Cohort 0xa51d1000-0xa51d103f; the n=32 contrasts cover its first 32 seeds. Evidence tier: public-development. Sources: RS-20260821T205102Z-d89df4b5. Paired contrasts in the depth x chance-resolution factorial (mean delta,one-sided 95% lower bound)-100,000-50,000050,000100,000150,000at depth 4at depth 5at 5 strataat 7 strataPaired mean score delta (points)Paired contrast (same seeds, whole games)Chance resolution: 7 strata minus 5 strata — at depth 4 | d4s7 - d4s5, W-T-L 41-0-23, 3.82x work | Paired mean score delta: +101,171 points | 95% lower bound: 47,447 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD4s7MinusD4s5 (101,170.8 / 47,446.8)Chance resolution: 7 strata minus 5 strata — at depth 4d4s7 - d4s5, W-T-L 41-0-23, 3.82x workPaired mean score delta: +101,171 points95% lower bound: 47,447 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD4s7MinusD4s5 (101,170.8 / 47,446.8)Chance resolution: 7 strata minus 5 strata — at depth 5 | d5s7 - d5s5, W-T-L 19-0-13, 5.85x work | Paired mean score delta: +123,613 points | 95% lower bound: 32,575 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s7MinusD5s5 (123,612.7 / 32,575.2)Chance resolution: 7 strata minus 5 strata — at depth 5d5s7 - d5s5, W-T-L 19-0-13, 5.85x workPaired mean score delta: +123,613 points95% lower bound: 32,575 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s7MinusD5s5 (123,612.7 / 32,575.2)Depth: 5 plies minus 4 plies — at 5 strata | d5s5 - d4s5, W-T-L 33-0-31, 23.29x work (below 47,052 floor) | Paired mean score delta: -8,624 points | 95% lower bound: -55,134 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s5MinusD4s5 (-8,623.7 / -55,133.7)Depth: 5 plies minus 4 plies — at 5 stratad5s5 - d4s5, W-T-L 33-0-31, 23.29x work (below 47,052 floor)Paired mean score delta: -8,624 points95% lower bound: -55,134 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s5MinusD4s5 (-8,623.7 / -55,133.7)Depth: 5 plies minus 4 plies — at 7 strata | d5s7 - d4s7, W-T-L 17-0-15, 35.62x work (below 107,988 floor) | Paired mean score delta: +23,367 points | 95% lower bound: -83,046 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s7MinusD4s7 (23,366.8 / -83,046.2)Depth: 5 plies minus 4 plies — at 7 stratad5s7 - d4s7, W-T-L 17-0-15, 35.62x work (below 107,988 floor)Paired mean score delta: +23,367 points95% lower bound: -83,046 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s7MinusD4s7 (23,366.8 / -83,046.2)Chance resolution: 7 strata minus 5 strataDepth: 5 plies minus 4 plies
Paired contrasts from the same factorial: seven strata minus five strata, and the fifth ply minus the fourth, with one-sided 95% lower bounds. The source table under each figure names the record every point was copied from.
Source data

Whisker is the one-sided 95% whole-game bootstrap lower bound; the upper end of the whisker is the point estimate itself (no upper bound is drawn). A contrast is significant when its lower bound clears zero: both stratum contrasts (7 minus 5 strata) do; neither depth contrast (5 minus 4 plies) does, and each sits below its own detection floor (47,052 at n=64, 107,988 at n=32). Cohort 0xa51d1000-0xa51d103f; the n=32 contrasts cover its first 32 seeds. Evidence tier: public-development.

SeriesPaired contrast (same seeds, whole games)Paired mean score deltaBoundsnSource
Chance resolution: 7 strata minus 5 strataat depth 4 d4s7 - d4s5, W-T-L 41-0-23, 3.82x work101,171 pointslower 47,44764RS-20260821T205102Z-d89df4b5 metrics.pairedD4s7MinusD4s5 (101,170.8 / 47,446.8)
Chance resolution: 7 strata minus 5 strataat depth 5 d5s7 - d5s5, W-T-L 19-0-13, 5.85x work123,613 pointslower 32,57532RS-20260821T205102Z-d89df4b5 metrics.pairedD5s7MinusD5s5 (123,612.7 / 32,575.2)
Depth: 5 plies minus 4 pliesat 5 strata d5s5 - d4s5, W-T-L 33-0-31, 23.29x work (below 47,052 floor)-8,624 pointslower -55,13464RS-20260821T205102Z-d89df4b5 metrics.pairedD5s5MinusD4s5 (-8,623.7 / -55,133.7)
Depth: 5 plies minus 4 pliesat 7 strata d5s7 - d4s7, W-T-L 17-0-15, 35.62x work (below 107,988 floor)23,367 pointslower -83,04632RS-20260821T205102Z-d89df4b5 metrics.pairedD5s7MinusD4s7 (23,366.8 / -83,046.2)

Spec: web/content/figures/strata-5-vs-7.json · 1 source record

The simulator and reference searches are mature enough to support reproducible research, but the strategy problem is unsolved. Corrected-score fair depth-4 expectimax is the strongest dependable reference found so far. Its average score across 64 games is 308,296 points, far below the goal of a one-million-point average.

Exploratory work on 2026-08-20/21 reproduced that reference on fresh seeds and completed the depth-by-chance-resolution factorial. Two results from it strongly suggest what to do next:

  • The fourth ply is worth a large, significant gain — but only with an exact chance model. With the approximate five-stratum model the same fourth ply is worth slightly less than nothing. Chance resolution does not merely add points; it changes the sign of the depth gradient.
  • The fifth ply is not measurable by this design. Both depth-5 contrasts sit far inside their own detection floors, and the seven-stratum estimate changed sign when the cohort grew from 16 games to 32, so the supported statement is that any depth-5 effect is smaller than a 64-game paired cohort can resolve — not that it is zero. What can be said is economic: even the optimistic estimate buys its points at 35.6x the work per move.
  • A 64-game paired cohort in this game cannot see an effect below roughly 50,000 points. Every null result in the depth factorial is below its own detection floor and is therefore a non-measurement rather than evidence of no effect (RS-20260821T205102Z-d89df4b5). Choose experiments whose predicted effect exceeds the floor, or find a lower-variance estimator than complete games.

A fourth day, 2026-08-23, went to learned models and label economics — the reveal-construction probe in live play, the optimistic-state D0 gate, the leaf-cost NNUE student C0, and the P-SOL G0 label-semantics guardrail — and closed each of those four directions as tested; the research log for that day tells the story, and the closed-directions table below carries the headline numbers.

No candidate has qualified for the protected validation protocol. The frozen record marks both the protected and one-shot final cohorts as unopened. See the research status evidence for details.

How the research progressed

Seven eras, one sentence each; every configuration and outcome below is tabled row by row in the experiment index.

  1. A trustworthy simulator. TypeScript and native Hardcore rules were aligned and a scoring audit corrected the five-move level award from 7,000 to 17,000; results made with 7,000-point scoring remain historical Sequence-scored evidence.
  2. Hand-written policies and shallow lookahead provided fast baselines and showed that immediate score alone is a poor guide.
  3. Fair expectimax became the reference: completed fair D4 consistently improved on D3 in the corrected-score small cohorts, while deeper variants were not automatically better.
  4. Learning public-state values repeatedly hit sibling extrapolation: a model learned the outcome of the action that was played, then deployment asked it to rank actions it had not observed equally well.
  5. Oracle and long-outcome teachers found predictive signal, but students failed held-out sibling ranking or were too slow; oracle strength is an upper-bound teaching signal, not a legal policy result.
  6. Constructive cycles and explicit reservoirs look useful as features or options, but no tested controller displaced D4.
  7. Offline policy improvement around D4 on a locked panel of 477 public roots underperformed or barely overrode D4; neither ranker justified a gameplay run.

The failure mode that recurs through eras 4-7 is easier drawn than described:

Sibling extrapolation: training labels the played action; deployment ranks all seven Mechanism diagram: sibling extrapolation, the repeated failure mode of learned Drop7 policies. Panel 1 (training): a root board with seven legal columns; only the played action's afterstate carries a label (an H40 return); the other six successors are unlabelled and marked with question marks. Panel 2 (deployment): the model must rank all seven siblings; the six it never observed equally are outlined as off-distribution. Boards are 7x7, row 0 at top, columns 1-7 left to right; solid gray discs are covered, crack-marked gray discs are cracked; successor boards are drawn as the root plus the dropped disc (cascades not simulated). Explains docs/research/status.md section 4 and docs/exploratory/audit-05-optimistic-curriculum.md section 4, class (iii) sibling coverage / within-root discrimination (experiments 5, 9, 11, 13, 14, 17). No measured numbers are shown. 3 2 5 1 4 2 3 1 Sibling extrapolation: training labels the played action; deployment ranks all seven audit-05 §4 class (iii) — sibling coverage / within-root discrimination: 6 of 17 experiments TRAINING — only the played action gets a label row 0 1234567 next disc: 2 · 3 moves to rise col 1col 2col 3col 4col 5col 6col 7 label: H40 return ?????? DEPLOYMENT — the model must rank all seven siblings row 0 1234567 col 1col 2col 3col 4col 5col 6col 7 rank ?rank ?observedrank ?rank ?rank ?rank ? training covers the played action; deployment ranks all siblings. low value error on the visited successor does not rank the unvisited ones — docs/research/status.md §4 Root position — the public stateiRoot position — the public statevisible board, next disc, moves to rise; nothing else.7×7, row 0 at top, columns 1–7 left to right. Solid gray = covered;crack-marked gray = cracked (one hit taken). Data collection playedcolumn 3 from this root. Unplayed siblings — no labelsiUnplayed siblings — no label was ever collected for these six afterstates;any value the model assigns them is extrapolated from other roots.audit-05 §4 class (iii): experiments 5, 9, 11, 13, 14, 17.source: docs/exploratory/audit-05-optimistic-curriculum.md §4 Played action — the one labelled siblingiPlayed action (column 3) — the only sibling with a training label:an H40 return under D1 continuation (the afterstate label panel ofthe 2026-08-20/21 experiments). source: docs/research/status.md §7The other six columns were never played here, so they carry no label. Deployment ranks all seven — six are off-distributioniDeployment asks the model to rank all seven columns; six of the seveninputs were never observed equally — within-root discrimination fails.Even successor-closed coverage with every legal sibling labelled and exactsearch-value labels ranked worse than one-ply exact search.source: docs/research/status.md, directions closed, 2026-08-21 The observed sibling at deploymentiThe observed sibling: ranking signal exists here —it is the other six that are extrapolation.source: docs/research/status.md §4 Caption — the failure in one lineistatus.md §4: low value error on visited states did not guarantee goodroot-action ranking. Sibling coverage is necessary — and demonstrably notsufficient: with coverage complete, the obstacle moved to compact-evaluatorcapacity (status.md §7 and directions closed, 2026-08-21).
The era-4 failure mode: training labels the played action's successor, deployment asks the model to rank all seven siblings — six of which it never observed equally.
Source

diagram-sibling-extrapolation.svg — source and reading guide

Mechanism diagram D1 of runs/RUN-20260823T191900Z-b9f8f80d/kimi-k3-figure-plan.md ("Diagrams (mechanisms, not charts)"). Hand-written, self-contained SVG; no measured numbers are drawn.

What it explains

The repeated failure mode named in docs/research/status.md §4 ("Learning public-state values"): a model is trained on the outcome of the action that was actually played, then deployment asks it to rank all seven legal columns — six of which it never observed equally. This is failure class (iii) sibling coverage / within-root discrimination of docs/exploratory/audit-05-optimistic-curriculum.md §4, the largest class in the census: 6 of 17 learned-policy experiments (experiments 5, 9, 11, 13, 14, 17).

Element-by-element

  • Root board (both panels): a small invented 7×7 position, drawn in engine orientation (row 0 at top, columns 1–7 left to right). Blue circles are numbered discs; solid gray circles are covered discs; the gray disc with a crack mark is cracked (one hit taken). "next disc: 2 · 3 moves to rise" is the rest of the public state a legal policy may use (docs/methodology.md, information boundary).
  • Seven arrows: one per legal column, fanning from the root to the seven successor afterstates.
  • Solid accent arrow + "label: H40 return" tag (column 3): the action that was played during data collection. Its successor is the only one carrying a training label — an H40 return under D1 continuation, the afterstate label panel of the 2026-08-20/21 experiments (docs/research/status.md §7).
  • Dashed arrows + "?" tags: the six unplayed columns. No label exists for these afterstates; any value the model assigns them is extrapolation.
  • Deployment panel: the same fan, but now the model must produce a ranking of all seven. The six never-observed successors carry dashed red outlines and "rank ?" tags; the observed one is marked "observed".
  • Caption strip: the failure in one line. The second line quotes the status.md §4 conclusion that low value error on visited states did not guarantee good root-action ranking.
  • Hover/focus popovers (the fig-pt/fig-pop convention of the chart generator): drill-downs on the root board (public-state definition and board orientation), the labelled sibling, the unlabelled siblings, the deployment fan, and the caption.

Simplifications (stated explicitly)

  1. Successors are drawn as root + dropped disc. Cascades, reveals and gravity after the drop are not simulated; the diagram is about label coverage, not mechanics. The landing cell of the dropped disc is the accent square in each mini-board.
  2. The root position is invented, not a recorded board; no recorded board image exists in the repository for this purpose, and the mechanism does not depend on the position.
  3. "H40 return" names the label family of the afterstate experiments (status.md §7); the diagram does not assert that every class-(iii) failure used that exact label — the class spans played-action-only labels, noisy labels, and labels that cannot separate siblings (audit-05 §4 definition).
  4. The 2026-08-21 result quoted in the deployment popover (successor-closed coverage, every legal sibling, exact search-value labels, still worse than one-ply exact search) is from docs/research/status.md "Directions closed", included because it relocates the obstacle from coverage to capacity.

Sources

  • docs/research/status.md §4 (sibling extrapolation paragraph) and §7; "Directions closed" table (the successor-closed coverage row).
  • docs/exploratory/audit-05-optimistic-curriculum.md §4 — class (iii) definition and the counts table (6 of 17; experiments 5, 9, 11, 13, 14, 17).
  • docs/methodology.md — the public-state definition.
  • Figure spec: runs/RUN-20260823T191900Z-b9f8f80d/kimi-k3-figure-plan.md, D1.

Conventions

CSS variables with light-theme fallbacks (var(--fig-fg, #222), var(--fig-muted, #888), var(--fig-accent, #2563eb), --fig-grid, --fig-pop-bg, --fig-pop-border, --fig-cover, --fig-danger) so the diagram renders standalone and adopts the console's dark theme inside .research-fig (web/app/globals.css). Popovers are pure SVG/CSS (fig-pt + fig-pop), the same convention as web/content/figures/score-vs-depth.svg; they need no JavaScript.

Source: web/content/figures/diagrams/diagram-sibling-extrapolation.source.md

And every learned evaluator's ranking quality, against the exact search it tried to reproduce:

Learned evaluators versus exact search: top-1, pairwise and regret on held-out roots Learned evaluators versus exact search: top-1, pairwise and regret on held-out roots. Panels differ across points and are named per category: the afterstate models and the fair D4 / fair D1 comparators are gated on the H40 D1-continuation panel; the D4-value student on a D4-ordering panel; the planner-distill student and its comparator on the fair-planner H5K256 panel; martingale-dual B0 and its comparator on the locked 477-root H200 panel; exact D1 / exact D2 on the historical panel of RS-20260821T104500Z-77d21e90. Compare within a panel, not across. Fair D1 pairwise and regret on the afterstate panel are not recorded, and exact D1 / exact D2 are recorded top-1 only, so those series points are omitted. Full-train pairwise 0.658 is from the record's summary prose; its per-half values are not recorded. Per-point labels are omitted to keep the dot grid readable; every popover carries the value and source. All eight students sit below their panel's fair-D4 comparator column. Evidence tier: pilot (afterstate and D4-value line); see each source for its panel definition. Sources: RS-20260820T094500Z-5c1e9a04, RS-20260820T114500Z-2b7c9e31, RS-20260820T142500Z-8f4a2d17, RS-20260821T094500Z-1a7e3c55, RS-20260821T134500Z-4b9d2f68, RS-20260821T104500Z-77d21e90, docs/exploratory/finding-11-planner-distillation.md, docs/research/history.md. Learned evaluators versus exact search: top-1, pairwise and regret onheld-out roots00.20.40.60.8afterstate K=8 · H40afterstate K=64 · H40afterstate K=256 · H40afterstate full-train · H40afterstate D2-teacher · H40D4-value student · D4 panelplanner-distill · planner panelmartingale-dual B0 · H200fairD4 · H40fairD1 · H40exactD1 · historical panelexactD2 · historical panelfairD4 · planner panelfairD4 · H200metric value (unitless)evaluator (evaluation panel after the dot)top-1 accuracy — afterstate K=8 · H40 | metric value: 0.25 unitless | source: RS-20260820T094500Z-5c1e9a04 · metrics.modelTop1Pooledtop-1 accuracy — afterstate K=8 · H40metric value: 0.25 unitlesssource: RS-20260820T094500Z-5c1e9a04 · metrics.modelTop1Pooledtop-1 accuracy — afterstate K=64 · H40 | metric value: 0.34 unitless | source: RS-20260820T114500Z-2b7c9e31 · metrics.modelTop1Pooledtop-1 accuracy — afterstate K=64 · H40metric value: 0.34 unitlesssource: RS-20260820T114500Z-2b7c9e31 · metrics.modelTop1Pooledtop-1 accuracy — afterstate K=256 · H40 | metric value: 0.42 unitless | source: RS-20260820T142500Z-8f4a2d17 · metrics.modelTop1Pooledtop-1 accuracy — afterstate K=256 · H40metric value: 0.42 unitlesssource: RS-20260820T142500Z-8f4a2d17 · metrics.modelTop1Pooledtop-1 accuracy — afterstate full-train · H40 | metric value: 0.36 unitless | source: RS-20260821T094500Z-1a7e3c55 · | metrics.diagnosticTop1VsIter3top-1 accuracy — afterstate full-train · H40metric value: 0.36 unitlesssource: RS-20260821T094500Z-1a7e3c55 ·metrics.diagnosticTop1VsIter3top-1 accuracy — afterstate D2-teacher · H40 | metric value: 0.34 unitless | source: RS-20260821T134500Z-4b9d2f68 · metrics.rankingTop1Pooledtop-1 accuracy — afterstate D2-teacher · H40metric value: 0.34 unitlesssource: RS-20260821T134500Z-4b9d2f68 · metrics.rankingTop1Pooledtop-1 accuracy — D4-value student · D4 panel | metric value: 0.38 unitless | source: RS-20260821T104500Z-77d21e90 · metrics.top1Pooledtop-1 accuracy — D4-value student · D4 panelmetric value: 0.38 unitlesssource: RS-20260821T104500Z-77d21e90 · metrics.top1Pooledtop-1 accuracy — planner-distill · planner panel | metric value: 0.49 unitless | source: docs/exploratory/finding-11-planner-distillation.md · | Summary tabletop-1 accuracy — planner-distill · planner panelmetric value: 0.49 unitlesssource: docs/exploratory/finding-11-planner-distillation.md ·Summary tabletop-1 accuracy — martingale-dual B0 · H200 | metric value: 0.29 unitless | source: docs/research/history.md · Martingale-dual B0 ranking | audit table (line 4526)top-1 accuracy — martingale-dual B0 · H200metric value: 0.29 unitlesssource: docs/research/history.md · Martingale-dual B0 rankingaudit table (line 4526)top-1 accuracy — fair D4 · H40 | metric value: 0.5 unitless | source: RS-20260820T142500Z-8f4a2d17 · metrics.d4Top1Pooledtop-1 accuracy — fair D4 · H40metric value: 0.5 unitlesssource: RS-20260820T142500Z-8f4a2d17 · metrics.d4Top1Pooledtop-1 accuracy — fair D1 · H40 | metric value: 0.32 unitless | source: RS-20260820T142500Z-8f4a2d17 · metrics.d1Top1Pooledtop-1 accuracy — fair D1 · H40metric value: 0.32 unitlesssource: RS-20260820T142500Z-8f4a2d17 · metrics.d1Top1Pooledtop-1 accuracy — exact D1 · historical panel | metric value: 0.49 unitless | source: RS-20260821T104500Z-77d21e90 · metrics.referenceD1Top1top-1 accuracy — exact D1 · historical panelmetric value: 0.49 unitlesssource: RS-20260821T104500Z-77d21e90 · metrics.referenceD1Top1top-1 accuracy — exact D2 · historical panel | metric value: 0.57 unitless | source: RS-20260821T104500Z-77d21e90 · metrics.referenceD2Top1top-1 accuracy — exact D2 · historical panelmetric value: 0.57 unitlesssource: RS-20260821T104500Z-77d21e90 · metrics.referenceD2Top1top-1 accuracy — fair D4 · planner panel | metric value: 0.61 unitless | source: docs/exploratory/finding-11-planner-distillation.md · | Summary tabletop-1 accuracy — fair D4 · planner panelmetric value: 0.61 unitlesssource: docs/exploratory/finding-11-planner-distillation.md ·Summary tabletop-1 accuracy — fair D4 · H200 | metric value: 0.38 unitless | source: docs/research/history.md · Martingale-dual B0 ranking | audit table (line 4527)top-1 accuracy — fair D4 · H200metric value: 0.38 unitlesssource: docs/research/history.md · Martingale-dual B0 rankingaudit table (line 4527)pairwise accuracy — afterstate K=8 · H40 | metric value: 0.57 unitless | source: RS-20260820T094500Z-5c1e9a04 · | metrics.modelPairwisePooledpairwise accuracy — afterstate K=8 · H40metric value: 0.57 unitlesssource: RS-20260820T094500Z-5c1e9a04 ·metrics.modelPairwisePooledpairwise accuracy — afterstate K=64 · H40 | metric value: 0.64 unitless | source: RS-20260820T114500Z-2b7c9e31 · | metrics.modelPairwisePooledpairwise accuracy — afterstate K=64 · H40metric value: 0.64 unitlesssource: RS-20260820T114500Z-2b7c9e31 ·metrics.modelPairwisePooledpairwise accuracy — afterstate K=256 · H40 | metric value: 0.69 unitless | source: RS-20260820T142500Z-8f4a2d17 · | metrics.modelPairwisePooledpairwise accuracy — afterstate K=256 · H40metric value: 0.69 unitlesssource: RS-20260820T142500Z-8f4a2d17 ·metrics.modelPairwisePooledpairwise accuracy — afterstate full-train · H40 | metric value: 0.66 unitless | source: RS-20260821T094500Z-1a7e3c55 · summary (pairwise 0.658 | vs 0.685)pairwise accuracy — afterstate full-train · H40metric value: 0.66 unitlesssource: RS-20260821T094500Z-1a7e3c55 · summary (pairwise 0.658vs 0.685)pairwise accuracy — afterstate D2-teacher · H40 | metric value: 0.63 unitless | source: RS-20260821T134500Z-4b9d2f68 · | metrics.rankingPairwisePooledpairwise accuracy — afterstate D2-teacher · H40metric value: 0.63 unitlesssource: RS-20260821T134500Z-4b9d2f68 ·metrics.rankingPairwisePooledpairwise accuracy — D4-value student · D4 panel | metric value: 0.64 unitless | source: RS-20260821T104500Z-77d21e90 · metrics.pairwisePooledpairwise accuracy — D4-value student · D4 panelmetric value: 0.64 unitlesssource: RS-20260821T104500Z-77d21e90 · metrics.pairwisePooledpairwise accuracy — planner-distill · planner panel | metric value: 0.74 unitless | source: docs/exploratory/finding-11-planner-distillation.md · | Summary tablepairwise accuracy — planner-distill · planner panelmetric value: 0.74 unitlesssource: docs/exploratory/finding-11-planner-distillation.md ·Summary tablepairwise accuracy — martingale-dual B0 · H200 | metric value: 0.6 unitless | source: docs/research/history.md · Martingale-dual B0 ranking | audit table (line 4526)pairwise accuracy — martingale-dual B0 · H200metric value: 0.6 unitlesssource: docs/research/history.md · Martingale-dual B0 rankingaudit table (line 4526)pairwise accuracy — fair D4 · H40 | metric value: 0.74 unitless | source: RS-20260820T142500Z-8f4a2d17 · metrics.d4PairwisePooledpairwise accuracy — fair D4 · H40metric value: 0.74 unitlesssource: RS-20260820T142500Z-8f4a2d17 · metrics.d4PairwisePooledpairwise accuracy — fair D4 · planner panel | metric value: 0.8 unitless | source: docs/exploratory/finding-11-planner-distillation.md · | Summary tablepairwise accuracy — fair D4 · planner panelmetric value: 0.8 unitlesssource: docs/exploratory/finding-11-planner-distillation.md ·Summary tablepairwise accuracy — fair D4 · H200 | metric value: 0.67 unitless | source: docs/research/history.md · Martingale-dual B0 ranking | audit table (line 4527)pairwise accuracy — fair D4 · H200metric value: 0.67 unitlesssource: docs/research/history.md · Martingale-dual B0 rankingaudit table (line 4527)normalized regret — afterstate K=8 · H40 | metric value: 0.41 unitless | source: RS-20260820T094500Z-5c1e9a04 · metrics.modelRegretPoolednormalized regret — afterstate K=8 · H40metric value: 0.41 unitlesssource: RS-20260820T094500Z-5c1e9a04 · metrics.modelRegretPoolednormalized regret — afterstate K=64 · H40 | metric value: 0.3 unitless | source: RS-20260820T114500Z-2b7c9e31 · metrics.modelRegretPoolednormalized regret — afterstate K=64 · H40metric value: 0.3 unitlesssource: RS-20260820T114500Z-2b7c9e31 · metrics.modelRegretPoolednormalized regret — afterstate K=256 · H40 | metric value: 0.24 unitless | source: RS-20260820T142500Z-8f4a2d17 · metrics.modelRegretPoolednormalized regret — afterstate K=256 · H40metric value: 0.24 unitlesssource: RS-20260820T142500Z-8f4a2d17 · metrics.modelRegretPoolednormalized regret — afterstate full-train · H40 | metric value: 0.28 unitless | source: RS-20260821T094500Z-1a7e3c55 · | metrics.diagnosticRegretVsIter3normalized regret — afterstate full-train · H40metric value: 0.28 unitlesssource: RS-20260821T094500Z-1a7e3c55 ·metrics.diagnosticRegretVsIter3normalized regret — afterstate D2-teacher · H40 | metric value: 0.31 unitless | source: RS-20260821T134500Z-4b9d2f68 · | metrics.rankingRegretPoolednormalized regret — afterstate D2-teacher · H40metric value: 0.31 unitlesssource: RS-20260821T134500Z-4b9d2f68 ·metrics.rankingRegretPoolednormalized regret — D4-value student · D4 panel | metric value: 0.29 unitless | source: RS-20260821T104500Z-77d21e90 · metrics.regretPoolednormalized regret — D4-value student · D4 panelmetric value: 0.29 unitlesssource: RS-20260821T104500Z-77d21e90 · metrics.regretPoolednormalized regret — planner-distill · planner panel | metric value: 0.19 unitless | source: docs/exploratory/finding-11-planner-distillation.md · | Summary tablenormalized regret — planner-distill · planner panelmetric value: 0.19 unitlesssource: docs/exploratory/finding-11-planner-distillation.md ·Summary tablenormalized regret — martingale-dual B0 · H200 | metric value: 0.35 unitless | source: docs/research/history.md · Martingale-dual B0 ranking | audit table (line 4526)normalized regret — martingale-dual B0 · H200metric value: 0.35 unitlesssource: docs/research/history.md · Martingale-dual B0 rankingaudit table (line 4526)normalized regret — fair D4 · H40 | metric value: 0.18 unitless | source: RS-20260820T142500Z-8f4a2d17 · metrics.d4RegretPoolednormalized regret — fair D4 · H40metric value: 0.18 unitlesssource: RS-20260820T142500Z-8f4a2d17 · metrics.d4RegretPoolednormalized regret — fair D4 · planner panel | metric value: 0.13 unitless | source: docs/exploratory/finding-11-planner-distillation.md · | Summary tablenormalized regret — fair D4 · planner panelmetric value: 0.13 unitlesssource: docs/exploratory/finding-11-planner-distillation.md ·Summary tablenormalized regret — fair D4 · H200 | metric value: 0.28 unitless | source: docs/research/history.md · Martingale-dual B0 ranking | audit table (line 4527)normalized regret — fair D4 · H200metric value: 0.28 unitlesssource: docs/research/history.md · Martingale-dual B0 rankingaudit table (line 4527)top-1 accuracypairwise accuracynormalized regret
Every learned evaluator against its panel's exact-search comparator on held-out roots: top-1, pairwise accuracy and normalized regret. Panels differ across points and are named per category; compare within a panel, not across.
Source data

Panels differ across points and are named per category: the afterstate models and the fair D4 / fair D1 comparators are gated on the H40 D1-continuation panel; the D4-value student on a D4-ordering panel; the planner-distill student and its comparator on the fair-planner H5K256 panel; martingale-dual B0 and its comparator on the locked 477-root H200 panel; exact D1 / exact D2 on the historical panel of RS-20260821T104500Z-77d21e90. Compare within a panel, not across. Fair D1 pairwise and regret on the afterstate panel are not recorded, and exact D1 / exact D2 are recorded top-1 only, so those series points are omitted. Full-train pairwise 0.658 is from the record's summary prose; its per-half values are not recorded. Per-point labels are omitted to keep the dot grid readable; every popover carries the value and source. All eight students sit below their panel's fair-D4 comparator column. Evidence tier: pilot (afterstate and D4-value line); see each source for its panel definition.

Seriesevaluator (evaluation panel after the dot)metric valueBoundsnSource
top-1 accuracyafterstate K=8 · H400.25 unitlessRS-20260820T094500Z-5c1e9a04 metrics.modelTop1Pooled
top-1 accuracyafterstate K=64 · H400.34 unitlessRS-20260820T114500Z-2b7c9e31 metrics.modelTop1Pooled
top-1 accuracyafterstate K=256 · H400.42 unitlessRS-20260820T142500Z-8f4a2d17 metrics.modelTop1Pooled
top-1 accuracyafterstate full-train · H400.36 unitlessRS-20260821T094500Z-1a7e3c55 metrics.diagnosticTop1VsIter3
top-1 accuracyafterstate D2-teacher · H400.34 unitlessRS-20260821T134500Z-4b9d2f68 metrics.rankingTop1Pooled
top-1 accuracyD4-value student · D4 panel0.38 unitlessRS-20260821T104500Z-77d21e90 metrics.top1Pooled
top-1 accuracyplanner-distill · planner panel0.49 unitlessdocs/exploratory/finding-11-planner-distillation.md Summary table
top-1 accuracymartingale-dual B0 · H2000.29 unitlessdocs/research/history.md Martingale-dual B0 ranking audit table (line 4526)
top-1 accuracyfair D4 · H400.5 unitlessRS-20260820T142500Z-8f4a2d17 metrics.d4Top1Pooled
top-1 accuracyfair D1 · H400.32 unitlessRS-20260820T142500Z-8f4a2d17 metrics.d1Top1Pooled
top-1 accuracyexact D1 · historical panel0.49 unitlessRS-20260821T104500Z-77d21e90 metrics.referenceD1Top1
top-1 accuracyexact D2 · historical panel0.57 unitlessRS-20260821T104500Z-77d21e90 metrics.referenceD2Top1
top-1 accuracyfair D4 · planner panel0.61 unitlessdocs/exploratory/finding-11-planner-distillation.md Summary table
top-1 accuracyfair D4 · H2000.38 unitlessdocs/research/history.md Martingale-dual B0 ranking audit table (line 4527)
pairwise accuracyafterstate K=8 · H400.57 unitlessRS-20260820T094500Z-5c1e9a04 metrics.modelPairwisePooled
pairwise accuracyafterstate K=64 · H400.64 unitlessRS-20260820T114500Z-2b7c9e31 metrics.modelPairwisePooled
pairwise accuracyafterstate K=256 · H400.69 unitlessRS-20260820T142500Z-8f4a2d17 metrics.modelPairwisePooled
pairwise accuracyafterstate full-train · H400.66 unitlessRS-20260821T094500Z-1a7e3c55 summary (pairwise 0.658 vs 0.685)
pairwise accuracyafterstate D2-teacher · H400.63 unitlessRS-20260821T134500Z-4b9d2f68 metrics.rankingPairwisePooled
pairwise accuracyD4-value student · D4 panel0.64 unitlessRS-20260821T104500Z-77d21e90 metrics.pairwisePooled
pairwise accuracyplanner-distill · planner panel0.74 unitlessdocs/exploratory/finding-11-planner-distillation.md Summary table
pairwise accuracymartingale-dual B0 · H2000.6 unitlessdocs/research/history.md Martingale-dual B0 ranking audit table (line 4526)
pairwise accuracyfair D4 · H400.74 unitlessRS-20260820T142500Z-8f4a2d17 metrics.d4PairwisePooled
pairwise accuracyfair D4 · planner panel0.8 unitlessdocs/exploratory/finding-11-planner-distillation.md Summary table
pairwise accuracyfair D4 · H2000.67 unitlessdocs/research/history.md Martingale-dual B0 ranking audit table (line 4527)
normalized regretafterstate K=8 · H400.41 unitlessRS-20260820T094500Z-5c1e9a04 metrics.modelRegretPooled
normalized regretafterstate K=64 · H400.3 unitlessRS-20260820T114500Z-2b7c9e31 metrics.modelRegretPooled
normalized regretafterstate K=256 · H400.24 unitlessRS-20260820T142500Z-8f4a2d17 metrics.modelRegretPooled
normalized regretafterstate full-train · H400.28 unitlessRS-20260821T094500Z-1a7e3c55 metrics.diagnosticRegretVsIter3
normalized regretafterstate D2-teacher · H400.31 unitlessRS-20260821T134500Z-4b9d2f68 metrics.rankingRegretPooled
normalized regretD4-value student · D4 panel0.29 unitlessRS-20260821T104500Z-77d21e90 metrics.regretPooled
normalized regretplanner-distill · planner panel0.19 unitlessdocs/exploratory/finding-11-planner-distillation.md Summary table
normalized regretmartingale-dual B0 · H2000.35 unitlessdocs/research/history.md Martingale-dual B0 ranking audit table (line 4526)
normalized regretfair D4 · H400.18 unitlessRS-20260820T142500Z-8f4a2d17 metrics.d4RegretPooled
normalized regretfair D4 · planner panel0.13 unitlessdocs/exploratory/finding-11-planner-distillation.md Summary table
normalized regretfair D4 · H2000.28 unitlessdocs/research/history.md Martingale-dual B0 ranking audit table (line 4527)

Spec: web/content/figures/learned-ranking-metrics.json · 8 source records

Most useful conclusions so far

  • Fair chance handling matters. Optimistic, worst-case, or tiny reused reveal samples can rank moves incorrectly.
  • More depth is not automatically more strength. Depth and chance resolution substitute rather than compound: with an approximate chance model the best depth is three, with an exact chance model it is four, and the budget frontier has a flat top at the fair-D4 operating point, reachable from either axis (RS-20260821T192140Z-189fe392, finding-16).

The substitution is visible as a crossed pair of lines:

Mean score by search depth and chance resolution (shared cohort) Mean score by search depth and chance resolution (shared cohort). Cohort 0xa51d1000-0xa51d103f, corrected 17,000-point Hardcore scoring, 64 games per cell except the depth-5 seven-stratum arm, which was stopped by decision after the first 32 seeds (0xa51d1000-0xa51d101f) and is final. The stopped arm is drawn as its own dashed one-point series so the two sample sizes are never joined by a line (same convention as score-vs-depth). No per-mean confidence intervals are recorded anywhere for these arms; the bounds on the paired contrasts live in detection-floor-map. The d2s5 mean is the coordinator-supplied figure printed in finding-10's position-mode table and in the 2026-08-21 log ArmTable (the log path is not an accepted sourceRecord scheme, so the finding is cited). Evidence tier: development / public-development. Sources: docs/exploratory/finding-10-suite-validation.md, docs/exploratory/finding-05-chance-strata.md, docs/exploratory/finding-15-depth5-exact-estimator.md, RS-20260821T205102Z-d89df4b5. Mean score by search depth and chance resolution (shared cohort)200,000250,000300,000350,000400,000450,0002345mean score (points)search depth (plies)5 strata (approximate chance), search depth 2 | d2s5 | mean score: 249,641 points | n = 64 games | source: docs/exploratory/finding-10-suite-validation.md · | position-mode table, d2s5 row (also log 2026-08-21 ArmTable)5 strata (approximate chance), search depth 2d2s5mean score: 249,641 pointsn = 64 gamessource: docs/exploratory/finding-10-suite-validation.md ·position-mode table, d2s5 row (also log 2026-08-21 ArmTable)5 strata (approximate chance), search depth 3 | d3s5 | mean score: 305,051 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · Cost | table, depth 3 / 5 strata row5 strata (approximate chance), search depth 3d3s5mean score: 305,051 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md · Costtable, depth 3 / 5 strata row5 strata (approximate chance), search depth 4 | d4s5 (frozen reference) | mean score: 297,327 points | n = 64 games | source: docs/exploratory/finding-15-depth5-exact-estimator.md · | 8.3 cohort means table5 strata (approximate chance), search depth 4d4s5 (frozen reference)mean score: 297,327 pointsn = 64 gamessource: docs/exploratory/finding-15-depth5-exact-estimator.md ·8.3 cohort means table5 strata (approximate chance), search depth 5 | d5s5 | mean score: 288,704 points | n = 64 games | source: docs/exploratory/finding-15-depth5-exact-estimator.md · | 8.3 cohort means table5 strata (approximate chance), search depth 5d5s5mean score: 288,704 pointsn = 64 gamessource: docs/exploratory/finding-15-depth5-exact-estimator.md ·8.3 cohort means table7 strata (exact chance), search depth 2 | d2s7 | mean score: 265,294 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · Cost | table, depth 2 / 7 strata row7 strata (exact chance), search depth 2d2s7mean score: 265,294 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md · Costtable, depth 2 / 7 strata row7 strata (exact chance), search depth 3 | d3s7 | mean score: 312,327 points | n = 64 games | source: docs/exploratory/finding-15-depth5-exact-estimator.md · | 8.3 cohort means table7 strata (exact chance), search depth 3d3s7mean score: 312,327 pointsn = 64 gamessource: docs/exploratory/finding-15-depth5-exact-estimator.md ·8.3 cohort means table7 strata (exact chance), search depth 4 | d4s7 | mean score: 398,498 points | n = 64 games | source: docs/exploratory/finding-15-depth5-exact-estimator.md · | 8.3 cohort means table7 strata (exact chance), search depth 4d4s7mean score: 398,498 pointsn = 64 gamessource: docs/exploratory/finding-15-depth5-exact-estimator.md ·8.3 cohort means table7 strata, first 32 seeds (arm stopped, final), search depth 5 | d5s7: 32 of 64 planned games, stopped by decision, final | mean score: 411,874 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · metrics.d5s7MeanScore | (411,873.65625)7 strata, first 32 seeds (arm stopped, final), search depth 5d5s7: 32 of 64 planned games, stopped by decision, finalmean score: 411,874 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 · metrics.d5s7MeanScore(411,873.65625)5 strata (approximate chance)7 strata (exact chance)7 strata, first 32 seeds (arm stopped, final)
Mean score by search depth and chance resolution on the shared cohort. The sign of the depth gradient flips with the stratum count.
Source data

Cohort 0xa51d1000-0xa51d103f, corrected 17,000-point Hardcore scoring, 64 games per cell except the depth-5 seven-stratum arm, which was stopped by decision after the first 32 seeds (0xa51d1000-0xa51d101f) and is final. The stopped arm is drawn as its own dashed one-point series so the two sample sizes are never joined by a line (same convention as score-vs-depth). No per-mean confidence intervals are recorded anywhere for these arms; the bounds on the paired contrasts live in detection-floor-map. The d2s5 mean is the coordinator-supplied figure printed in finding-10's position-mode table and in the 2026-08-21 log ArmTable (the log path is not an accepted sourceRecord scheme, so the finding is cited). Evidence tier: development / public-development.

Seriessearch depthmean scoreBoundsnSource
5 strata (approximate chance)2 d2s5249,641 points64docs/exploratory/finding-10-suite-validation.md position-mode table, d2s5 row (also log 2026-08-21 ArmTable)
5 strata (approximate chance)3 d3s5305,051 points64docs/exploratory/finding-05-chance-strata.md Cost table, depth 3 / 5 strata row
5 strata (approximate chance)4 d4s5 (frozen reference)297,327 points64docs/exploratory/finding-15-depth5-exact-estimator.md 8.3 cohort means table
5 strata (approximate chance)5 d5s5288,704 points64docs/exploratory/finding-15-depth5-exact-estimator.md 8.3 cohort means table
7 strata (exact chance)2 d2s7265,294 points64docs/exploratory/finding-05-chance-strata.md Cost table, depth 2 / 7 strata row
7 strata (exact chance)3 d3s7312,327 points64docs/exploratory/finding-15-depth5-exact-estimator.md 8.3 cohort means table
7 strata (exact chance)4 d4s7398,498 points64docs/exploratory/finding-15-depth5-exact-estimator.md 8.3 cohort means table
7 strata, first 32 seeds (arm stopped, final)5 d5s7: 32 of 64 planned games, stopped by decision, final411,874 points32RS-20260821T205102Z-d89df4b5 metrics.d5s7MeanScore (411,873.65625)

Spec: web/content/figures/depth-chance-factorial.json · 4 source records

And the budget view of the same arms shows why neither axis wins on its own:

Mean score against logical work per move — the flat-topped budget frontier Mean score against logical work per move — the flat-topped budget frontier. Cohort 0xa51d1000-0xa51d103f, corrected 17,000-point Hardcore scoring, n=64 per arm except the d5s7 point (32 of 64 planned games, stopped by decision, final), which is drawn as its own dashed one-point series so the two sample sizes are never joined by a line. Cache capacity differs across arms (60,000 entries for the recorded depth-4 comparators, 200,000 for the depth-5 arms, auto-sized 960,695 / 346,921 for the new reveal-sampling arms); capacity provably cannot change play but work per move is not comparable across capacities (finding-16 Limitations; RS-20260821T205102Z-d89df4b5 limitations). Equal-work pair: d3 M=6 (4,244,020; 376,442) vs d4 M=1 (4,956,614; 398,498) is statistically indistinguishable (paired -22,056 [-89,867, +46,009], RS-20260821T192140Z-189fe392 pairedD3M6MinusD4M1). The x axis is linear: the plan's log-scale recommendation is not expressible by the generator, and the d5s7 point dominates the range. Evidence tier: development / public-development. Sources: docs/exploratory/finding-05-chance-strata.md, RS-20260821T192140Z-189fe392, docs/exploratory/finding-15-depth5-exact-estimator.md, RS-20260821T205102Z-d89df4b5. Mean score against logical work per move — the flat-topped budgetfrontier250,000300,000350,000400,000450,0004,13954,429156,8341,045,7191,296,0344,244,0204,956,61413,506,43420,178,32730,183,227176,536,117mean score (points)logical work per move (work units)exact chance (7 strata), logical work per move 4139 | d2 M=1 | mean score: 265,294 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · Cost | table, depth 2 / 7 strata rowexact chance (7 strata), logical work per move 4139d2 M=1mean score: 265,294 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md · Costtable, depth 2 / 7 strata rowexact chance (7 strata), logical work per move 156834 | d3 M=1 | mean score: 312,327 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · | metrics.ladderWorkPerMoveD3N7.M1 / metrics.ladderD3N7.M1exact chance (7 strata), logical work per move 156834d3 M=1mean score: 312,327 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 ·metrics.ladderWorkPerMoveD3N7.M1 / metrics.ladderD3N7.M1exact chance (7 strata), logical work per move 1045719 | d3 M=3 | mean score: 337,306 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · | metrics.ladderWorkPerMoveD3N7.M3 / metrics.ladderD3N7.M3exact chance (7 strata), logical work per move 1045719d3 M=3mean score: 337,306 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 ·metrics.ladderWorkPerMoveD3N7.M3 / metrics.ladderD3N7.M3exact chance (7 strata), logical work per move 4244020 | d3 M=6 | mean score: 376,442 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · | metrics.ladderWorkPerMoveD3N7.M6 / metrics.ladderD3N7.M6exact chance (7 strata), logical work per move 4244020d3 M=6mean score: 376,442 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 ·metrics.ladderWorkPerMoveD3N7.M6 / metrics.ladderD3N7.M6exact chance (7 strata), logical work per move 4956614 | d4 M=1 (fair D4 reference) | mean score: 398,498 points | n = 64 games | source: docs/exploratory/finding-15-depth5-exact-estimator.md · | 8.3 cohort means tableexact chance (7 strata), logical work per move 4956614d4 M=1 (fair D4 reference)mean score: 398,498 pointsn = 64 gamessource: docs/exploratory/finding-15-depth5-exact-estimator.md ·8.3 cohort means tableexact chance (7 strata), logical work per move 13506434 | d3 M=12 (full joint coverage) | mean score: 349,345 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · | metrics.ladderWorkPerMoveD3N7.M12 / metrics.ladderD3N7.M12exact chance (7 strata), logical work per move 13506434d3 M=12 (full joint coverage)mean score: 349,345 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 ·metrics.ladderWorkPerMoveD3N7.M12 / metrics.ladderD3N7.M12exact chance (7 strata), logical work per move 20178327 | d4 M=2 | mean score: 356,548 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · metrics.d4N7M2WorkPerMove | / metrics.d4N7M2MeanScoreexact chance (7 strata), logical work per move 20178327d4 M=2mean score: 356,548 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 · metrics.d4N7M2WorkPerMove/ metrics.d4N7M2MeanScoreapproximate chance (5 strata), logical work per move 54429 | d3 M=1 | mean score: 305,051 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · Cost | table, depth 3 / 5 strata rowapproximate chance (5 strata), logical work per move 54429d3 M=1mean score: 305,051 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md · Costtable, depth 3 / 5 strata rowapproximate chance (5 strata), logical work per move 1296034 | d4 M=1 (frozen reference) | mean score: 297,327 points | n = 64 games | source: docs/exploratory/finding-05-chance-strata.md · Cost | table, depth 4 / 5 strata rowapproximate chance (5 strata), logical work per move 1296034d4 M=1 (frozen reference)mean score: 297,327 pointsn = 64 gamessource: docs/exploratory/finding-05-chance-strata.md · Costtable, depth 4 / 5 strata rowapproximate chance (5 strata), logical work per move 30183227 | d5 M=1 | mean score: 288,704 points | n = 64 games | source: docs/exploratory/finding-15-depth5-exact-estimator.md · | 8.3 cohort means tableapproximate chance (5 strata), logical work per move 30183227d5 M=1mean score: 288,704 pointsn = 64 gamessource: docs/exploratory/finding-15-depth5-exact-estimator.md ·8.3 cohort means tableexact chance, d5 M=1 (n=32, stopped, final), logical work per move 176536117 | d5s7: 32 of 64 planned games, stopped by decision, final | mean score: 411,874 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · metrics.d5s7WorkPerMove / | metrics.d5s7MeanScoreexact chance, d5 M=1 (n=32, stopped, final), logical work per move 176536117d5s7: 32 of 64 planned games, stopped by decision, finalmean score: 411,874 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 · metrics.d5s7WorkPerMove /metrics.d5s7MeanScoreexact chance (7 strata)approximate chance (5 strata)exact chance, d5 M=1 (n=32, stopped, final)
Mean score against logical work per move. Two different ways of spending the same budget — depth or reveal sampling — land in the same place, and the frontier is flat at the fair-D4 operating point.
Source data

Cohort 0xa51d1000-0xa51d103f, corrected 17,000-point Hardcore scoring, n=64 per arm except the d5s7 point (32 of 64 planned games, stopped by decision, final), which is drawn as its own dashed one-point series so the two sample sizes are never joined by a line. Cache capacity differs across arms (60,000 entries for the recorded depth-4 comparators, 200,000 for the depth-5 arms, auto-sized 960,695 / 346,921 for the new reveal-sampling arms); capacity provably cannot change play but work per move is not comparable across capacities (finding-16 Limitations; RS-20260821T205102Z-d89df4b5 limitations). Equal-work pair: d3 M=6 (4,244,020; 376,442) vs d4 M=1 (4,956,614; 398,498) is statistically indistinguishable (paired -22,056 [-89,867, +46,009], RS-20260821T192140Z-189fe392 pairedD3M6MinusD4M1). The x axis is linear: the plan's log-scale recommendation is not expressible by the generator, and the d5s7 point dominates the range. Evidence tier: development / public-development.

Serieslogical work per movemean scoreBoundsnSource
exact chance (7 strata)4139 d2 M=1265,294 points64docs/exploratory/finding-05-chance-strata.md Cost table, depth 2 / 7 strata row
exact chance (7 strata)156834 d3 M=1312,327 points64RS-20260821T192140Z-189fe392 metrics.ladderWorkPerMoveD3N7.M1 / metrics.ladderD3N7.M1
exact chance (7 strata)1045719 d3 M=3337,306 points64RS-20260821T192140Z-189fe392 metrics.ladderWorkPerMoveD3N7.M3 / metrics.ladderD3N7.M3
exact chance (7 strata)4244020 d3 M=6376,442 points64RS-20260821T192140Z-189fe392 metrics.ladderWorkPerMoveD3N7.M6 / metrics.ladderD3N7.M6
exact chance (7 strata)4956614 d4 M=1 (fair D4 reference)398,498 points64docs/exploratory/finding-15-depth5-exact-estimator.md 8.3 cohort means table
exact chance (7 strata)13506434 d3 M=12 (full joint coverage)349,345 points64RS-20260821T192140Z-189fe392 metrics.ladderWorkPerMoveD3N7.M12 / metrics.ladderD3N7.M12
exact chance (7 strata)20178327 d4 M=2356,548 points64RS-20260821T192140Z-189fe392 metrics.d4N7M2WorkPerMove / metrics.d4N7M2MeanScore
approximate chance (5 strata)54429 d3 M=1305,051 points64docs/exploratory/finding-05-chance-strata.md Cost table, depth 3 / 5 strata row
approximate chance (5 strata)1296034 d4 M=1 (frozen reference)297,327 points64docs/exploratory/finding-05-chance-strata.md Cost table, depth 4 / 5 strata row
approximate chance (5 strata)30183227 d5 M=1288,704 points64docs/exploratory/finding-15-depth5-exact-estimator.md 8.3 cohort means table
exact chance, d5 M=1 (n=32, stopped, final)176536117 d5s7: 32 of 64 planned games, stopped by decision, final411,874 points32RS-20260821T205102Z-d89df4b5 metrics.d5s7WorkPerMove / metrics.d5s7MeanScore

Spec: web/content/figures/score-vs-work-frontier.json · 4 source records

  • Flow matters before spectacular chains. The task record repeatedly used roughly 2.4 numbered clears and 1.4 reveals per move as the region associated with stable long games. Treat these as diagnostic targets from limited runs, not proven universal thresholds.
  • Static board potential is insufficient. Similar-looking boards can have very different futures depending on how reachable triggers and covered discs evolve across rises.
  • State-value accuracy is not action-ranking accuracy. Training data must cover legal siblings or use an objective designed for relative action value. As of 2026-08-21 this is necessary but demonstrably not sufficient: a student given successor-closed coverage, every legal sibling, and exact search-value labels still ranked worse than a one-ply exact search.
  • Score is heavy-tailed, and this sets a hard measurement floor. One million-point game can coexist with a much lower average, and pairing does not rescue small effects: before running a cohort, state the effect size the mechanism predicts and check it against the detection floor (RS-20260821T205102Z-d89df4b5).

Each contrast of the depth factorial, drawn against the smallest effect its cohort could have seen:

Paired contrasts against their detection floors — every significant result is above its floor, every null below it Paired contrasts against their detection floors — every significant result is above its floor, every null below it. All six contrasts from RS-20260821T205102Z-d89df4b5, cohort 0xa51d1000-0xa51d103f, corrected 17,000-point Hardcore scoring; the n=32 contrasts cover the cohort's first 32 seeds (d5s7 arm stopped by decision, final). Whiskers are one-sided 95% percentile-bootstrap lower bounds (20,000 resamples, Mulberry32 domain 0xb0075eed); upper bounds are NOT recorded for these contrasts in the record and are omitted, not zero. Detection floor = 1.645 x sd / sqrt(n), the smallest true effect whose one-sided 95% bound would clear zero (metrics.detectionFloorDefinition). The d4s7-d3s7 lower bound is carried as +26,468 from metrics.bootstrapVersusNormalApproximation; status.md and the 2026-08-21 log print +26,605, which is prose drift — the record and finding-05's interaction table both say +26,468. Evidence tier: public-development. Sources: RS-20260821T205102Z-d89df4b5. Paired contrasts against their detection floors — every significantresult is above its floor, every null below it-100,000-50,000050,000100,000150,000d4s7-d4s5(n=64)d4s7-d3s7(n=64)d5s7-d5s5(n=32)d5s7-d3s7(n=32)d5s5-d4s5(n=64)d5s7-d4s7(n=32)points (points)paired contrast (same seeds, whole games)paired mean delta (whisker: 95% lower bound) — d4s7-d4s5 (n=64) | strata at depth 4; W-T-L 41-0-23; 3.82x work; floor 55,192: | above | points: +101,171 points | 95% lower bound: 47,447 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD4s7MinusD4s5paired mean delta (whisker: 95% lower bound) — d4s7-d4s5 (n=64)strata at depth 4; W-T-L 41-0-23; 3.82x work; floor 55,192:abovepoints: +101,171 points95% lower bound: 47,447 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD4s7MinusD4s5paired mean delta (whisker: 95% lower bound) — d4s7-d3s7 (n=64) | depth 4 minus 3 at 7 strata; W-T-L 40-0-24; floor 61,457: above | points: +86,172 points | 95% lower bound: 26,468 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · metrics.powerTable[1]; | metrics.bootstrapVersusNormalApproximationpaired mean delta (whisker: 95% lower bound) — d4s7-d3s7 (n=64)depth 4 minus 3 at 7 strata; W-T-L 40-0-24; floor 61,457: abovepoints: +86,172 points95% lower bound: 26,468 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 · metrics.powerTable[1];metrics.bootstrapVersusNormalApproximationpaired mean delta (whisker: 95% lower bound) — d5s7-d5s5 (n=32) | strata at depth 5; W-T-L 19-0-13; 5.85x work; floor 95,207: | above | points: +123,613 points | 95% lower bound: 32,575 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s7MinusD5s5paired mean delta (whisker: 95% lower bound) — d5s7-d5s5 (n=32)strata at depth 5; W-T-L 19-0-13; 5.85x work; floor 95,207:abovepoints: +123,613 points95% lower bound: 32,575 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s7MinusD5s5paired mean delta (whisker: 95% lower bound) — d5s7-d3s7 (n=32) | two plies at 7 strata; W-T-L 20-0-12; floor 97,211: below | points: +86,397 points | 95% lower bound: -6,303 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s7MinusD3s7paired mean delta (whisker: 95% lower bound) — d5s7-d3s7 (n=32)two plies at 7 strata; W-T-L 20-0-12; floor 97,211: belowpoints: +86,397 points95% lower bound: -6,303 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s7MinusD3s7paired mean delta (whisker: 95% lower bound) — d5s5-d4s5 (n=64) | depth at 5 strata; W-T-L 33-0-31; 23.29x work; floor 47,052: | below | points: -8,624 points | 95% lower bound: -55,134 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s5MinusD4s5paired mean delta (whisker: 95% lower bound) — d5s5-d4s5 (n=64)depth at 5 strata; W-T-L 33-0-31; 23.29x work; floor 47,052:belowpoints: -8,624 points95% lower bound: -55,134 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s5MinusD4s5paired mean delta (whisker: 95% lower bound) — d5s7-d4s7 (n=32) | depth at 7 strata; W-T-L 17-0-15; 35.62x work; floor 107,988: | below (22% of floor) | points: +23,367 points | 95% lower bound: -83,046 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s7MinusD4s7paired mean delta (whisker: 95% lower bound) — d5s7-d4s7 (n=32)depth at 7 strata; W-T-L 17-0-15; 35.62x work; floor 107,988:below (22% of floor)points: +23,367 points95% lower bound: -83,046 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s7MinusD4s7detection floor 1.645*sd/sqrt(n) — d4s7-d4s5 (n=64) | floor for d4s7-d4s5 | points: +55,192 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.powerTable[0].detectionFloordetection floor 1.645*sd/sqrt(n) — d4s7-d4s5 (n=64)floor for d4s7-d4s5points: +55,192 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.powerTable[0].detectionFloordetection floor 1.645*sd/sqrt(n) — d4s7-d3s7 (n=64) | floor for d4s7-d3s7 | points: +61,457 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.powerTable[1].detectionFloordetection floor 1.645*sd/sqrt(n) — d4s7-d3s7 (n=64)floor for d4s7-d3s7points: +61,457 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.powerTable[1].detectionFloordetection floor 1.645*sd/sqrt(n) — d5s7-d5s5 (n=32) | floor for d5s7-d5s5 | points: +95,207 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.powerTable[2].detectionFloordetection floor 1.645*sd/sqrt(n) — d5s7-d5s5 (n=32)floor for d5s7-d5s5points: +95,207 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.powerTable[2].detectionFloordetection floor 1.645*sd/sqrt(n) — d5s7-d3s7 (n=32) | floor for d5s7-d3s7 | points: +97,211 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.powerTable[3].detectionFloordetection floor 1.645*sd/sqrt(n) — d5s7-d3s7 (n=32)floor for d5s7-d3s7points: +97,211 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.powerTable[3].detectionFloordetection floor 1.645*sd/sqrt(n) — d5s5-d4s5 (n=64) | floor for d5s5-d4s5 | points: +47,052 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.powerTable[4].detectionFloordetection floor 1.645*sd/sqrt(n) — d5s5-d4s5 (n=64)floor for d5s5-d4s5points: +47,052 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.powerTable[4].detectionFloordetection floor 1.645*sd/sqrt(n) — d5s7-d4s7 (n=32) | floor for d5s7-d4s7 | points: +107,988 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.powerTable[5].detectionFloordetection floor 1.645*sd/sqrt(n) — d5s7-d4s7 (n=32)floor for d5s7-d4s7points: +107,988 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.powerTable[5].detectionFloorpaired mean delta (whisker: 95% lower bound)detection floor 1.645*sd/sqrt(n)
The six paired contrasts of the depth factorial against their detection floors. Every significant result clears its floor; every null sits below it and is a non-measurement, not evidence of no effect.
Source data

All six contrasts from RS-20260821T205102Z-d89df4b5, cohort 0xa51d1000-0xa51d103f, corrected 17,000-point Hardcore scoring; the n=32 contrasts cover the cohort's first 32 seeds (d5s7 arm stopped by decision, final). Whiskers are one-sided 95% percentile-bootstrap lower bounds (20,000 resamples, Mulberry32 domain 0xb0075eed); upper bounds are NOT recorded for these contrasts in the record and are omitted, not zero. Detection floor = 1.645 x sd / sqrt(n), the smallest true effect whose one-sided 95% bound would clear zero (metrics.detectionFloorDefinition). The d4s7-d3s7 lower bound is carried as +26,468 from metrics.bootstrapVersusNormalApproximation; status.md and the 2026-08-21 log print +26,605, which is prose drift — the record and finding-05's interaction table both say +26,468. Evidence tier: public-development.

Seriespaired contrast (same seeds, whole games)pointsBoundsnSource
paired mean delta (whisker: 95% lower bound)d4s7-d4s5 (n=64) strata at depth 4; W-T-L 41-0-23; 3.82x work; floor 55,192: above101,171 pointslower 47,44764RS-20260821T205102Z-d89df4b5 metrics.pairedD4s7MinusD4s5
paired mean delta (whisker: 95% lower bound)d4s7-d3s7 (n=64) depth 4 minus 3 at 7 strata; W-T-L 40-0-24; floor 61,457: above86,172 pointslower 26,46864RS-20260821T205102Z-d89df4b5 metrics.powerTable[1]; metrics.bootstrapVersusNormalApproximation
paired mean delta (whisker: 95% lower bound)d5s7-d5s5 (n=32) strata at depth 5; W-T-L 19-0-13; 5.85x work; floor 95,207: above123,613 pointslower 32,57532RS-20260821T205102Z-d89df4b5 metrics.pairedD5s7MinusD5s5
paired mean delta (whisker: 95% lower bound)d5s7-d3s7 (n=32) two plies at 7 strata; W-T-L 20-0-12; floor 97,211: below86,397 pointslower -6,30332RS-20260821T205102Z-d89df4b5 metrics.pairedD5s7MinusD3s7
paired mean delta (whisker: 95% lower bound)d5s5-d4s5 (n=64) depth at 5 strata; W-T-L 33-0-31; 23.29x work; floor 47,052: below-8,624 pointslower -55,13464RS-20260821T205102Z-d89df4b5 metrics.pairedD5s5MinusD4s5
paired mean delta (whisker: 95% lower bound)d5s7-d4s7 (n=32) depth at 7 strata; W-T-L 17-0-15; 35.62x work; floor 107,988: below (22% of floor)23,367 pointslower -83,04632RS-20260821T205102Z-d89df4b5 metrics.pairedD5s7MinusD4s7
detection floor 1.645*sd/sqrt(n)d4s7-d4s5 (n=64) floor for d4s7-d4s555,192 points64RS-20260821T205102Z-d89df4b5 metrics.powerTable[0].detectionFloor
detection floor 1.645*sd/sqrt(n)d4s7-d3s7 (n=64) floor for d4s7-d3s761,457 points64RS-20260821T205102Z-d89df4b5 metrics.powerTable[1].detectionFloor
detection floor 1.645*sd/sqrt(n)d5s7-d5s5 (n=32) floor for d5s7-d5s595,207 points32RS-20260821T205102Z-d89df4b5 metrics.powerTable[2].detectionFloor
detection floor 1.645*sd/sqrt(n)d5s7-d3s7 (n=32) floor for d5s7-d3s797,211 points32RS-20260821T205102Z-d89df4b5 metrics.powerTable[3].detectionFloor
detection floor 1.645*sd/sqrt(n)d5s5-d4s5 (n=64) floor for d5s5-d4s547,052 points64RS-20260821T205102Z-d89df4b5 metrics.powerTable[4].detectionFloor
detection floor 1.645*sd/sqrt(n)d5s7-d4s7 (n=32) floor for d5s7-d4s7107,988 points32RS-20260821T205102Z-d89df4b5 metrics.powerTable[5].detectionFloor

Spec: web/content/figures/detection-floor-map.json · 1 source record

  • D4 is useful but not the answer. It is a strong tactical fallback and teacher, yet its average is not close enough to qualify.
  • The leaf evaluation prices what search cannot see. Refitting it toward a quantity the search already computes exactly — achievable clears over the next few moves — loses monotonically, and loses by destroying the upper tail rather than the lower one. A leaf that correlates weakly with short-horizon clear rate is behaving correctly, not failing.
  • Cheap proxies for this game invert. Three independent times a short-horizon screen has ranked configurations in the opposite order from complete games: the suite-h9-v1 scenario suite, a depth-3 screening run, and an eight-move achievable-clear label. Screen at the depth you intend to deploy, or do not screen.

Directions closed by exploratory evidence, 2026-08-20/23

Each entry rejects the exact configuration tested, at development or pilot tier. None of them is a proof that the underlying idea is impossible, and each names what would reopen it.

The screen and confirmation deltas behind these closures, with their recorded bounds:

Every screen and confirmation paired delta with its recorded 95% bound Every screen and confirmation paired delta with its recorded 95% bound. Whole-game paired deltas against each experiment's reference arm, corrected 17,000-point Hardcore scoring. Bounds are one-sided 95% bootstrap lower bounds; upper bounds are drawn only where recorded (the three reveal-sampling contrasts, CMA-ES, both survival-instinct arms). finding-03, finding-14, both learned-leaf arms and A900 record no upper bound — the whisker is absent there, not zero. finding-05's confirmation lower bound is printed +47,457 in the finding and +47,446.8 in RS-20260821T205102Z-d89df4b5; the figure uses the record (+47,447). A900's rejection rests on flat flow statistics and opposite-sign halves, not the score delta, which sits inside its 30,957 detection floor (RS-20260823T131226Z-16564ed9). Cohorts differ per bar (n carried per point; the survival-instinct arms share 128 fresh seeds, A900 256 fresh seeds); this chart assembles separate experiments, it does not pool them. Evidence tier: development / public-development screens. Sources: RS-20260821T205102Z-d89df4b5, docs/exploratory/finding-08-learned-leaf.md, RS-20260821T192140Z-189fe392, docs/exploratory/finding-03-rollout-veto-17k.md, docs/exploratory/finding-14-leaf-reweight.md, RS-20260822T120736Z-662b39ca, RS-20260822T233343Z-12becce9, RS-20260823T131226Z-16564ed9. Every screen and confirmation paired delta with its recorded 95% bound-300,000-200,000-100,0000100,000200,000f05confirm d4s7-d4s5f08learned leaf s5f08learned leaf s7f16 d3M6-M1f16 d4M2-M1f16 d3M12-M6f03rollout vetof14full leaf refitCMA-ESleafsurvival STRICTsurvival LITERALA900reveal constr.paired mean score delta (points)screen or confirmation (candidate minus reference)paired mean score delta vs reference — f05 confirm d4s7-d4s5 | finding-05 confirmation, chance strata at depth 4; W-T-L 41-0-23 | paired mean score delta: +101,171 points | 95% lower bound: 47,447 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD4s7MinusD4s5paired mean score delta vs reference — f05 confirm d4s7-d4s5finding-05 confirmation, chance strata at depth 4; W-T-L 41-0-23paired mean score delta: +101,171 points95% lower bound: 47,447 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD4s7MinusD4s5paired mean score delta vs reference — f08 learned leaf s5 | learned leaf minus reference, 5 strata; W-T-L 37-0-27; | significant | paired mean score delta: +39,105 points | 95% lower bound: 1,138 points | n = 64 games | source: docs/exploratory/finding-08-learned-leaf.md · section 7 | paired deltas tablepaired mean score delta vs reference — f08 learned leaf s5learned leaf minus reference, 5 strata; W-T-L 37-0-27;significantpaired mean score delta: +39,105 points95% lower bound: 1,138 pointsn = 64 gamessource: docs/exploratory/finding-08-learned-leaf.md · section 7paired deltas tablepaired mean score delta vs reference — f08 learned leaf s7 | learned leaf minus reference, 7 strata; W-T-L 34-0-30; not | significant | paired mean score delta: +17,281 points | 95% lower bound: -55,892 points | n = 64 games | source: docs/exploratory/finding-08-learned-leaf.md · section 7 | paired deltas tablepaired mean score delta vs reference — f08 learned leaf s7learned leaf minus reference, 7 strata; W-T-L 34-0-30; notsignificantpaired mean score delta: +17,281 points95% lower bound: -55,892 pointsn = 64 gamessource: docs/exploratory/finding-08-learned-leaf.md · section 7paired deltas tablepaired mean score delta vs reference — f16 d3 M6-M1 | reveal sampling at depth 3, M=6 minus M=1; W-T-L 36-0-28 | paired mean score delta: +64,116 points | bounds: 7,475 points to 121,776 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · | metrics.pairedD3M6MinusD3M1paired mean score delta vs reference — f16 d3 M6-M1reveal sampling at depth 3, M=6 minus M=1; W-T-L 36-0-28paired mean score delta: +64,116 pointsbounds: 7,475 points to 121,776 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 ·metrics.pairedD3M6MinusD3M1paired mean score delta vs reference — f16 d4 M2-M1 | reveal sampling at depth 4, M=2 minus M=1; W-T-L 28-0-36 | paired mean score delta: -41,950 points | bounds: -100,137 points to 17,541 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · | metrics.pairedD4M2MinusD4M1paired mean score delta vs reference — f16 d4 M2-M1reveal sampling at depth 4, M=2 minus M=1; W-T-L 28-0-36paired mean score delta: -41,950 pointsbounds: -100,137 points to 17,541 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 ·metrics.pairedD4M2MinusD4M1paired mean score delta vs reference — f16 d3 M12-M6 | full joint coverage minus M=6 at depth 3; W-T-L 28-0-36; | saturation | paired mean score delta: -27,097 points | bounds: -83,807 points to 31,209 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · | metrics.pairedD3M12MinusD3M6paired mean score delta vs reference — f16 d3 M12-M6full joint coverage minus M=6 at depth 3; W-T-L 28-0-36;saturationpaired mean score delta: -27,097 pointsbounds: -83,807 points to 31,209 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 ·metrics.pairedD3M12MinusD3M6paired mean score delta vs reference — f03 rollout veto | 25-move rollout veto, 17,000-point port; W-T-L 9-4-19; rejected | paired mean score delta: -46,511 points | 95% lower bound: -91,925 points | n = 32 games | source: docs/exploratory/finding-03-rollout-veto-17k.md · | section 5.2 paired comparisonpaired mean score delta vs reference — f03 rollout veto25-move rollout veto, 17,000-point port; W-T-L 9-4-19; rejectedpaired mean score delta: -46,511 points95% lower bound: -91,925 pointsn = 32 gamessource: docs/exploratory/finding-03-rollout-veto-17k.md ·section 5.2 paired comparisonpaired mean score delta vs reference — f14 full leaf refit | leaf reweighted fully toward the achievable-clear direction; | W-T-L 7-0-57; rejected | paired mean score delta: -237,182 points | 95% lower bound: -290,406 points | n = 64 games | source: docs/exploratory/finding-14-leaf-reweight.md · Summary / | dose table (t2-fair-a1)paired mean score delta vs reference — f14 full leaf refitleaf reweighted fully toward the achievable-clear direction;W-T-L 7-0-57; rejectedpaired mean score delta: -237,182 points95% lower bound: -290,406 pointsn = 64 gamessource: docs/exploratory/finding-14-leaf-reweight.md · Summary /dose table (t2-fair-a1)paired mean score delta vs reference — CMA-ES leaf | CMA-ES-tuned leaf screen, held out; W-T-L 28-0-36 | paired mean score delta: -30,300 points | bounds: -70,928 points to 9,786 points | n = 64 games | source: RS-20260822T120736Z-662b39ca · | metrics.heldOutD4S5.pairedScorepaired mean score delta vs reference — CMA-ES leafCMA-ES-tuned leaf screen, held out; W-T-L 28-0-36paired mean score delta: -30,300 pointsbounds: -70,928 points to 9,786 pointsn = 64 gamessource: RS-20260822T120736Z-662b39ca ·metrics.heldOutD4S5.pairedScorepaired mean score delta vs reference — survival STRICT | survival-instinct root filter, strict rule; W-T-L 41-26-61; | inconclusive by its coverage rule | paired mean score delta: -1,970 points | bounds: -27,738 points to 22,313 points | n = 128 games | source: RS-20260822T233343Z-12becce9 · | metrics.strict.pairedScorepaired mean score delta vs reference — survival STRICTsurvival-instinct root filter, strict rule; W-T-L 41-26-61;inconclusive by its coverage rulepaired mean score delta: -1,970 pointsbounds: -27,738 points to 22,313 pointsn = 128 gamessource: RS-20260822T233343Z-12becce9 ·metrics.strict.pairedScorepaired mean score delta vs reference — survival LITERAL | survival-instinct root filter, literal rule; W-T-L 43-0-85; | rejected | paired mean score delta: -97,064 points | bounds: -136,887 points to -59,500 points | n = 128 games | source: RS-20260822T233343Z-12becce9 · | metrics.literal.pairedScorepaired mean score delta vs reference — survival LITERALsurvival-instinct root filter, literal rule; W-T-L 43-0-85;rejectedpaired mean score delta: -97,064 pointsbounds: -136,887 points to -59,500 pointsn = 128 gamessource: RS-20260822T233343Z-12becce9 ·metrics.literal.pairedScorepaired mean score delta vs reference — A900 reveal constr. | reveal-construction term at threshold 900; W-T-L 99-53-104; | inside its 30,957 floor | paired mean score delta: +3,204 points | 95% lower bound: -26,860 points | n = 256 games | source: RS-20260823T131226Z-16564ed9 · metrics.A900_minus_frozenpaired mean score delta vs reference — A900 reveal constr.reveal-construction term at threshold 900; W-T-L 99-53-104;inside its 30,957 floorpaired mean score delta: +3,204 points95% lower bound: -26,860 pointsn = 256 gamessource: RS-20260823T131226Z-16564ed9 · metrics.A900_minus_frozenpaired mean score delta vs reference
The score-valued closures at a glance: every screen and confirmation paired delta with its recorded 95% bound. Whiskers are absent, not zero, where a record carries no upper bound.
Source data

Whole-game paired deltas against each experiment's reference arm, corrected 17,000-point Hardcore scoring. Bounds are one-sided 95% bootstrap lower bounds; upper bounds are drawn only where recorded (the three reveal-sampling contrasts, CMA-ES, both survival-instinct arms). finding-03, finding-14, both learned-leaf arms and A900 record no upper bound — the whisker is absent there, not zero. finding-05's confirmation lower bound is printed +47,457 in the finding and +47,446.8 in RS-20260821T205102Z-d89df4b5; the figure uses the record (+47,447). A900's rejection rests on flat flow statistics and opposite-sign halves, not the score delta, which sits inside its 30,957 detection floor (RS-20260823T131226Z-16564ed9). Cohorts differ per bar (n carried per point; the survival-instinct arms share 128 fresh seeds, A900 256 fresh seeds); this chart assembles separate experiments, it does not pool them. Evidence tier: development / public-development screens.

Seriesscreen or confirmation (candidate minus reference)paired mean score deltaBoundsnSource
paired mean score delta vs referencef05 confirm d4s7-d4s5 finding-05 confirmation, chance strata at depth 4; W-T-L 41-0-23101,171 pointslower 47,44764RS-20260821T205102Z-d89df4b5 metrics.pairedD4s7MinusD4s5
paired mean score delta vs referencef08 learned leaf s5 learned leaf minus reference, 5 strata; W-T-L 37-0-27; significant39,105 pointslower 1,13864docs/exploratory/finding-08-learned-leaf.md section 7 paired deltas table
paired mean score delta vs referencef08 learned leaf s7 learned leaf minus reference, 7 strata; W-T-L 34-0-30; not significant17,281 pointslower -55,89264docs/exploratory/finding-08-learned-leaf.md section 7 paired deltas table
paired mean score delta vs referencef16 d3 M6-M1 reveal sampling at depth 3, M=6 minus M=1; W-T-L 36-0-2864,116 points7,475 to 121,77664RS-20260821T192140Z-189fe392 metrics.pairedD3M6MinusD3M1
paired mean score delta vs referencef16 d4 M2-M1 reveal sampling at depth 4, M=2 minus M=1; W-T-L 28-0-36-41,950 points-100,137 to 17,54164RS-20260821T192140Z-189fe392 metrics.pairedD4M2MinusD4M1
paired mean score delta vs referencef16 d3 M12-M6 full joint coverage minus M=6 at depth 3; W-T-L 28-0-36; saturation-27,097 points-83,807 to 31,20964RS-20260821T192140Z-189fe392 metrics.pairedD3M12MinusD3M6
paired mean score delta vs referencef03 rollout veto 25-move rollout veto, 17,000-point port; W-T-L 9-4-19; rejected-46,511 pointslower -91,92532docs/exploratory/finding-03-rollout-veto-17k.md section 5.2 paired comparison
paired mean score delta vs referencef14 full leaf refit leaf reweighted fully toward the achievable-clear direction; W-T-L 7-0-57; rejected-237,182 pointslower -290,40664docs/exploratory/finding-14-leaf-reweight.md Summary / dose table (t2-fair-a1)
paired mean score delta vs referenceCMA-ES leaf CMA-ES-tuned leaf screen, held out; W-T-L 28-0-36-30,300 points-70,928 to 9,78664RS-20260822T120736Z-662b39ca metrics.heldOutD4S5.pairedScore
paired mean score delta vs referencesurvival STRICT survival-instinct root filter, strict rule; W-T-L 41-26-61; inconclusive by its coverage rule-1,970 points-27,738 to 22,313128RS-20260822T233343Z-12becce9 metrics.strict.pairedScore
paired mean score delta vs referencesurvival LITERAL survival-instinct root filter, literal rule; W-T-L 43-0-85; rejected-97,064 points-136,887 to -59,500128RS-20260822T233343Z-12becce9 metrics.literal.pairedScore
paired mean score delta vs referenceA900 reveal constr. reveal-construction term at threshold 900; W-T-L 99-53-104; inside its 30,957 floor3,204 pointslower -26,860256RS-20260823T131226Z-16564ed9 metrics.A900_minus_frozen

Spec: web/content/figures/screen-deltas-with-bounds.json · 8 source records

The closures themselves, one dot per direction, each drawn against the detection floor its record states — the geometric difference between "rejected" and "not measurable":

Directions closed 2026-08-20/23 — headline paired delta with recorded bounds, and the detection floor where one is recorded Directions closed 2026-08-20/23 — headline paired delta with recorded bounds, and the detection floor where one is recorded. One dot per closed direction whose headline result is a points number, copied verbatim from its record; whiskers are the recorded bounds — one-sided 95% bootstrap lower bounds where that is all the record carries (both depth rows, reveal-construction A900, rollout veto, leaf refit), two-sided 95% where recorded (reveal M=2, CMA-ES, survival-instinct STRICT), and none for terminal utility, whose zero is EXACT (byte-identical games at four magnitudes from -1,000,000 to -50,000,000, 0-64-0; no statistical bound exists or is needed). Detection-floor markers (1.645*sd/sqrt(n)) are drawn only where the record carries one; no floor is recorded for reveal M=2, the rollout veto, or the leaf refit — markers omitted, not zero. Every closure rejects the exact tested configuration at development/pilot tier; none proves the underlying idea impossible, and each row's 'reopens if' condition stays in docs/research/status.md. The plan's 'depth beyond four plies' row is charted as its two recorded contrasts (s7 n=32 stopped/final, s5 n=64). Not chartable on a points axis, recorded here verbatim instead: compact afterstate model reproducing D4's ordering — top-1 0.375 against the 0.60 gate (RS-20260821T104500Z-77d21e90); the suite-h9-v1 benchmark as a strength measure — Spearman -0.257 (finding-10). status.md's table gained three further non-points closures after this figure's plan was written, also not chartable here: leaf-cost NNUE student top-1 0.296/0.3011 by half-fold against the 0.60 gate (RS-20260823T194142Z-946e3cd1), H-pool optimistic states tau -0.959 [cluster bootstrap -3.069, -0.390] (RS-20260823T205143Z-ead14c9d), and P-SOL G0 cheap-label proxy mean tau 0.370, LB95 0.283 (RS-20260823T225753Z-0fbd48c3). Cohorts differ per point (n on every marker); this chart assembles separate experiments, it does not pool them. Sources: RS-20260821T205102Z-d89df4b5, RS-20260821T192140Z-189fe392, RS-20260823T131226Z-16564ed9, docs/exploratory/finding-03-rollout-veto-17k.md, docs/exploratory/finding-14-leaf-reweight.md, RS-20260822T120736Z-662b39ca, RS-20260822T233343Z-12becce9, docs/exploratory/finding-04-terminal-utility-saturated.md. Directions closed 2026-08-20/23 — headline paired delta with recorded bounds, and the detection floor where one is recordedDirections closed 2026-08-20/23 — headline paired delta with recorded b…Directions closed 2026-08-20/23 — headline paired delta with recorded bounds, and the detection floor where one is recorded-300,000-200,000-100,0000100,000200,000depth 5 vs 4 · s7depth 5 vs 4 · s5reveal M=2 on ply 4reveal-constr. A900rollout veto 17kleaf refit to clearsCMA-ES leafsurvival-instinct STRICTsurvival-instinct ST…survival-instinct STRICTterminal utility ×50paired delta vs reference (points)headline paired delta (points) — depth 5 vs 4 · s7 | d5s7 - d4s7; arm stopped at 32 of 64 games, final; 35.62x work; | below its 107,988 floor: a non-measurement | paired delta vs reference: 23,367 points | 95% lower bound: -83,046 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s7MinusD4s7headline paired delta (points) — depth 5 vs 4 · s7d5s7 - d4s7; arm stopped at 32 of 64 games, final; 35.62x work;below its 107,988 floor: a non-measurementpaired delta vs reference: 23,367 points95% lower bound: -83,046 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s7MinusD4s7headline paired delta (points) — depth 5 vs 4 · s5 | d5s5 - d4s5; 23.29x work; below its 47,052 floor | paired delta vs reference: -8,624 points | 95% lower bound: -55,134 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.pairedD5s5MinusD4s5headline paired delta (points) — depth 5 vs 4 · s5d5s5 - d4s5; 23.29x work; below its 47,052 floorpaired delta vs reference: -8,624 points95% lower bound: -55,134 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.pairedD5s5MinusD4s5headline paired delta (points) — reveal M=2 on ply 4 | d4 M=2 - M=1; 4.07x work; only 28.6% joint coverage was | affordable at depth 4; no floor recorded | paired delta vs reference: -41,950 points | bounds: -100,137 points to 17,541 points | n = 64 games | source: RS-20260821T192140Z-189fe392 · | metrics.pairedD4M2MinusD4M1headline paired delta (points) — reveal M=2 on ply 4d4 M=2 - M=1; 4.07x work; only 28.6% joint coverage wasaffordable at depth 4; no floor recordedpaired delta vs reference: -41,950 pointsbounds: -100,137 points to 17,541 pointsn = 64 gamessource: RS-20260821T192140Z-189fe392 ·metrics.pairedD4M2MinusD4M1headline paired delta (points) — reveal-constr. A900 | aligned_double_hit +900 on 256 fresh paired games; inside its | 30,957 floor; halves +48,762 / -42,354 opposite signs; | reveals/move flat 1.1520 vs 1.1536 | paired delta vs reference: 3,204 points | 95% lower bound: -26,860 points | n = 256 games | source: RS-20260823T131226Z-16564ed9 · metrics.A900_minus_frozenheadline paired delta (points) — reveal-constr. A900aligned_double_hit +900 on 256 fresh paired games; inside its30,957 floor; halves +48,762 / -42,354 opposite signs;reveals/move flat 1.1520 vs 1.1536paired delta vs reference: 3,204 points95% lower bound: -26,860 pointsn = 256 gamessource: RS-20260823T131226Z-16564ed9 · metrics.A900_minus_frozenheadline paired delta (points) — rollout veto 17k | 25-move 7-scenario D2-continuation veto, SCREEN; -21,887 points | per veto taken; all five gate conditions failed; no floor | recorded | paired delta vs reference: -46,511 points | 95% lower bound: -91,925 points | n = 32 games | source: docs/exploratory/finding-03-rollout-veto-17k.md · | section 5.2 paired whole-game comparisonheadline paired delta (points) — rollout veto 17k25-move 7-scenario D2-continuation veto, SCREEN; -21,887 pointsper veto taken; all five gate conditions failed; no floorrecordedpaired delta vs reference: -46,511 points95% lower bound: -91,925 pointsn = 32 gamessource: docs/exploratory/finding-03-rollout-veto-17k.md ·section 5.2 paired whole-game comparisonheadline paired delta (points) — leaf refit to clears | t2-fair-a1, the full fair-play refit; monotone loss across six | arms; 7-0-57; no floor recorded | paired delta vs reference: -237,182 points | 95% lower bound: -290,406 points | n = 64 games | source: docs/exploratory/finding-14-leaf-reweight.md · section 5 | paired-delta tableheadline paired delta (points) — leaf refit to clearst2-fair-a1, the full fair-play refit; monotone loss across sixarms; 7-0-57; no floor recordedpaired delta vs reference: -237,182 points95% lower bound: -290,406 pointsn = 64 gamessource: docs/exploratory/finding-14-leaf-reweight.md · section 5paired-delta tableheadline paired delta (points) — CMA-ES leaf | distribution mean after 40 generations, held-out d4s5 screen on | never-read games; below its 40,597 floor | paired delta vs reference: -30,300 points | bounds: -70,928 points to 9,786 points | n = 64 games | source: RS-20260822T120736Z-662b39ca · | metrics.heldOutD4S5.pairedScoreheadline paired delta (points) — CMA-ES leafdistribution mean after 40 generations, held-out d4s5 screen onnever-read games; below its 40,597 floorpaired delta vs reference: -30,300 pointsbounds: -70,928 points to 9,786 pointsn = 64 gamessource: RS-20260822T120736Z-662b39ca ·metrics.heldOutD4S5.pairedScoreheadline paired delta (points) — survival-instinct STRICT | 128 paired games; inside its 24,999 floor; outcome inconclusive | (LITERAL arm -97,064 [-136,887, -59,500] is the rejection) | paired delta vs reference: -1,970 points | bounds: -27,738 points to 22,313 points | n = 128 games | source: RS-20260822T233343Z-12becce9 · | metrics.strict.pairedScoreheadline paired delta (points) — survival-instinct STRICT128 paired games; inside its 24,999 floor; outcome inconclusive(LITERAL arm -97,064 [-136,887, -59,500] is the rejection)paired delta vs reference: -1,970 pointsbounds: -27,738 points to 22,313 pointsn = 128 gamessource: RS-20260822T233343Z-12becce9 ·metrics.strict.pairedScoreheadline paired delta (points) — terminal utility ×50 | EXACT zero: magnitudes -3M/-10M/-50M play byte-identical games | to the frozen -1M (0-64-0); parameter saturated; no bound exists | paired delta vs reference: 0 points | n = 64 games | source: | docs/exploratory/finding-04-terminal-utility-saturated.md · | Result tableheadline paired delta (points) — terminal utility ×50EXACT zero: magnitudes -3M/-10M/-50M play byte-identical gamesto the frozen -1M (0-64-0); parameter saturated; no bound existspaired delta vs reference: 0 pointsn = 64 gamessource:docs/exploratory/finding-04-terminal-utility-saturated.md ·Result tabledetection floor 1.645·sd/√n (where recorded) — depth 5 vs 4 · s7 | floor for d5s7 - d4s7 | paired delta vs reference: 107,988 points | n = 32 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.powerTable[5].detectionFloordetection floor 1.645·sd/√n (where recorded) — depth 5 vs 4 · s7floor for d5s7 - d4s7paired delta vs reference: 107,988 pointsn = 32 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.powerTable[5].detectionFloordetection floor 1.645·sd/√n (where recorded) — depth 5 vs 4 · s5 | floor for d5s5 - d4s5 | paired delta vs reference: 47,052 points | n = 64 games | source: RS-20260821T205102Z-d89df4b5 · | metrics.powerTable[4].detectionFloordetection floor 1.645·sd/√n (where recorded) — depth 5 vs 4 · s5floor for d5s5 - d4s5paired delta vs reference: 47,052 pointsn = 64 gamessource: RS-20260821T205102Z-d89df4b5 ·metrics.powerTable[4].detectionFloordetection floor 1.645·sd/√n (where recorded) — reveal-constr. | A900 | floor for the A900 paired delta | paired delta vs reference: 30,957 points | n = 256 games | source: RS-20260823T131226Z-16564ed9 · | metrics.A900_minus_frozen.detectionFloordetection floor 1.645·sd/√n (where recorded) — reveal-constr.A900floor for the A900 paired deltapaired delta vs reference: 30,957 pointsn = 256 gamessource: RS-20260823T131226Z-16564ed9 ·metrics.A900_minus_frozen.detectionFloordetection floor 1.645·sd/√n (where recorded) — CMA-ES leaf | floor for the CMA-ES held-out paired delta | paired delta vs reference: 40,597 points | n = 64 games | source: RS-20260822T120736Z-662b39ca · | metrics.heldOutD4S5.pairedScore.detectionFloordetection floor 1.645·sd/√n (where recorded) — CMA-ES leaffloor for the CMA-ES held-out paired deltapaired delta vs reference: 40,597 pointsn = 64 gamessource: RS-20260822T120736Z-662b39ca ·metrics.heldOutD4S5.pairedScore.detectionFloordetection floor 1.645·sd/√n (where recorded) — survival-instinct | STRICT | floor for the STRICT paired delta | paired delta vs reference: 24,999 points | n = 128 games | source: RS-20260822T233343Z-12becce9 · | metrics.strict.pairedScore.detectionFloordetection floor 1.645·sd/√n (where recorded) — survival-instinctSTRICTfloor for the STRICT paired deltapaired delta vs reference: 24,999 pointsn = 128 gamessource: RS-20260822T233343Z-12becce9 ·metrics.strict.pairedScore.detectionFloorheadline paired delta (points)detection floor 1.645·sd/√n (where recorded)
Every closed direction whose headline is a points number, one row each, with its recorded bound and, where the record states one, its detection floor. A dot whose magnitude sits inside its floor is a non-measurement, not a rejection. Hover a marker for the recorded note; the non-score closures are quoted verbatim in the figure's notes.
Source data

One dot per closed direction whose headline result is a points number, copied verbatim from its record; whiskers are the recorded bounds — one-sided 95% bootstrap lower bounds where that is all the record carries (both depth rows, reveal-construction A900, rollout veto, leaf refit), two-sided 95% where recorded (reveal M=2, CMA-ES, survival-instinct STRICT), and none for terminal utility, whose zero is EXACT (byte-identical games at four magnitudes from -1,000,000 to -50,000,000, 0-64-0; no statistical bound exists or is needed). Detection-floor markers (1.645*sd/sqrt(n)) are drawn only where the record carries one; no floor is recorded for reveal M=2, the rollout veto, or the leaf refit — markers omitted, not zero. Every closure rejects the exact tested configuration at development/pilot tier; none proves the underlying idea impossible, and each row's 'reopens if' condition stays in docs/research/status.md. The plan's 'depth beyond four plies' row is charted as its two recorded contrasts (s7 n=32 stopped/final, s5 n=64). Not chartable on a points axis, recorded here verbatim instead: compact afterstate model reproducing D4's ordering — top-1 0.375 against the 0.60 gate (RS-20260821T104500Z-77d21e90); the suite-h9-v1 benchmark as a strength measure — Spearman -0.257 (finding-10). status.md's table gained three further non-points closures after this figure's plan was written, also not chartable here: leaf-cost NNUE student top-1 0.296/0.3011 by half-fold against the 0.60 gate (RS-20260823T194142Z-946e3cd1), H-pool optimistic states tau -0.959 [cluster bootstrap -3.069, -0.390] (RS-20260823T205143Z-ead14c9d), and P-SOL G0 cheap-label proxy mean tau 0.370, LB95 0.283 (RS-20260823T225753Z-0fbd48c3). Cohorts differ per point (n on every marker); this chart assembles separate experiments, it does not pool them.

Seriesclosed directionpaired delta vs referenceBoundsnSource
headline paired delta (points)depth 5 vs 4 · s7 d5s7 - d4s7; arm stopped at 32 of 64 games, final; 35.62x work; below its 107,988 floor: a non-measurement23,367 pointslower -83,04632RS-20260821T205102Z-d89df4b5 metrics.pairedD5s7MinusD4s7
headline paired delta (points)depth 5 vs 4 · s5 d5s5 - d4s5; 23.29x work; below its 47,052 floor-8,624 pointslower -55,13464RS-20260821T205102Z-d89df4b5 metrics.pairedD5s5MinusD4s5
headline paired delta (points)reveal M=2 on ply 4 d4 M=2 - M=1; 4.07x work; only 28.6% joint coverage was affordable at depth 4; no floor recorded-41,950 points-100,137 to 17,54164RS-20260821T192140Z-189fe392 metrics.pairedD4M2MinusD4M1
headline paired delta (points)reveal-constr. A900 aligned_double_hit +900 on 256 fresh paired games; inside its 30,957 floor; halves +48,762 / -42,354 opposite signs; reveals/move flat 1.1520 vs 1.15363,204 pointslower -26,860256RS-20260823T131226Z-16564ed9 metrics.A900_minus_frozen
headline paired delta (points)rollout veto 17k 25-move 7-scenario D2-continuation veto, SCREEN; -21,887 points per veto taken; all five gate conditions failed; no floor recorded-46,511 pointslower -91,92532docs/exploratory/finding-03-rollout-veto-17k.md section 5.2 paired whole-game comparison
headline paired delta (points)leaf refit to clears t2-fair-a1, the full fair-play refit; monotone loss across six arms; 7-0-57; no floor recorded-237,182 pointslower -290,40664docs/exploratory/finding-14-leaf-reweight.md section 5 paired-delta table
headline paired delta (points)CMA-ES leaf distribution mean after 40 generations, held-out d4s5 screen on never-read games; below its 40,597 floor-30,300 points-70,928 to 9,78664RS-20260822T120736Z-662b39ca metrics.heldOutD4S5.pairedScore
headline paired delta (points)survival-instinct STRICT 128 paired games; inside its 24,999 floor; outcome inconclusive (LITERAL arm -97,064 [-136,887, -59,500] is the rejection)-1,970 points-27,738 to 22,313128RS-20260822T233343Z-12becce9 metrics.strict.pairedScore
headline paired delta (points)terminal utility ×50 EXACT zero: magnitudes -3M/-10M/-50M play byte-identical games to the frozen -1M (0-64-0); parameter saturated; no bound exists0 points64docs/exploratory/finding-04-terminal-utility-saturated.md Result table
detection floor 1.645·sd/√n (where recorded)depth 5 vs 4 · s7 floor for d5s7 - d4s7107,988 points32RS-20260821T205102Z-d89df4b5 metrics.powerTable[5].detectionFloor
detection floor 1.645·sd/√n (where recorded)depth 5 vs 4 · s5 floor for d5s5 - d4s547,052 points64RS-20260821T205102Z-d89df4b5 metrics.powerTable[4].detectionFloor
detection floor 1.645·sd/√n (where recorded)reveal-constr. A900 floor for the A900 paired delta30,957 points256RS-20260823T131226Z-16564ed9 metrics.A900_minus_frozen.detectionFloor
detection floor 1.645·sd/√n (where recorded)CMA-ES leaf floor for the CMA-ES held-out paired delta40,597 points64RS-20260822T120736Z-662b39ca metrics.heldOutD4S5.pairedScore.detectionFloor
detection floor 1.645·sd/√n (where recorded)survival-instinct STRICT floor for the STRICT paired delta24,999 points128RS-20260822T233343Z-12becce9 metrics.strict.pairedScore.detectionFloor

Spec: web/content/figures/closed-directions-map.json · 8 source records

DirectionResultReopens if
Search depth beyond four pliesNot measurable: +23,367 (7 strata, n=32) and −8,624 (5 strata, n=64), both inside their detection floors (RS-20260821T205102Z-d89df4b5)A design that can resolve sub-50k effects, or a mechanism predicting an effect above the floor
Stacking reveal sampling on top of the fourth ply−41,950 at M=2 (RS-20260821T192140Z-189fe392)A wider dose is tested; only 28.6% joint coverage was affordable at depth 4, and the dose that worked at depth 3 was 85.7%
Pricing the same-wave double hit on a solid gray (reveal construction) as a leaf termCorpus: partial r −0.044 with 60% of live setups uncollected; in play, +3,204 over 256 fresh paired games, inside its floor (RS-20260823T131226Z-16564ed9)A term that raises reveals per move at all; or a per-root counterfactual showing the uncollected setups were worth collecting
Harsher terminal (death) utilitySaturated: byte-identical play at every magnitude past the current value (finding-04)Never, for this parameterization
A leaf-cost NNUE student reproducing D4's within-root ordering from successor-closed D4 valuesTop-1 0.296/0.301 by half-fold against the 0.60 gate (RS-20260823T194142Z-946e3cd1)A different architecture class at leaf cost, e.g. a cross-sibling set ranker; the self-play loop's leaf form stays blocked and its redesign is root-prior shaped
Optimistic states with fair labels (H-pool, the salvageable core of the oracle curriculum)Stage D0: tau = −0.959 [−3.069, −0.390], with a degenerate denominator — 63/64 oracle games hit the 500-move cap (RS-20260823T205143Z-ead14c9d)A D0 rerun with an uncapped horizon shows tau ≥ 0.25 and the oracle's action at least matching D4's under fair futures
Cheap M=1 continuation labels standing in for D3 N7M6 label semantics (P-SOL G0)Guardrail kill: within-root orderings agree at only mean tau 0.370 (LB95 0.283) on 6 CRN-matched roots; the divergence is the reveal quadrature, not the engine (RS-20260823T225753Z-0fbd48c3)Reopened 2026-08-24: E-FAST-M6 passed every equivalence gate (RS-20260824T010000Z-8f3e9b4f) at a realised 5.6x speedup (0.177 CPU-s/move in continuation duty), so a powered M=6 guardrail is now affordable; a full M=6 label corpus still costs thousands of CPU-hours and needs a P-SOL-3 design or a scale-out lease
25-move rollout veto−21,887 points per veto (finding-03)A lower-variance rollout estimator
Refitting the leaf toward achievable clear rateMonotone loss across six arms; fully-fitted vector −237,182 (finding-14)A refit against remaining lifetime rather than achievable clears
Direct action override by the compact afterstate modelOverride gate failed narrowly, and full training was harmful in one half-fold (RS-20260820T184500Z-63c0a8e2, RS-20260821T094500Z-1a7e3c55)A materially larger student, or a stronger teacher than D1-continuation
Compact afterstate model reproducing D4's orderingTop-1 0.375 against a 0.60 gate; exact D1 scores 0.486 (RS-20260821T104500Z-77d21e90)A materially larger model; the claim is only established for compact evaluators
The suite-h9-v1 scenario benchmark as a strength measureRanked policies backwards, Spearman −0.257 (finding-10)A validated longer horizon; it remains usable as a diagnostic

The two afterstate rows matter more than their tier suggests. The program's standing explanation for every failed learned policy was insufficient sibling coverage. The 2026-08-21 experiment supplied successor-closed coverage, every legal sibling, exact search-value labels, and completeness 1.0 — and the student still ranked worse than a one-ply exact search. That relocates the obstacle from the training data to the capacity of a compact board evaluator, and it makes "train a bigger student" a falsifiable next step rather than a hopeful one.

Open work

The next research step should be evidence-driven rather than another broad architecture sweep:

  1. Re-register any resumed SHA-locked experiment against the reorganized source tree and rerun cross-engine parity.
  2. Refit the search leaf against remaining lifetime, the quantity that correlates with score at r = 0.9995, rather than against achievable clears.
  3. Test whether student capacity is the binding constraint, since that is now the explicit claim on the table and the cheapest one to falsify.
  4. Test long-cycle features as bounded corrections to D4 before allowing them to control an entire game.
  5. Resolve the three recorded divergences between this simulator and the shipped game, which is an owner decision and not an agent decision.

Items 1, 4 and 5 are unchanged. Items 2 and 3 replace the AFBR-40 closure sequence, whose data-feasibility question was answered directly: closure was achieved and did not rescue the model.

AFBR-40 as originally proposed has no implementation, protocol, model, or measurements in this repository and must not appear in a list of attempted results. The action-free afterstate models that were built and measured on 2026-08-20/21 are separate, separately recorded, and are not AFBR-40.

For configurations and every retained entry point, use the experiment index. For the unabridged chronological record, use the experiment history. For alternatives and research priority, use the strategy landscape and the staged research roadmap. Future agents must use the standard benchmark contract and machine-readable records under research/; these do not retroactively upgrade the historical evidence above.

For a walkthrough with board animations, start at how the game works and the concepts primer; every term is defined in the glossary.