On this page
Dates
Created
Updated
Record idEX-20260906-ntuple-scale-replication-wide-plateau-f627f07a

No explanation has been written for this record yet.

Technical recordThe registered protocolEX-20260906-ntuple-scale-replication-wide-plateau-f627f07a
Hypothesis
Candidate: the stock fair expectimax search at depth 3, seven chance strata, terminal utility -1,000,000, policy seed 0xd7075eed, completion-guaranteeing work bound work_bound_for(3,7)+1, 64k-entry direct-mapped table (the deployment configuration of EX-20260905-ntuple-scale-tc-td-leaf-d3-535b2620, unchanged), with the n-tuple lookup-table leaf of that experiment made wider: the sum, over the active patterns of the reflection-canonical board, of learned f32 entries in rise units (x 17,000 points) from the layout rows,cols,win23,win32,win24,win42,phase=all, that is the first experiment's selected layout (absolute-position seven-cell row and column tuples, 10^7 patterns each, and the thirty 2x3 and thirty 3x2 window tuples, 10^6 patterns each, every one of these families conditioned on the moves remaining until the rise, x5) plus two new families that are never phase-conditioned: win24, the 24 absolute placements of a 2-wide x 4-tall window (10^8 patterns each), and win42, the 24 placements of a 4-wide x 2-tall window (10^8 patterns each). 5.8 x 10^9 table entries, 122 active per state, 23.2 GB frozen and 69.6 GB with the coherence accumulators. The visible next disc is not an input. Weights: fresh tables initialised to 20 rise units spread over the active entries, trained on-policy by TD(0) with temporal-coherence per-entry step sizes (alpha 1.0 x |E|/A, error clamped to +-30 rise units) from complete games played by the greedy one-ply chance-state policy over the same tables (seven stratified reveal samples per column from the public sampler, epsilon 0) on the Rust bitboard engine, 32 asynchronous workers sharing the tables without locks (Hogwild; not bit-reproducible across runs, disclosed; the gate binary proves the serial update rule is deterministic). Pattern indices are 64-bit (the wide layout exceeds 2^32 entries). Stage 0 (CHECK, no leased seed): gate --probe-start 0xa5277000 --layout <the wide layout> on the already-opened Rust-engine probe block, 19 gates including feature-index references for win24 and win42 (all passed 2026-09-06T01:38Z before any leased seed was read), and a 4 x 10^7-move throughput smoke run on the same probe block with the tables discarded, which measured 1.27 x 10^6 training moves per second on 32 threads and a peak resident set of 42.7 GB after 4 x 10^7 moves (the accumulators' pages become resident as they are touched; the bound is 69.6 GB of tables plus working memory). No pilot: the configuration is fixed here from the first experiment's pilot and its widening is the treatment. Stage B (main run, training-role seeds): the wide layout trained from fresh tables on the never-read training lease SL-20260906T013222Z-ce319097 (0xa5500000, 2,031,616 seeds read in order, wrapping, wrap count recorded), chunk 5 x 10^7 moves, a validation point every 5 x 10^8 moves on the never-read 256-game training-role validation block SL-20260906T013222Z-6cce6192 (0xa52f2280) with the paired line-up ntuple-d3s7 / ntuple-1ply / fair-d3s7 on identical seeds, a quick direct-play check every 10^8 moves, a checkpoint every fourth validation point and at the end. Stop rule, fixed here: the run stops at the first validation point k (counting from 1) with k >= 8 at which the mean paired margin (ntuple-d3s7 minus fair-d3s7) of the last four points is not above the mean of the four points before them (train.rs plateau_check, window 4, minimum 8 points); or at 4 x 10^10 moves; or at 43,200 s of training wall time; or when a STOP file appears in the run directory (an operator abort, recorded as such). The candidate is best-weights.bin, the tables at the validation point with the largest paired margin (ties to the later point). Stage C (freeze): SHA-256 of best-weights.bin recorded as candidate-weights.sha256, and the SHA-256 of the first experiment's frozen candidate (runs/RUN-20260905T193006Z-4fbeb4e5/ntuple-scale/main/best-weights.bin) verified equal to 0ade9d4e4080ebdd52a1474b1a13410dc8dfb77f5eba24b078aa7703c92ace0b and recorded as prior-weights.sha256, both before the screen lease opens. Stage D (SCREEN, public-development, one-shot): 512 paired games on SL-20260906T013222Z-f34b23c0 (0xa52f2380-0xa52f2580) with six arms on identical seeds: candidate-d3s7 (the wide candidate as the leaf), candidate-1ply (the same tables played directly, diagnostic), prior-d3s7 (the first experiment's frozen tables as the same leaf: the fresh-block replication arm), prior-1ply (diagnostic), fair-d3s7 (the comparator: the identical search with the frozen fair leaf) and fair-d4s7 (the program's standing reference, diagnostic). Three readings are fixed here: the candidate gate (candidate-d3s7 vs fair-d3s7, the pass criteria below), the replication (prior-d3s7 vs fair-d3s7, the same four statistical criteria, reported as pass or fail beside the candidate gate and not folded into it) and the scale verdict (candidate-d3s7 vs prior-d3s7: supported when both the bootstrap and the Student-t 95% lower bounds are above zero, refuted when the bootstrap 95% upper bound is below zero, inconclusive otherwise; reported, not a gate). Every stage runs through approaches/ntuple-rl/ntuple-scale/scripts/pipeline2.sh, which hard-codes the leased sub-blocks and records per-stage wall, CPU and peak-RSS usage; scripts/pipeline.sh, the first experiment's driver, is not touched. Expected cost from the smoke run: about 6.6 minutes of training per validation interval plus about a minute of validation, so the eight-point minimum is about an hour and the 4 x 10^10-move cap about eleven hours; the screen's fair-d4s7 arm is about 50 minutes.
Arms
ArmNameEntry pointManifest
Candidated3s7-ntuple-scale-wide-leafapproaches/ntuple-rl/ntuple-scale/src/bin/train.rs
Comparatorfair-d3s7approaches/fair-expectimax/rust-engine/src/leaf.rsresearch/benchmarks/baselines-v1.json
Classification
algorithmic
Information boundary
public-policy
Benchmark tier
SCREEN
Lifecycle
completed
Primary metric
paired mean whole-game score delta, candidate-d3s7 minus fair-d3s7, both at the identical d3s7 configuration, 512 held-out games
Secondary metrics
  • replication: paired mean score delta prior-d3s7 minus fair-d3s7 on the same 512 games, with the same bounds, halves, W-T-L and lower quartile
  • scale: paired mean score delta candidate-d3s7 minus prior-d3s7 on the same 512 games, with bootstrap and Student-t bounds, detection floor and W-T-L, and the fixed three-way verdict
  • paired mean moves delta; numbered clears per move and cover reveals per move; mean occupied cells; lower quartile, median, maximum score; first-half (seeds 0-255) and second-half (256-511) paired mean deltas
  • diagnostic arms on the same seeds: candidate-1ply and prior-1ply (the tables played directly one ply, seven reveal strata) and fair-d4s7 (the standing reference)
  • training-curve diagnostics: per-validation-point paired margin of ntuple-d3s7 and ntuple-1ply over fair-d3s7 on the 256-game training-role block with bootstrap bounds, the plateau rule's two window means at every point from the eighth, training-game mean score and length, mean |TD error|, mean coherence, touched table entries, moves per second, seed wraps
  • resource observations: per-stage wall, CPU and peak resident set; the smoke run's throughput
Statistical unit
whole-game
Uncertainty method
one-sided 95% percentile bootstrap over whole games, 20,000 resamples, RNG seed 0xb0071eaf (the unchanged compare.py of the leaf-evolution screen), plus a one-sided 95% Student-t lower bound; detection floor 1.645*sd/sqrt(n) reported; the same statistics for the replication and scale contrasts
Data role
public-development
Seed leases
  • SL-20260906T013222Z-ce319097
  • SL-20260906T013222Z-6cce6192
  • SL-20260906T013222Z-f34b23c0
Whole-origin split
yes
Reuse disclosure
CHECK gates and the throughput smoke run read only the already-opened development probe block 0xa5276000-0xa5277fff (the Rust engine's gate/benchmark range); no new seed is opened for mechanics or timing. Training reads the never-read training lease 0xa5500000-0xa56f0000 in order and wraps when the block is exhausted (wrap counts recorded); the never-read validation block 0xa52f2280-0xa52f2380 (training role) is re-read at every validation point and drives the plateau rule and the candidate choice; the held-out screen opens 0xa52f2380-0xa52f2580 exactly once, after both candidate hashes are recorded. All three ranges are disjoint from each other and from every block the first experiment read (training 0xa5300000-0xa54f0000, validation 0xa52f2240-0xa52f2280, screen 0xa52f2140-0xa52f2240), so the prior arm's tables have never seen any seed of this experiment. The prior tables themselves are re-used as a frozen artifact (SHA-256 verified), not re-trained.
Pass criteria
  1. All CHECK gates passed before the first leased seed was read: codec against Horner, row gather against the cell accessor, feature indices against the independent reference for every layout including win24, win42 and the wide layout, information-boundary blindness (score/level/moves_played/next_disc), reflection consistency of values and decisions, direct-policy legality, leaf-in-search determinism and worker-count independence, serial training determinism, finiteness.
  2. Every validation artifact and the screen artifact have illegalDecisions 0 and incompleteDecisions 0; censored games are reported as censored.
  3. Held-out screen, 512 paired games on 0xa52f2380-0xa52f257f, candidate-d3s7 vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0.
  4. Held-out screen: candidate-d3s7 minus fair-d3s7 paired mean score delta > 0 in both halves (seeds 0-255 and 256-511).
  5. Held-out screen: candidate-d3s7 lower-quartile score >= fair-d3s7 lower-quartile score.
  6. The screened candidate is the exact frozen best-weights.bin of the main run (candidate-weights.sha256 recorded before the screen lease opens) and the prior arm is the exact frozen candidate of the first experiment (prior-weights.sha256 verified before the screen lease opens); no other table is screened.
  7. Reported beside the gate, not part of it: the replication passes when prior-d3s7 vs fair-d3s7 meets the same four statistical criteria on the same 512 games; the scale verdict is supported / refuted / inconclusive as defined in the protocol.
On pass
Freeze best-weights.bin with its SHA-256; record the replication and scale readings; register a fresh-development STANDARD evaluation of whichever frozen candidate the scale verdict favours (the first when inconclusive) as the successor; do not open protected or final seeds.
On fail
Record a valid run with scientific outcome fail for this exact configuration (wide layout, optimiser, plateau rule, budget, deployment depth); do not adopt the tables; report the replication and scale readings and the diagnostic arms as recorded; if the replication also fails, mark the theory's assessment as not replicated on a fresh block; open no further cohort for this candidate.
Gate fixed before controlled data
yes
Resources
Wall seconds57600
CPU threads32
Max host bytes103079215104
Max GPU bytes
GPU devices
Stop conditions
  1. Stop on any rules, information-boundary, legality, determinism, or parity failure.
  2. The main run stops by the plateau rule (window 4, from the eighth validation point, last-four mean not above the previous-four mean), at 4 x 10^10 moves, at 43,200 s of training wall time, or on a STOP file, whichever comes first; the reason is written to stop.json and the candidate is then the best validation point so far.
  3. A validation or screen artifact with any illegal or incomplete decision voids the run (invalid), not the candidate.
  4. The screen is evaluated exactly once; no re-run on the same or a different held-out block without a new experiment record.
  5. The theory's training-signal falsifier (no validation point with a positive paired margin) is evaluated and reported separately; the screen is opened once on the frozen candidate regardless.
  6. Host memory above the declared bound (96 GiB resident) aborts the main run through the operating system or the operator; such a run is recorded as interrupted with its artifacts.
Expected artifacts
  • runs/<run-id>/ntuple-scale/gates.log (CHECK)
  • runs/<run-id>/ntuple-scale/smoke/{config.json,progress.jsonl,DONE} (CHECK throughput; tables discarded)
  • runs/<run-id>/ntuple-scale/main/{config.json,progress.jsonl,val-*.json,best.json,best-weights.bin,latest-weights.bin,checkpoint.bin,stop.json,DONE} (stage B)
  • runs/<run-id>/ntuple-scale/main/{candidate-weights.sha256,prior-weights.sha256} (stage C)
  • runs/<run-id>/ntuple-scale/screen/heldout.json + compare-*.json (stage D)
  • runs/<run-id>/ntuple-scale/{pipeline.log,rusage.jsonl,analysis.json,analysis.md}
  • artifacts emitted by train and screen carry this experiment id in their config field; table files (23.2 GB weights, 69.6 GB checkpoint) are too large for the repository and are retained on the workstation with their SHA-256 in the result record
Amendments
TimestampBefore controlled dataReason
2026-09-06T04:01:14Znolifecycle advanced running -> completed after run RUN-20260906T013222Z-ba0ee34f and result RS-20260906T040113Z-6ba93171 were written; protocol content otherwise unchanged (frozen hash 0284abf064ef24fe47ae1ee1b3a7d63245b9f470bdcdcbf00976576bb1d67df8 retained in the run and lease records).
Technical recordResults recorded against this protocol1 record
valid runoutcome: passsupported-as-testedtier: public-developmentRS-20260906T040113Z-6ba93171

Held-out screen, 512 never-read paired public-development games (0xa52f2380+): the frozen n-tuple tables as the leaf of the depth-3 seven-stratum fair search averaged 481,869 points and 139.39 moves against 326,717 points and 95.87 moves for the identical search with the frozen fair leaf: paired +155,153 points (bootstrap 95% lower bound +126,819, Student-t lower bound +126,919, upper bound +183,307, detection floor 28,185), W-T-L 333-0-179, halves +159,105 / +151,201, lower quartile 241,610 vs 192,040, moves +43.52. The preregistered gate PASSES. Replication: the first experiment's frozen tables (SHA-256 0ade9d4e4080ebdd52a1474b1a13410dc8dfb77f5eba24b078aa7703c92ace0b) as the same leaf on these fresh seeds averaged 487,066 points and 140.83 moves; prior-d3s7 minus fair-d3s7: paired +160,349 (bootstrap 95% lower bound +129,753, Student-t lower bound +129,626, upper bound +191,264, detection floor 30,670), W-T-L 330-0-182, halves +133,349 / +187,349. The replication criteria PASS. Scale: the wider, longer-trained candidate against the first candidate on the same seeds is -5,196 paired (bootstrap LB -40,535, t LB -40,060, UB +29,158, floor 34,803, W-T-L 267-1-244): the preregistered scale verdict is inconclusive. Against the program's standing reference, the fair leaf at depth 4 on the same seeds (397,154 points), the candidate is +84,716 paired (bootstrap LB +54,794, UB +114,355, W-T-L 298-0-214); diagnostic only. The first candidate against fair-d4s7 on these seeds: +89,912 (LB +58,243); diagnostic. The same tables played directly one ply averaged 328,039 points, +1,323 paired against fair-d3s7 (LB -17,545); diagnostic. Fair-d4s7 minus fair-d3s7 on these seeds: +70,437 (LB +46,917). Candidate: the tables of the main run's validation point at 2,000,153,332 training moves (layout rows,cols,win23,win32,win24,win42,phase=all, alpha 1, 5,800,000,000 entries), whose paired margin on the 256-game training-role validation block was +187,500; SHA-256 824b0a39a90d8a5aae63438c1538d4c5022f0e0f09c75fb7e2d758a1d6c6fb8a. Main run: 4,500,370,590 training moves, 47,477,538 games, 9 validation points, mean 1,184,973 moves per second, stopped by the plateau rule (last 4 validation points mean +165,666, the 4 before them +173,783). Training-signal check: at least one validation point of the main run had a positive paired margin.

What it had to pass
  • All CHECK gates passed before the first training seed was read — observed: 18 gate lines, all PASS
  • Every validation artifact and the screen artifact have illegalDecisions 0 and incompleteDecisions 0 — observed: main validations illegal 0 / incomplete 0; screen arms candidate-d3s7: 0/0, candidate-1ply: 0/0, prior-d3s7: 0/0, prior-1ply: 0/0, fair-d3s7: 0/0, fair-d4s7: 0/0
  • bootstrap 95% lower bound of candidate-d3s7 minus fair-d3s7 > 0 — observed: 126818.78720703124
  • Student-t 95% lower bound > 0 — observed: 126919.162189214
  • paired mean delta > 0 in both halves — observed: [159104.70703125, 151200.68359375]
  • candidate-d3s7 Q25 >= fair-d3s7 Q25 — observed: [241609.5, 192039.5]
  • The screened candidate is the exact frozen best-weights.bin (SHA-256 recorded before the screen lease opened) — observed: 824b0a39a90d8a5aae63438c1538d4c5022f0e0f09c75fb7e2d758a1d6c6fb8a
  • Replication: bootstrap 95% lower bound of prior-d3s7 minus fair-d3s7 > 0 — observed: 129753.3208984375
  • Replication: Student-t 95% lower bound > 0 — observed: 129625.57751269455
  • Replication: paired mean delta > 0 in both halves — observed: [133348.80078125, 187349.16796875]
  • Replication: prior-d3s7 Q25 >= fair-d3s7 Q25 — observed: [232825.0, 192039.5]
  • Replication: the prior arm is the first experiment's exact frozen tables (SHA-256 verified before the screen lease opened) — observed: 0ade9d4e4080ebdd52a1474b1a13410dc8dfb77f5eba24b078aa7703c92ace0b
Technical recordRecorded metricsRS-20260906T040113Z-6ba93171
screen
games
512
seedStartHex
0xa52f2380
arms
candidate-d3s7
meanScore
481869.3887
medianScore
393,375
q25Score
241609.5000
minScore
85,504
maxScore
2,487,485
sdScore
327121.3810
meanMoves
139.3867
q25Moves
70.7500
numberedClearsPerMove
2.1088
coverRevealsPerMove
1.1927
meanOccupiedCells
22.9657
censoredGames
0
illegalDecisions
0
incompleteDecisions
0
gamesAtOrAboveMillion
36
meanWallSecondsPerGame
5.4114
candidate-1ply
meanScore
328039.3203
medianScore
282096.5000
q25Score
178989.7500
minScore
85,684
maxScore
1,224,664
sdScore
190096.3800
meanMoves
97.2695
q25Moves
55
numberedClearsPerMove
1.9762
coverRevealsPerMove
1.0946
meanOccupiedCells
24.1436
censoredGames
0
illegalDecisions
0
incompleteDecisions
0
gamesAtOrAboveMillion
2
meanWallSecondsPerGame
0.0013
prior-d3s7
meanScore
487065.6777
medianScore
385960.5000
q25Score
232,825
minScore
85,504
maxScore
3,236,265
sdScore
381394.4123
meanMoves
140.8340
q25Moves
70
numberedClearsPerMove
2.1123
coverRevealsPerMove
1.1949
meanOccupiedCells
23.0757
censoredGames
0
illegalDecisions
0
incompleteDecisions
0
gamesAtOrAboveMillion
43
meanWallSecondsPerGame
4.2504
prior-1ply
meanScore
294323.1445
medianScore
247,261
q25Score
173120.7500
minScore
85,362
maxScore
1,270,003
sdScore
179592.4792
meanMoves
87.9121
q25Moves
54.7500
numberedClearsPerMove
1.9353
coverRevealsPerMove
1.0665
meanOccupiedCells
24.3244
censoredGames
0
illegalDecisions
0
incompleteDecisions
0
gamesAtOrAboveMillion
5
meanWallSecondsPerGame
0.0009
fair-d3s7
meanScore
326716.6934
medianScore
269,648
q25Score
192039.5000
minScore
103,026
maxScore
1,714,793
sdScore
204736.3872
meanMoves
95.8652
q25Moves
56.7500
numberedClearsPerMove
2.0033
coverRevealsPerMove
1.1138
meanOccupiedCells
24.0979
censoredGames
0
illegalDecisions
0
incompleteDecisions
0
gamesAtOrAboveMillion
7
meanWallSecondsPerGame
5.1310
fair-d4s7
meanScore
397153.8184
medianScore
320,815
q25Score
202359.7500
minScore
102,858
maxScore
1,707,841
sdScore
264186.5566
meanMoves
114.5469
q25Moves
60
numberedClearsPerMove
2.0598
coverRevealsPerMove
1.1550
meanOccupiedCells
23.7587
censoredGames
0
illegalDecisions
0
incompleteDecisions
0
gamesAtOrAboveMillion
19
meanWallSecondsPerGame
196.6131
contrasts
candidate-d3s7-vs-fair-d3s7
score
n
512
meanDelta
155152.6953
pairedSd
387690.2710
bootstrapLower95
126818.7872
bootstrapUpper95
183307.3728
studentTLower95
126919.1622
detectionFloor
28184.8563
wins
333
ties
0
losses
179
firstHalfMeanDelta
159104.7070
secondHalfMeanDelta
151200.6836
q25Delta
49,570
candidateQ25
241609.5000
referenceQ25
192039.5000
moves
n
512
meanDelta
43.5215
pairedSd
106.8674
bootstrapLower95
35.7206
bootstrapUpper95
51.2736
studentTLower95
35.7389
detectionFloor
7.7692
wins
331
ties
9
losses
172
firstHalfMeanDelta
44.6016
secondHalfMeanDelta
42.4414
q25Delta
14
candidateQ25
70.7500
referenceQ25
56.7500
candidate-1ply-vs-fair-d3s7
score
n
512
meanDelta
1322.6270
pairedSd
257859.8080
bootstrapLower95
-17,545
bootstrapUpper95
20096.1094
studentTLower95
-17456.0063
detectionFloor
18746.2574
wins
268
ties
0
losses
244
firstHalfMeanDelta
-21520.6602
secondHalfMeanDelta
24165.9141
q25Delta
-13049.7500
candidateQ25
178989.7500
referenceQ25
192039.5000
moves
n
512
meanDelta
1.4043
pairedSd
71.1180
bootstrapLower95
-3.7833
bootstrapUpper95
6.5742
studentTLower95
-3.7749
detectionFloor
5.1702
wins
263
ties
19
losses
230
firstHalfMeanDelta
-4.8789
secondHalfMeanDelta
7.6875
q25Delta
-1.7500
candidateQ25
55
referenceQ25
56.7500
candidate-d3s7-vs-fair-d4s7
score
n
512
meanDelta
84715.5703
pairedSd
412546.1390
bootstrapLower95
54794.3239
bootstrapUpper95
114355.1267
studentTLower95
54671.9093
detectionFloor
29991.8634
wins
298
ties
0
losses
214
firstHalfMeanDelta
92758.7422
secondHalfMeanDelta
76672.3984
q25Delta
39249.7500
candidateQ25
241609.5000
referenceQ25
202359.7500
moves
n
512
meanDelta
24.8398
pairedSd
113.2232
bootstrapLower95
16.6386
bootstrapUpper95
32.9903
studentTLower95
16.5944
detectionFloor
8.2313
wins
293
ties
12
losses
207
firstHalfMeanDelta
27.1406
secondHalfMeanDelta
22.5391
q25Delta
10.7500
candidateQ25
70.7500
referenceQ25
60
fair-d4s7-vs-fair-d3s7
score
n
512
meanDelta
70437.1250
pairedSd
319133.9625
bootstrapLower95
46916.5703
bootstrapUpper95
93486.4364
studentTLower95
47196.2031
detectionFloor
23200.8527
wins
302
ties
0
losses
210
firstHalfMeanDelta
66345.9648
secondHalfMeanDelta
74528.2852
q25Delta
10320.2500
candidateQ25
202359.7500
referenceQ25
192039.5000
moves
n
512
meanDelta
18.6816
pairedSd
86.7909
bootstrapLower95
12.2983
bootstrapUpper95
24.9356
studentTLower95
12.3611
detectionFloor
6.3096
wins
289
ties
21
losses
202
firstHalfMeanDelta
17.4609
secondHalfMeanDelta
19.9023
q25Delta
3.2500
candidateQ25
60
referenceQ25
56.7500
candidate-d3s7-vs-candidate-1ply
score
n
512
meanDelta
153830.0684
pairedSd
361441.8280
bootstrapLower95
127664.9001
bootstrapUpper95
180356.5437
studentTLower95
127508.0774
detectionFloor
26276.6098
wins
333
ties
0
losses
179
firstHalfMeanDelta
180625.3672
secondHalfMeanDelta
127034.7695
q25Delta
62619.7500
candidateQ25
241609.5000
referenceQ25
178989.7500
moves
n
512
meanDelta
42.1172
pairedSd
100.1258
bootstrapLower95
34.8573
bootstrapUpper95
49.4670
studentTLower95
34.8255
detectionFloor
7.2791
wins
330
ties
8
losses
174
firstHalfMeanDelta
49.4805
secondHalfMeanDelta
34.7539
q25Delta
15.7500
candidateQ25
70.7500
referenceQ25
55
prior-d3s7-vs-fair-d3s7
score
n
512
meanDelta
160348.9844
pairedSd
421880.1055
bootstrapLower95
129753.3209
bootstrapUpper95
191264.1133
studentTLower95
129625.5775
detectionFloor
30670.4373
wins
330
ties
0
losses
182
firstHalfMeanDelta
133348.8008
secondHalfMeanDelta
187349.1680
q25Delta
40785.5000
candidateQ25
232,825
referenceQ25
192039.5000
moves
n
512
meanDelta
44.9688
pairedSd
116.3619
bootstrapLower95
36.5253
bootstrapUpper95
53.5156
studentTLower95
36.4947
detectionFloor
8.4594
wins
326
ties
21
losses
165
firstHalfMeanDelta
37.6875
secondHalfMeanDelta
52.2500
q25Delta
13.2500
candidateQ25
70
referenceQ25
56.7500
candidate-d3s7-vs-prior-d3s7
score
n
512
meanDelta
-5196.2891
pairedSd
478728.0439
bootstrapLower95
-40535.3815
bootstrapUpper95
29157.6797
studentTLower95
-40059.6454
detectionFloor
34803.2492
wins
267
ties
1
losses
244
firstHalfMeanDelta
25755.9063
secondHalfMeanDelta
-36148.4844
q25Delta
8784.5000
candidateQ25
241609.5000
referenceQ25
232,825
moves
n
512
meanDelta
-1.4473
pairedSd
132.4515
bootstrapLower95
-11.2169
bootstrapUpper95
8.0607
studentTLower95
-11.0930
detectionFloor
9.6291
wins
261
ties
23
losses
228
firstHalfMeanDelta
6.9141
secondHalfMeanDelta
-9.8086
q25Delta
0.7500
candidateQ25
70.7500
referenceQ25
70
prior-d3s7-vs-fair-d4s7
score
n
512
meanDelta
89911.8594
pairedSd
434700.2797
bootstrapLower95
58243.2331
bootstrapUpper95
121566.8728
studentTLower95
58254.8237
detectionFloor
31602.4564
wins
290
ties
0
losses
222
firstHalfMeanDelta
67002.8359
secondHalfMeanDelta
112820.8828
q25Delta
30465.2500
candidateQ25
232,825
referenceQ25
202359.7500
moves
n
512
meanDelta
26.2871
pairedSd
119.4578
bootstrapLower95
17.5995
bootstrapUpper95
34.9923
studentTLower95
17.5876
detectionFloor
8.6845
wins
285
ties
20
losses
207
firstHalfMeanDelta
20.2266
secondHalfMeanDelta
32.3477
q25Delta
10
candidateQ25
70
referenceQ25
60
prior-1ply-vs-fair-d3s7
score
n
512
meanDelta
-32393.5488
pairedSd
256646.8681
bootstrapLower95
-51189.7492
bootstrapUpper95
-13947.6580
studentTLower95
-51083.8498
detectionFloor
18658.0774
wins
219
ties
0
losses
293
firstHalfMeanDelta
-50575.9883
secondHalfMeanDelta
-14211.1094
q25Delta
-18918.7500
candidateQ25
173120.7500
referenceQ25
192039.5000
moves
n
512
meanDelta
-7.9531
pairedSd
70.7517
bootstrapLower95
-13.1446
bootstrapUpper95
-2.8632
studentTLower95
-13.1056
detectionFloor
5.1436
wins
212
ties
27
losses
273
firstHalfMeanDelta
-12.9609
secondHalfMeanDelta
-2.9453
q25Delta
-2
candidateQ25
54.7500
referenceQ25
56.7500
replication
checks
  1. criterion
    screen artifact: illegalDecisions 0 and incompleteDecisions 0 in every arm
    passed
    true
  2. criterion
    bootstrap 95% lower bound of prior-d3s7 minus fair-d3s7 > 0
    passed
    true
    observed
    129753.3209
  3. criterion
    Student-t 95% lower bound > 0
    passed
    true
    observed
    129625.5775
  4. criterion
    paired mean delta > 0 in both halves
    passed
    true
    observed
    1. 133348.8008
    2. 187349.1680
  5. criterion
    prior-d3s7 Q25 >= fair-d3s7 Q25
    passed
    true
    observed
    1. 232,825
    2. 192039.5000
passed
true
scale
verdict
inconclusive
meanDelta
-5196.2891
bootstrapLower95
-40535.3815
bootstrapUpper95
29157.6797
studentTLower95
-40059.6454
detectionFloor
34803.2492
wins
267
ties
1
losses
244
main
layout
rows,cols,win23,win32,win24,win42,phase=all
alpha
1
entries
5,800,000,000
activePerState
122
validateGames
256
movesTotal
4,500,370,590
gamesTotal
47,477,538
wallSeconds
4357.8000
meanMovesPerSecond
1184973.3333
stop
reason
plateau
movesTotal
4,500,370,590
gamesTotal
47,477,538
wallSeconds
4462.4000
validationPoints
9
plateauWindow
4
recentWindowMean
165665.8408
previousWindowMean
173783.3174
bestMargin
187500.0781
best
moves
2,000,153,332
artifact
val-002000153332.json
pairedDeltaD3
187500.0781
ntupleD3Mean
525039.1602
fairD3Mean
337539.0820
validations
  1. point
    1
    movesTrained
    500,034,572
    ntupleD3Mean
    492111.3789
    directMean
    323747.6328
    fairD3Mean
    337539.0820
    pairedDeltaD3
    154572.2969
    bootstrapLower95
    118307.4332
    wins
    173
    losses
    83
    plateau
  2. point
    2
    movesTrained
    1,000,072,627
    ntupleD3Mean
    483528.0859
    directMean
    310861.9219
    fairD3Mean
    337539.0820
    pairedDeltaD3
    145989.0039
    bootstrapLower95
    106061.5297
    wins
    160
    losses
    96
    plateau
  3. point
    3
    movesTrained
    1,500,111,255
    ntupleD3Mean
    515398.6211
    directMean
    316633.8633
    fairD3Mean
    337539.0820
    pairedDeltaD3
    177859.5391
    bootstrapLower95
    135891.1139
    wins
    171
    losses
    85
    plateau
  4. point
    4
    movesTrained
    2,000,153,332
    ntupleD3Mean
    525039.1602
    directMean
    335958.7422
    fairD3Mean
    337539.0820
    pairedDeltaD3
    187500.0781
    bootstrapLower95
    144166.2436
    wins
    168
    losses
    88
    plateau
  5. point
    5
    movesTrained
    2,500,198,069
    ntupleD3Mean
    521323.7305
    directMean
    336159.7148
    fairD3Mean
    337539.0820
    pairedDeltaD3
    183784.6484
    bootstrapLower95
    138928.7887
    wins
    171
    losses
    85
    plateau
  6. point
    6
    movesTrained
    3,000,240,102
    ntupleD3Mean
    524169.5898
    directMean
    338270.4297
    fairD3Mean
    337539.0820
    pairedDeltaD3
    186630.5078
    bootstrapLower95
    142912.5244
    wins
    169
    losses
    87
    plateau
  7. point
    7
    movesTrained
    3,500,281,390
    ntupleD3Mean
    509645.5117
    directMean
    348955.2148
    fairD3Mean
    337539.0820
    pairedDeltaD3
    172106.4297
    bootstrapLower95
    130275.7250
    wins
    167
    losses
    89
    plateau
  8. point
    8
    movesTrained
    4,000,326,496
    ntupleD3Mean
    480053.3672
    directMean
    363824.5039
    fairD3Mean
    337539.0820
    pairedDeltaD3
    142514.2852
    bootstrapLower95
    106047.1164
    wins
    164
    losses
    92
    plateau
    points
    8
    window
    4
    recentMean
    171258.9678
    previousMean
    166480.2295
    stop
    false
  9. point
    9
    movesTrained
    4,500,370,590
    ntupleD3Mean
    498951.2227
    directMean
    362132.4141
    fairD3Mean
    337539.0820
    pairedDeltaD3
    161412.1406
    bootstrapLower95
    120448.6258
    wins
    159
    losses
    97
    plateau
    points
    9
    window
    4
    recentMean
    165665.8408
    previousMean
    173783.3174
    stop
    true
anyPositiveMargin
true
candidateSha256
824b0a39a90d8a5aae63438c1538d4c5022f0e0f09c75fb7e2d758a1d6c6fb8a
priorSha256
0ade9d4e4080ebdd52a1474b1a13410dc8dfb77f5eba24b078aa7703c92ace0b
smoke
layout
rows,cols,win23,win32,win24,win42,phase=all
entries
5,800,000,000
movesTotal
40,002,639
wallSeconds
33.8000
meanMovesPerSecond
1,240,481
Limitations
  • Public-development SCREEN tier, 512 paired games opened once; nothing here is a qualification claim, and protected and final cohorts stay sealed.
  • The candidate is the validation point with the largest paired margin on a 256-game training-role block that was read at every validation point and drove the plateau stop rule; that selection is upward-biased, which is why the held-out screen exists.
  • Training used lock-free asynchronous updates from 32 threads, so the training run is not bit-reproducible; the frozen tables are hashed and every gameplay arm is deterministic and worker-count independent.
  • The fair-d4s7 arm is the program's standing reference for context only; the preregistered comparator is the identical depth-3 search with the frozen fair leaf.
  • Table files (23.2 GB weights, 69.6 GB with accumulators) are retained on the workstation with their SHA-256 and are not committed.
  • Wall times were measured on a shared workstation; ratios between arms on the same seeds are the trustworthy quantity.
  • The replication arm re-screens tables frozen by the first experiment on a block that experiment never read; it is a fresh-block replication by the same runner on the same machine, not by an independent runner.

Open the result record

Agent contextHow to extend this record

To add a reader-facing explanation, write web/content/research/EX-20260906-ntuple-scale-replication-wide-plateau-f627f07a.mdx; it renders above this record on the next request. The registered protocol itself is in the technical record above.

Record file: research/experiments/EX-20260906-ntuple-scale-replication-wide-plateau-f627f07a.json, validated against research/schemas/experiment-v1.schema.json. Protocol hash: 940282974d1db2720b64d49fba535f43e3e6099dd41723d506971bb65842d3e8.

Registered by Claude Code / claude-fable-5-1 (claude-code-ntuple-scale-replication).