Drop7 Research
← Experiments

C0: paired D3 N7M6 minus D4 s5 contrast from retained per-game records

completedtier CHECKdiagnosticpreviously-evaluated-developmentseed-freeEX-20260823-d3n7m6-vs-d4s5-paired-reanalysis-ea66f4ec
No explanation has been written for this experiment yet. Add web/content/research/EX-20260823-d3n7m6-vs-d4s5-paired-reanalysis-ea66f4ec.mdx and it will appear here. The registered protocol is shown below.

The registered protocol

On the retained drop7-lifetime-cohort-v1 per-game records for D3 N7M6 (runs/RUN-A525-reveal/d3-n7-m6.json) and D4 s5 (runs/RUN-A51D-s7confirm/fresh-s5.json), both covering seeds 0xa51d1000-0xa51d103f, the paired score delta has a one-sided 95% bootstrap lower bound and a Student-t lower bound both above zero.

Candidate
D3 N7M6: factored-chance-fair-search, depth 3, discSamples 7, revealSamples 6, terminalUtility -1000000, maximumWork 51084852, 2000-move cap, corrected 17,000-point Hardcore scoring (finding-16 arm, RUN-20260821T035407Z-00483c6c)
runs/RUN-20260823T191900Z-b9f8f80d/c0/c0_compare.py
Comparator
D4 s5: parameterized-fair-search, depth 4, chanceSamples 5 (N5 M1), terminalUtility -1000000, maximumWork 3200000, 2000-move cap, corrected 17,000-point Hardcore scoring (finding-05 fresh-s5 arm, runs/RUN-A51D-s7confirm/fresh-s5.json; byte-identical games to runs/RUN-A52-LEAF/eval/s5-w000.json)
approaches/fair-expectimax/reference/fair-only-depth4.cpp

Primary metric

paired mean whole-game score delta, D3 N7M6 minus D4 s5, 64 games, seeds 0xa51d1000-0xa51d103f

Statistical unit: whole-game

Pass criteria, fixed in advance

  • Both arms are drop7-lifetime-cohort-v1 artifacts covering exactly the same 64 seeds 0xa51d1000-0xa51d103f in order, 0 censored games, 2000-move cap, terminalUtility -1000000, seedLease SEEDLEASE-A51D.
  • D3 N7M6 minus D4 s5 paired score delta: percentile bootstrap 95% lower bound > 0 AND Student-t 95% lower bound > 0.

On pass: Record valid + pass at public-development tier on already-read data; no tier above CHECK/diagnostic is claimed because the point estimate preceded the gate; any promotion requires a fresh-development cohort under a new protocol.

On fail: Record valid + fail; the theory is not-supported-as-tested at public-development tier; the +79,115 difference of means may not be quoted as a head start in the K2 program without the bound attached; open no cohort.

Data and reuse

Reanalysis only: no game is played, no seed is opened, no lease is consumed. Both arms were played under SEEDLEASE-A51D on the already-read shared development cohort 0xa51d1000-0xa51d103f (exploratory-development-diagnostic). The arm MEANS (376,442 and 297,327) and their difference (+79,115) were published in finding-16 and finding-05 and were known to the author before this gate was written, as was the finding-16 contrast D3 N7M6 minus D4 N7 M1 (-22,056). What was never computed or printed is the paired per-game D3 N7M6 minus D4 s5 delta and its bounds. This is therefore a bound computation on already-read data whose sign is already known; it is diagnostic tier only and its pass cannot be promoted above public-development on this cohort. The bootstrap seed, resample count and analysis code are pinned in expectedArtifacts before the wrapper is run.

seed leases: none

What happened

valid run · outcome: passsupported-as-testedpublic-developmentRS-20260823T194200Z-42b113db

C0 reanalysis (K2 program section 8 row 0): the paired per-game bound behind the +79,115 head start is now printed, and it clears zero. On the 64 shared seeds 0xa51d1000-0xa51d103f, depth 3 with seven disc strata and six reveal samples (fair leaf, finding-16) scores 376,442 against fair depth 4 at five strata's 297,327 (finding-05): paired +79,115, one-sided 95% percentile bootstrap lower bound +29,033 (upper +129,722; 20,000 resamples, seed 0xb0071eaf), Student-t lower bound +27,548, W-T-L 35-0-29, halves +88,138 / +70,092, Q25 +34,872 (227,224 vs 192,352), median 322,859 vs 260,415, median paired delta +32,926, moves +22.30 (LB +8.67), paired sd 247,113, detection floor 50,813, at 3.27x the logical work per move (4,244,020 vs 1,296,034). Both gate bounds are positive, so the experiment passes and the theory is supported-as-tested at public-development tier. READ THE CAVEATS: the means and their difference were published before this gate was written, so this is a bound attached to a known sign on already-read development data, not a discovery; it is not promotable above this tier on this cohort. The delta is heavy-tailed: the five largest paired gains (seeds 0xa51d1012 +747,416, 0xa51d1020 +684,129, 0xa51d1008 +643,799, 0xa51d1001 +643,344, 0xa51d1033 +558,163) carry 64.7% of the summed delta, the sixteen largest carry 134% (the remaining 48 games net negative), and the worst loss is -332,950 (0xa51d103c); the minimum leave-one-out mean is still +68,507, so no single game is load-bearing, but against a detection floor of 50,813 the mean sits 1.56 floors above zero and the bootstrap lower bound only 0.57 floors above it. The flow statistics move with the score: numbered clears per move 2.0447 vs 1.9489, cover reveals per move 1.1423 vs 1.0697 (paired clears +53.9 per game, LB +21.1; reveals +31.8, LB +12.6), mean occupancy 23.49 vs 24.29. SECONDARY, not gated: against fair depth 4 at seven strata (398,498) the same arm is -22,056 with bounds (-92,005, +45,490), t lower bound -92,330, W-T-L 30-0-34, halves -27,121 / -16,991, Q25 +14,360, moves -5.20, at 0.86x the work; this reproduces finding-16's -22,056 point estimate exactly and its bootstrap lower bound to within 2,138 (finding-16 printed -89,867 from its own resampler; this run's pinned seed gives -92,005), and remains a wash. So the head start the K2 fallback route (c) stands on is real against D4 s5 but the same arm does not separate from D4 s7, which is the stronger comparator and the one the program's end state must beat on the strength-cost frontier.

What it had to pass
  • Both arms are drop7-lifetime-cohort-v1 artifacts covering exactly the same 64 seeds 0xa51d1000-0xa51d103f in order, 0 censored games, 2000-move cap, terminalUtility -1000000, seedLease SEEDLEASE-A51D — observed: seed lists identical and in order for d3-n7-m6.json, fresh-s5.json, fresh-s7.json and s5-w000.json; censoredGames 0 in all; maximumMoves 2000; terminalUtility -1000000; seedLease SEEDLEASE-A51D; dataRole exploratory-development-diagnostic. Work bounds differ by design (51,084,852 vs 3,200,000 vs 16,000,000): they are the arms' own budgets, not a scoring setting. fresh-s5.json and s5-w000.json carry identical scores and moves on every game.
  • D3 N7M6 minus D4 s5: percentile bootstrap 95% lower bound > 0 AND Student-t 95% lower bound > 0 — observed: bootstrap LB +29,033, Student-t LB +27,548 (t quantile 1.6694, df 63); mean +79,115; W-T-L 35-0-29; halves +88,138 / +70,092
Recorded metrics
cohort
0xa51d1000-0xa51d103f, 64 games, 2,000-move cap, corrected 17,000-point Hardcore scoring, terminalUtility -1000000, SEEDLEASE-A51D, exploratory-development-diagnostic, already read
d3n7m6
artifact
runs/RUN-A525-reveal/d3-n7-m6.json
sha256
36aa1d0f0ec29b0ec70bd540e7d06758e37147a521f6891c0fde930943aa83fc
meanScore
376442.1406
median
322,859
q25
227224.5000
min
103,015
max
1,002,557
sd
223365.7773
meanMoves
109.4531
workPerMove
4244020.0805
maximumWork
51,084,852
censoredGames
0
gamesAtOrAboveMillion
1
d4s5
artifact
runs/RUN-A51D-s7confirm/fresh-s5.json
sha256
d6c1dbd403dbf5b477f1b0905b4aa440e443c2dfba8eef29085b7f5bee024d9a
meanScore
297327.3906
median
260,415
q25
192,352
min
86,935
max
836,427
sd
150549.5574
meanMoves
87.1563
workPerMove
1296033.6877
maximumWork
3,200,000
censoredGames
0
gamesAtOrAboveMillion
0
d4s7
artifact
runs/RUN-A51D-s7confirm/fresh-s7.json
sha256
0c95442620a78ef5a928a4b8cdab9f57b3bdb8a3d7df5fcd1a6d818785bd4209
meanScore
398498.2344
median
344630.5000
q25
212,864
meanMoves
114.6563
workPerMove
4956614.2652
maximumWork
16,000,000
censoredGames
0
pairedD3n7m6MinusD4s5
n
64
meanScoreDelta
79114.7500
pairedSd
247112.6017
bootstrapLower95
29032.5102
bootstrapUpper95
129721.6797
studentTLower95
27548.4592
studentTQuantile
1.6694
detectionFloor
50812.5287
winTieLoss
35-0-29
firstHalfMeanDelta
88137.8438
secondHalfMeanDelta
70091.6563
q25Delta
34872.5000
medianPairedDelta
32925.5000
minLeaveOneOutMeanDelta
68506.7900
top5Share
0.6470
top16Share
1.3400
meanMovesDelta
22.2969
movesBootstrapLower95
8.6719
movesStudentTLower95
8.2664
numberedClearedDelta
53.9375
numberedClearedLower95
21.1406
coversRevealedDelta
31.7969
coversRevealedLower95
12.6094
workRatio
3.2746
bootstrapSeedHex
0xb0071eaf
resamples
20,000
pairedD3n7m6MinusD4s7
n
64
meanScoreDelta
-22056.0938
pairedSd
336763.8340
bootstrapLower95
-92004.8930
bootstrapUpper95
45490.1641
studentTLower95
-92330.3803
detectionFloor
69247.0634
winTieLoss
30-0-34
firstHalfMeanDelta
-27121.1250
secondHalfMeanDelta
-16991.0625
q25Delta
14360.5000
meanMovesDelta
-5.2031
movesBootstrapLower95
-24.0313
workRatio
0.8562
finding16PointEstimate
-22,056
finding16BootstrapLower95
-89,867
Limitations
  • Bound computation on already-read development data whose means and sign were published (finding-16, finding-05) before the gate was written; diagnostic, not promotable above public-development on this cohort. A fresh-development replication under a new protocol is required before the +79k is used as anything but a planning prior.
  • The two arms were played in different runs (RUN-20260821T035407Z-00483c6c for D3 N7M6; the finding-05 fresh-s5 arm) by different binaries (factored-chance-fair-search vs parameterized-fair-search); the seed lists, move cap, terminal utility and scoring are identical but the pairing is across builds, not within one runner invocation. Engine parity between these families was established separately (finding-09/16 checks; finding-15 engine control), not re-run here.
  • Heavy tail: five games carry 64.7% of the summed delta; the bootstrap lower bound sits 0.57 detection floors above zero. 64 games cannot resolve anything below about 50,800.
  • The arm does not separate from D4 s7 (-22,056, bounds -92,005 to +45,490), so 'beats D4 s5' does not transfer to 'beats the current best fair comparator'.
  • The wrapper pinned at freeze failed to import (module-name collision) and was amended before any output existed; the amendment changed no statistic. The first pinned wrapper hash is retained in expectedArtifacts and the amended hash in amendments.