Linear leaf refit to the depth-4 root value: held-out fit, depth-5 agreement proxy, and gameplay pilot
Stage 1 (no new seeds): on the EX-4 panel, compute the eighteen fair-leaf terms, the frozen leaf value and the best immediate six-feature drop value of every root; fit ridge regressions of the depth-4 root value (max over legal columns of v4) on (i) the eighteen terms plus intercept and (ii) the same plus the drop term, using fit games 0-31 restricted to roots with root value above -50,000, lambda from {0.001, 0.01, 0.1, 1, 10} by 4-fold game-grouped CV; report held-out R^2 (games 32-47, same filter) against the frozen-leaf affine baseline.
On this page
- Created
- Updated
No explanation has been written for this record yet.
Technical recordThe registered protocol
- Hypothesis
- Stage 1 (no new seeds): on the EX-4 panel, compute the eighteen fair-leaf terms, the frozen leaf value and the best immediate six-feature drop value of every root; fit ridge regressions of the depth-4 root value (max over legal columns of v4) on (i) the eighteen terms plus intercept and (ii) the same plus the drop term, using fit games 0-31 restricted to roots with root value above -50,000, lambda from {0.001, 0.01, 0.1, 1, 10} by 4-fold game-grouped CV; report held-out R^2 (games 32-47, same filter) against the frozen-leaf affine baseline. Stage 2 (no new seeds): on EX-4's 128-root depth-5 sub-panel, decide every root with full-width d4s7 using the refit-18 leaf and report its regret under exact d5s7 (frozen leaf) next to frozen d4s7's, as a disclosed biased proxy. Stage 3 (pilot cohort, only if P-L1 passes): 256 games of full-width d4s7 with the refit-18 leaf on 0xa5277000-0xa52770ff, paired against EX-5's full-width frozen-leaf arm; the refit-19 leaf is a second arm only if its held-out R^2 exceeds refit-18's by at least 0.01.
- Arms
Arm Name Entry point Manifest Candidate full-width fair d4s7 with a refit LinearLeaf (drop7-oneply-q leaf.rs; evaluate --arm leaf:4:7:FILE) approaches/value-policy-learning/oneply-q-prune/rust/src/bin/leaf_terms.rs– Comparator full-width fair d4s7 with the frozen fair leaf: EX-5's 256-game arm on the same seeds, and exact d5s7 values on the sub-panel approaches/value-policy-learning/oneply-q-prune/rust/src/bin/evaluate.rs– - Classification
- algorithmic
- Information boundary
- public-policy
- Benchmark tier
- PILOT
- Lifecycle
- preregistered
- Theories tested
- Primary metric
- paired per-game corrected-score delta of the refit-18 arm over the frozen-leaf arm on 256 games, one-sided 95% percentile-bootstrap lower and upper bounds (10,000 resamples, RNG seed 0x6b660001), detection floor stated
- Secondary metrics
- held-out R^2 of the depth-4 root value for the frozen-leaf affine baseline, refit-18 and refit-19; CV curves and selected lambdas; the refit weight vectors next to the frozen constants
- sub-panel: regret under exact d5s7 of d4s7-refit-18 versus d4s7-frozen, top-1 agreement with d5, and the d4-refit-vs-d4-frozen disagreement rate
- paired lifetime delta, work per move, wall per game, score quartiles, mechanics counts
- Statistical unit
- whole-game
- Uncertainty method
- paired per-game deltas with one-sided 95% percentile bootstrap bounds; root-level proxies with game-clustered bootstrap
- Data role
- previously-evaluated-development
- Seed leases
- none recorded
- Whole-origin split
- yes
- Reuse disclosure
- Zero new seeds. Stages 1 and 2 read only the EX-4 panel (training-role seeds 0xa5200000-0xa520002f, whole-game split 0-31 fit / 32-47 held out). Stage 3 reads the previously evaluated pilot cohort 0xa5277000-0xa52770ff once more, as EX-5 does, and pairs against EX-5's rows. The coordinator's confirmation requested in the 2026-09-02 log covers this record too.
- Pass criteria
- P-L1 (theory a): held-out R^2 of refit-18 exceeds the frozen-leaf affine baseline's by at least 0.05; otherwise no gameplay arm is run and the theory's clause (a) fails.
- P-L2 (theory b): the refit-18 arm's paired score delta over the frozen-leaf arm has LB95 above zero; fail if UB95 is below zero; otherwise inconclusive with the floor stated.
- P-L3 (variant): refit-19's held-out R^2 exceeds refit-18's by at least 0.01 (expected to fail; a pass adds the refit-19 arm).
- Mechanics: LinearLeaf at the frozen constants is bit-identical to FairLeaf (crate test); zero illegal and incomplete decisions.
- On pass
- Record a pilot-tier assessment; a P-L2 pass warrants a SCREEN on fresh seeds under a coordinator lease and a second TreeStrap iteration (re-panel with the refit leaf).
- On fail
- Record valid + fail for the exact refit; the negative result closes one-step linear TreeStrap on these eighteen terms at pilot tier.
- Gate fixed before controlled data
- yes
- Resources
Wall seconds 14400 CPU threads 12 Max host bytes 8589934592 Max GPU bytes – GPU devices – - Stop conditions
- Stop at 14,400 s wall over all stages.
- No seed outside 0xa5200000-0xa520002f (panel, already read) and 0xa5277000-0xa52770ff (pilot cohort) may be read under this record.
- Any non-finite leaf value, illegal decision or runner failure marks the affected stage invalid.
- Expected artifacts
runs/<run-id>/oneply-q/root-terms.ndjsonruns/<run-id>/oneply-q/fit-leaf.json, weights-leaf18.txt, weights-leaf19.txtruns/<run-id>/oneply-q/leaf-d5.ndjson and leaf-d5-summary.jsonruns/<run-id>/oneply-q/evaluate-leaf.json and leaf-pilot-summary.json
- Amendments
- none recorded
Technical recordResults recorded against this protocol
Stage 1: on the EX-4 panel's fit games (2,913 roots with depth-4 root value above -50,000; 425 dropped), ridge regression of the depth-4 root value on the eighteen fair-leaf terms plus intercept (lambda 0.001 by game-grouped CV, the grid edge) reaches held-out R^2 0.852 on 1,877 held-out roots, against 0.729 for an affine rescaling of the frozen leaf (gain +0.123; P-L1 passes) and 0.852 with the best-immediate-six-feature-drop term added (gain -0.0001; P-L3 fails as predicted: the six drop features carry no leaf-relevant information the fair terms lack). The refit reweights the terms wildly relative to the frozen constants (open columns -4,232 against +180, solid exposure +3,273 against +40, intercept +45,744), which correlated terms permit. Stage 2, the disclosed biased proxy: full-width d4s7 with the refit leaf agrees with exact frozen-leaf d5s7 on 68.0% of the 128 sub-panel roots against 79.7% for the frozen leaf and loses 9,350 points per decision under those values against 146 (paired change -9,204, bound -21,308); it changes 28 of 128 decisions. Stage 3: in play on the 256 pilot games, full-width fair d4s7 with the refit leaf averages 321,045 points and 93.74 moves against 404,497 and 116.41 for the frozen leaf on the same seeds: paired delta -83,451 points (one-sided 95% bounds -118,022 and -49,021; floor 34,345) and -22.66 moves (bounds -31.98, -13.32), 104 wins to 152; upper bound below zero, so P-L2 FAILS with a measurable loss of about a fifth of the score. The leaf that predicts the depth-4 value best is a worse leaf: it explains more variance among live roots by absorbing four plies of expected score, and inside the search that reweighting mis-prices the positions the search actually has to choose between, above all near death, which the fit deliberately excluded. Unlike three earlier proxies, this one pointed the right way. Valid negative result at pilot tier; one-step linear TreeStrap on these eighteen terms is closed as tested.
- ✓P-L1 (theory a): held-out R^2 of refit-18 exceeds the frozen-leaf affine baseline by >= 0.05 — observed: 0.852 vs 0.729, gain +0.123 (lambda 0.001, the grid edge)
- ✕P-L2 (theory b): refit-18 arm's paired score delta over the frozen-leaf arm has LB95 > 0; fail if UB95 < 0 — observed: -83,451 [LB -118,022, UB -49,021], floor 34,345; moves -22.66 [LB -31.98, UB -13.32]; W/T/L 104/0/152
- ✕P-L3 (variant): refit-19's held-out R^2 exceeds refit-18's by >= 0.01 — observed: gain -0.0001 (kf_beta 326 on a term the eighteen already explain); no refit-19 arm run
- ✓Mechanics: LinearLeaf at the frozen constants bit-identical to FairLeaf; zero illegal and incomplete decisions — observed: crate test frozen_linear_leaf_matches_fair_leaf_bits (56 boards); 0 illegal, 0 incomplete, 0 censored
Technical recordRecorded metrics
- floor
- -50,000
- fitRoots
- 2,913
- fitRootsDropped
- 425
- heldOutRoots
- 1,877
- heldOutRootsDropped
- 183
- baselineAffine
- slope
- 0.9202
- intercept
- 12988.0593
- lambda
- 0.0010
- cv
- 0.001
- 73401582.5580
- 0.01
- 73410223.4741
- 0.1
- 74906572.4486
- 1.0
- 121426385.8128
- 10.0
- 233085590.2247
- r2HeldOut
- 0.7293
- refit18
- weights
- open_columns
- -4232.4797
- height_load
- 6.3982
- solid_cells
- -1201.5322
- cracked_cells
- 284.9428
- numbered_cells
- -338.5627
- high_low_numbers
- -149.0987
- direct_potential
- 863.4063
- latent_chain_potential
- 564.1850
- cracked_exposure
- 1538.3437
- solid_exposure
- 3272.8873
- adjacent_ones
- -2840.1211
- triple_twos
- -4222.7704
- dead_low_numbers
- -884.3974
- covered_height_risk
- -158.0583
- low_number_height_risk
- 17.6718
- danger_height_squared
- -333.3924
- rise_pressure
- 4.8411
- next_disc_vertical_options
- 58.6726
- bias
- 45744.3099
- lambda
- 0.0010
- cv
- 0.001
- 37957398.9705
- 0.01
- 37958704.1872
- 0.1
- 38523932.0167
- 1.0
- 54358654.2899
- 10.0
- 153452998.5746
- r2HeldOut
- 0.8521
- r2Fit
- 0.8654
- refit19
- kfBeta
- 325.9629
- bias
- 45637.6966
- lambda
- 0.0010
- r2HeldOut
- 0.8520
- gainOverBaseline
- 0.1228
- gain19Over18
- -0.0001
- roots
- 128
- truth
- exact frozen-leaf d5s7
- configs
- exact:4:7
- roots
- 128
- games
- 16
- top1Agreement
- 0.7969
- meanNormalisedRegret
- 0.0285
- meanRawRegret
- 145.6493
- maxNormalisedRegret
- 0.6959
- meanWork
- 5105613.2969
- workRatioOfMeans
- 1
- meanWorkRatio
- 1
- priorWorkShare
- 0
- meanPrunedNodes
- 0
- meanWallMs
- 4962.1700
- monotoneViolations
- 218
- top1AgreementLb95
- 0.7524
- meanNormalisedRegretUb95
- 0.0398
- weights-leaf18.txt
- roots
- 128
- games
- 16
- top1Agreement
- 0.6797
- meanNormalisedRegret
- 0.0846
- meanRawRegret
- 9349.9620
- maxNormalisedRegret
- 1
- meanWork
- 5105613.2969
- workRatioOfMeans
- 1
- meanWorkRatio
- 1
- priorWorkShare
- 0
- meanPrunedNodes
- 0
- meanWallMs
- 4950.8656
- monotoneViolations
- 861
- top1AgreementLb95
- 0.6129
- meanNormalisedRegretUb95
- 0.1321
- pairedVsexact:4:7
- meanNormalisedRegretReduction
- -0.0560
- normalisedLb95
- -0.0977
- meanRawRegretReduction
- -9204.3127
- rawLb95
- -21308.0222
- roots
- 128
- weights-leaf19.txt
- roots
- 128
- games
- 16
- top1Agreement
- 0.6875
- meanNormalisedRegret
- 0.0866
- meanRawRegret
- 9352.2231
- maxNormalisedRegret
- 1
- meanWork
- 5105613.2969
- workRatioOfMeans
- 1
- meanWorkRatio
- 1
- priorWorkShare
- 0
- meanPrunedNodes
- 0
- meanWallMs
- 5901.2160
- monotoneViolations
- 861
- top1AgreementLb95
- 0.6179
- meanNormalisedRegretUb95
- 0.1345
- pairedVsexact:4:7
- meanNormalisedRegretReduction
- -0.0581
- normalisedLb95
- -0.0986
- meanRawRegretReduction
- -9206.5737
- rawLb95
- -21291.2906
- roots
- 128
- refitVsFrozenD4SameAction
- 100/128
- cohort
- seedsStartHex
- 0xa5277000
- games
- 256
- moveCap
- 2,000
- arms
- fair-d4s7
- games
- 256
- score
- mean
- 404496.8555
- sd
- 282958.1647
- median
- 315157.5000
- q25
- 211976.2500
- min
- 103,308
- max
- 1,746,375
- moves
- mean
- 116.4063
- median
- 90
- q25
- 65
- min
- 35
- max
- 475
- censoredGames
- 0
- numberedClearsPerMove
- 2.0715
- coveredRevealsPerMove
- 1.1652
- meanChainDepth
- 2.1475
- maximumChainDepth
- 13
- illegalDecisions
- 0
- incompleteDecisions
- 0
- logicalWorkPerMove
- 5221078.2738
- wallSecondsPerGame
- 464.7704
- leaf-d4s7-weights-leaf18
- games
- 256
- score
- mean
- 321045.4141
- sd
- 203027.9780
- median
- 264787.5000
- q25
- 173,833
- min
- 85,331
- max
- 1,275,013
- moves
- mean
- 93.7422
- median
- 80
- q25
- 55
- min
- 30
- max
- 355
- censoredGames
- 0
- numberedClearsPerMove
- 2.0144
- coveredRevealsPerMove
- 1.1403
- meanChainDepth
- 2.1226
- maximumChainDepth
- 12
- illegalDecisions
- 0
- incompleteDecisions
- 0
- logicalWorkPerMove
- 5236533.4317
- wallSecondsPerGame
- 358.2558
- pairedVsFairD4s7
- leaf-d4s7-weights-leaf18
- score
- n
- 256
- meanDelta
- -83451.4414
- sdDelta
- 334053.5228
- lb95
- -118021.6523
- ub95
- -49021.2969
- detectionFloor
- 34344.8778
- wins
- 104
- ties
- 0
- losses
- 152
- moves
- n
- 256
- meanDelta
- -22.6641
- sdDelta
- 90.7159
- lb95
- -31.9805
- ub95
- -13.3164
- detectionFloor
- 9.3267
- wins
- 98
- ties
- 9
- losses
- 149
- workPerMoveRatio
- 1.0030
- meanScore
- 321045.4141
- referenceMeanScore
- 404496.8555
- wallClock
- 2026-09-02T20:25:12Z to 2026-09-03T00:42:41Z on 6 threads, 15,449 s, shared with three EX-5 arms
- Pilot tier; the comparator rows are EX-5's full-width arm on the same seeds (RS RS-20260903T013022Z-d32ee053), read once more here.
- The refit excluded death-dominated roots by design (root value above -50,000) and fitted the smallest lambda of the preregistered grid; a refit that keeps or reweights death-dominated roots, or a shrinkage toward the frozen constants, was not tested and is what would reopen the direction.
- The stage-2 proxy scores the refit leaf's decisions with the frozen leaf's depth-5 values and is biased toward the frozen leaf; it is reported because it predicted the gameplay sign correctly, not as evidence on its own.
- The leaf arm ran concurrently with three other arms; its 15,449 s wall plus the earlier stages exceeded the preregistered 14,400 s stop by about 1,300 s, which was not enforced; every row is deterministic and no number is affected. Wall times are contended; work per move (5.24M vs 5.22M, ratio 1.003) is the cost comparison.
- Panel-derived R^2 values are on training-role roots of fair-d4s7 play and describe the depth-4 root value, not score.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/EX-20260902-treestrap-leaf-refit-panel-and-pilot-4024f360.mdx; it renders above this record on the next request. The registered protocol itself is in the technical record above.
Record file: research/experiments/EX-20260902-treestrap-leaf-refit-panel-and-pilot-4024f360.json, validated against research/schemas/experiment-v1.schema.json. Protocol hash: 79d85f7ea7273d92382d12b0b8e33fd9fe2a1cd2efce4cf5d43257d8c3503134.
Registered by Claude Code / claude-fable-5-1 (claude-q-learning).