On this page
Dates
Created
Updated
Record idEX-20260902-oneply-q-fit-and-pruned-search-panel-46c75cdf

No explanation has been written for this record yet.

Technical recordThe registered protocolEX-20260902-oneply-q-fit-and-pruned-search-panel-46c75cdf
Hypothesis
Build a sibling-complete panel of every root state visited by fair d4s7 in 48 complete games on training-role seeds, with exact column values at depths one to four, per-sibling one-ply features under the depth-4 chance scenarios, and remaining lifetime. Fit a ridge-regression linear action value on the first 32 games and evaluate every ranker (centre order, the CEM six-feature Q of RS-20260902T084356Z-784ebf14, the exact d1 value, the exact d3 value, and the fitted Q) on the 16 held-out games for top-1 agreement, recall@k and normalised regret against exact d4s7. Then run the pruned search (interior max nodes with two or more plies remaining expand only the prior's top-w siblings; the root and the one-ply-remaining layer stay full width; chance scenarios, leaf and terminal utility unchanged) at depth four on every held-out root and at depth five on a 128-root held-out sub-panel with exact d5s7 values, recording chosen action, root values and logical work. Tests theory clauses (a)-(d).
Arms
ArmNameEntry pointManifest
Candidatepruned fair expectimax (drop7-oneply-q crate: prune.rs) with priors centre, kf-cem, d1, lqapproaches/value-policy-learning/oneply-q-prune/rust/src/bin/prune_eval.rs
Comparatorfull-width fair d4s7 (drop7-rs Searcher, FairLeaf, 65,536-entry depth table from depth 1, terminal -1,000,000, policy seed 0xd7075eed) and exact full-width d5s7 on the sub-panelapproaches/fair-expectimax/rust-engine/src/search.rs
Classification
algorithmic
Information boundary
public-policy
Benchmark tier
PILOT
Lifecycle
preregistered
Primary metric
held-out top-1 agreement and mean normalised regret ((V*(best)-V*(chosen))/(V*(best)-V*(worst)), 0 when all legal values are equal) of each pruned-d4s7 configuration against full-width d4s7, with the ratio of logical work (move calls plus leaf calls, prior calls included) per root
Secondary metrics
  • fit quality of the linear Q on held-out roots: R^2 on raw values, within-root Spearman, top-1, recall@2, recall@3, recall@4, normalised regret; the same ranking metrics for centre order, kf-cem, exact d1, exact d2 and exact d3
  • the fitted weight vector, the CV-selected ridge lambda and the CV curve
  • on the 128-root sub-panel with exact d5s7 values: mean normalised and raw regret under d5 of exact d4s7 and of every pruned-d5s7 configuration, their top-1 agreement with d5, work per root relative to full-width d4s7, and the d5-vs-d4 disagreement rate
  • monotonicity check: every pruned root value is at most the full-width root value on the same root (bit-level lower bound)
  • panel descriptives: roots per game, legal columns per root, remaining-lifetime distribution
  • wall seconds per stage
Statistical unit
root
Uncertainty method
root-level proportions with game-clustered (16 held-out games) standard errors; paired per-root regret differences with a one-sided 95% percentile bootstrap over games (10,000 resamples, RNG seed 0x6b660002); detection floors reported
Data role
previously-evaluated-development
Seed leases
none recorded
Whole-origin split
yes
Reuse disclosure
Zero new seeds. The panel games read 0xa5200000-0xa520002f, the first 48 seeds of the already-opened TRAINING-role block SEEDLEASE-A52 d2 (docs/exploratory/lease-map.md), a block re-read for training by RUN-20260902T081659Z-d2aa6375 and RUN-20260902T083357Z-196a7b87; the role stays training and nothing read here can become evaluation evidence. The whole-origin split is by game: cohort ordinals 0-31 fit, 32-47 held out; no fit, lambda choice or configuration choice may read a held-out game. CHECK probes may read 0xa5278000-0xa52784ff. No evaluation-role seed is read under this record; a gameplay pilot, if warranted, is a separate preregistered experiment. The coordinator's confirmation of the training-block reuse requested in the 2026-09-02 log covers this record too.
Pass criteria
  1. P1 (theory a): the fitted Q's held-out top-1 agreement with exact d4s7 exceeds the exact d1 value's on the same roots.
  2. P2 (theory b): the fitted Q's held-out recall@3 is at least 0.95 (fail below 0.90).
  3. P3 (theory c): a preregistered pruned-d4s7 configuration with work ratio at most 0.35 reaches top-1 agreement at least 0.95 and mean normalised regret at most 0.02 on held-out roots (fail if none reaches 0.90 and 0.05).
  4. P4 (theory d): a preregistered pruned-d5s7 configuration with work at most 2.0 times full-width d4s7 has mean regret under exact d5s7 below exact d4s7's, with a game-bootstrap one-sided 95% lower bound on the paired difference above zero (inconclusive if the point estimate is favourable but the bound is not).
  5. Mechanics (CHECK, before any panel game): the pruned searcher with all widths seven reproduces drop7-rs Searcher per-column values, chosen actions and work bit-for-bit on probe states; determinism across worker counts; every pruned decision is a legal column; the leaf-term accessor reconstructs the frozen fair leaf bit-for-bit.
  6. Preregistered configuration set. Depth 4, widths (three plies remaining, two plies remaining) in {(2,2),(3,3),(4,4),(3,2)} times priors {centre, kf-cem, d1, lq}. Depth 5, widths (four, three, two plies remaining) in {(2,2,2),(3,3,2),(3,3,3)} times priors {d1, lq}. The fitted Q used as the lq prior is the per-root-centred ridge fit at the CV-selected lambda from the grid {0.001, 0.01, 0.1, 1, 10} on standardised features (4-fold, grouped by game, within the 32 fit games); the raw-value fit is reported but not deployed.
On pass
If P3 passes, preregister a separate PILOT gameplay experiment on the previously evaluated cohort 0xa5277000-0xa52770ff (256 paired games; arms full-width fair d4s7, the lowest-regret d4 configuration at work ratio at most 0.35, and, if P4 passes, the lowest-regret d5 configuration at work at most 2.0 times d4s7). Selection uses only held-out panel metrics recorded here.
On fail
Record valid + fail for the failed clause(s) with the exact configuration set; a failed P2 closes the pruning clauses without a gameplay pilot; a failed P3 records that one-ply pruning changes d4 decisions too often to be an engineering speed-up and must be treated as a new algorithmic candidate.
Gate fixed before controlled data
yes
Resources
Wall seconds7200
CPU threads12
Max host bytes8589934592
Max GPU bytes
GPU devices
Stop conditions
  1. Stop at 7,200 s wall for the panel and pruning stages together.
  2. No seed outside 0xa5200000-0xa520002f (panel games) and 0xa5278000-0xa52784ff (CHECK probes) may be read under this record.
  3. Any illegal decision, incomplete search, or non-finite value aborts the affected stage as invalid.
  4. The 128-root sub-panel is every floor(N/128)-th held-out root in cohort order (first 128 such roots); its size is not enlarged after reading.
Expected artifacts
Amendments
none recorded
Technical recordResults recorded against this protocol1 record
valid runoutcome: failmixedtier: pilotRS-20260902T202705Z-75d87947

A sibling-complete panel of 5,398 roots from 48 fair-d4s7 games on already-read training seeds (mean 390,143 points, 112.5 moves, no game censored) was split 32 games for fitting and 16 held out (2,060 roots). The linear action value over 32 one-ply features, fitted by ridge regression to the exact depth-4 sibling values (lambda 0.001 at the grid edge for both the raw and the per-root-centred fit; held-out R^2 0.725 raw, 0.539 within root), is a WORSE ranker of depth-4's choice than the exact one-ply value it can represent: top-1 0.308 against 0.524 (paired difference -0.216, one-sided 95% bound -0.233 over games) and recall@3 0.604 against 0.809, so clauses (a) and (b) fail. A post-hoc check shows why: the features rebuild the exact one-ply value to 6e-11, and least squares halves the held-out error (7.0e9 against 1.4e10 for an affine map of the one-ply value) by predicting deaths within four plies, which cost a million points each, while ranking live siblings worse; refitting on per-root-centred targets clipped at -50,000 recovers exactly the one-ply value's ranking (top-1 0.535, recall@3 0.809) and nothing more. Depth-three exact values rank depth-4's choice into the top three on 95.0% of roots, depth-two on 90.0%, the six-feature CEM policy on 70.1%, centre order on 40.0%. Pruning the depth-4 search at interior max nodes with the exact one-ply value as prior at width three (both interior layers) reproduces the full-width decision on 92.3% of held-out roots (lower bound 91.8%) with mean normalised regret 0.0087 (upper bound 0.0098) and mean raw regret 74 points at 26.4% of the full-width logical work (prior calls included, 5.9% of it) and 3.8x faster wall time; width four gives 93.8% / 0.0059 at 40.9% work, width two 89.3% / 0.0145 at 13.8%. The six-feature prior, which costs no engine calls, reaches 88.4% / 0.0147 at width four (34.6% work); the fitted Q prunes worst (78.8% / 0.037 at width three); centre order is useless (69.1%). Every pruned root value was at or below the full-width value on every root (0 violations of the lower-bound property), and the width-seven searcher reproduced drop7-rs values, actions and work bit for bit. Clause (c) is inconclusive by the preregistered rule: the best configuration at or under 35% work passes the regret threshold (0.0087 <= 0.02) but its agreement of 92.3% sits between the 0.90 fail line and the 0.95 pass line. On the 128-root sub-panel with exact depth-5 values (mean 1.9e8 work per root, 37x depth 4), full-width depth 4 agrees with depth 5 on 79.7% of roots with normalised regret 0.0285 and raw regret 146 points per decision; pruned depth 5 with the one-ply prior at width two on all three interior layers (1.68x the depth-4 work) agrees on 80.5% with normalised regret 0.0174 (paired reduction +0.011, bound -0.0004) but HIGHER raw regret, 189 points (paired change -43, bound -207), so clause (d) is inconclusive at best and unfavourable in points: pruning three interior layers loses the high-stakes continuations that the extra ply was supposed to find. Wider depth-5 pruning (3,3,3 at 4.5x work) reaches only 82.8%. The theory's learned-prior claim is not supported; its mechanism claim is supported for the search's own one-ply value at depth 4 and not at depth 5. Pilot tier, training-role and held-out panel data, no new seeds.

What it had to pass
  • P1 (theory a): fitted Q's held-out top-1 exceeds the exact one-ply value's — observed: lq-centred 0.308 (bounds 0.292-0.327) vs d1@4 0.524; paired lq-d1 top-1 difference -0.216, one-sided 95% bound -0.233 over 16 games; lq-raw 0.311
  • P2 (theory b): fitted Q's held-out recall@3 >= 0.95 (fail < 0.90) — observed: 0.604 (bounds 0.589-0.622); d1@4 0.809; v3 0.950
  • P3 (theory c): a pruned-d4s7 configuration at <= 0.35 work with agreement >= 0.95 and regret <= 0.02 (fail if none reaches 0.90 / 0.05) — observed: inconclusive: pruned:4:7:3,3:d1 at 0.264 work has agreement 0.923 [LB 0.918] (between 0.90 and 0.95) and regret 0.0087 [UB 0.0098] (passes); pruned:4:7:3,2:d1 at 0.192 work 0.919 / 0.0102; pruned:4:7:4,4:kf at 0.346 work 0.884 / 0.0147
  • P4 (theory d): a pruned-d5s7 configuration at <= 2.0x d4 work with regret under exact d5 below d4's, paired LB > 0 — observed: inconclusive and unfavourable in points: pruned:5:7:2,2,2:d1 at 1.678x work has normalised regret 0.0174 vs d4's 0.0285 (paired reduction +0.0111, LB -0.0004) but raw regret 189 vs 146 points per decision (paired change -43, LB -207); agreement with exact d5 0.805 vs d4's 0.797; pruned:5:7:2,2,2:lq at 1.495x is worse on both (0.0486 / 217)
  • Mechanics (CHECK): width-seven searcher bit-identical to drop7-rs (values, actions, work); one-ply d1 bit-identical to the depth-1 search; determinism across worker counts; legality; leaf-term accessor bit-exact — observed: crate tests full_width_matches_drop7_rs_bit_for_bit (d3s5, d3s7, d4s5 on 12 probe states), oneply_d1_value_matches_depth_one_search_bit_for_bit, pruned_values_are_lower_bounds_and_legal, choose_action_replays_and_stays_legal, panel_game_is_deterministic_and_sibling_complete; engine test terms_dot_frozen_weights_match_fair_leaf_bits (100 boards); 0 monotonicity violations over 2,060 x 16 pruned decisions; 0 illegal decisions
Technical recordRecorded metricsRS-20260902T202705Z-75d87947
panel
games
48
roots
5,398
fitGames
32
fitRoots
3,338
fitRows
22,746
heldOutGames
16
heldOutRoots
2,060
heldOutRows
14,184
meanScore
390142.7083
meanMoves
112.4583
censoredGames
0
maxScore
1,170,447
panelWallSeconds
1433.6413
datasetId
DS-20260902-fair-d4s7-sibling-panel-e4e06b7c
fit
lambdaGrid
  1. 0.0010
  2. 0.0100
  3. 0.1000
  4. 1
  5. 10
lambdaRaw
0.0010
lambdaCentred
0.0010
cvMse
raw
0.001
9423855972.5112
0.01
9450988319.4366
0.1
9807155675.8847
1.0
12311060654.1782
10.0
23314691498.9662
centred
0.001
2799392427.0973
0.01
2842132115.3311
0.1
3109252073.8982
1.0
4216113803.9430
10.0
5401136647.9315
r2HeldOutRaw
0.7248
r2HeldOutCentredWithinRoot
0.5388
bootstrapOverGames
resamples
10,000
seed
1,801,846,786
lqMinusD1Top1
lb95
-0.2333
median
-0.2159
lqMinusD1Recall3
lb95
-0.2186
median
-0.2055
lqTop1
lb95
0.2923
ub95
0.3271
lqRecall3
lb95
0.5887
ub95
0.6221
weightsCentred
leaf_open_columns
137837.0504
leaf_height_load
378.1549
leaf_solid_cells
-6488.3764
leaf_cracked_cells
-1347.9892
leaf_numbered_cells
-2718.0940
leaf_high_low_numbers
-14309.6997
leaf_direct_potential
7438.2469
leaf_latent_chain_potential
3999.3394
leaf_cracked_exposure
-3834.7383
leaf_solid_exposure
-1230.2104
leaf_adjacent_ones
-8994.5464
leaf_triple_twos
-17894.6672
leaf_dead_low_numbers
-1394.8837
leaf_covered_height_risk
138.1181
leaf_low_number_height_risk
2044.6067
leaf_danger_height_squared
-11905.0013
leaf_rise_pressure
-250.7718
leaf_next_disc_vertical_options
7194.2342
delta_mean
-1.7023
terminal_frac
-271985.9091
kf_min_eq_elem_True
-715.6169
kf_row_dets
-995.1247
kf_col_dets
-87.1152
kf_max_eq_elem
-893.1574
kf_1_dets
1819.6338
kf_elem_det
818.3358
rise_1
0.0000
rise_2
0
rise_3
0.0000
rise_4
0
rise_5
0.0000
bias
0
weightsRaw
leaf_open_columns
114553.5553
leaf_height_load
508.1005
leaf_solid_cells
2795.2408
leaf_cracked_cells
4991.3561
leaf_numbered_cells
-5063.6683
leaf_high_low_numbers
45.2395
leaf_direct_potential
8265.7972
leaf_latent_chain_potential
5428.1432
leaf_cracked_exposure
-1649.6328
leaf_solid_exposure
1930.5985
leaf_adjacent_ones
-6884.2806
leaf_triple_twos
-21627.6831
leaf_dead_low_numbers
1928.6702
leaf_covered_height_risk
-608.2756
leaf_low_number_height_risk
355.8755
leaf_danger_height_squared
-11326.8416
leaf_rise_pressure
-402.6838
leaf_next_disc_vertical_options
16664.4817
delta_mean
-0.7481
terminal_frac
-165847.2588
kf_min_eq_elem_True
3741.9930
kf_row_dets
-965.7629
kf_col_dets
680.9917
kf_max_eq_elem
-848.5231
kf_1_dets
-799.9510
kf_elem_det
1564.8586
rise_1
3334.7789
rise_2
68092.4791
rise_3
-10514.0847
rise_4
-36234.0715
rise_5
-24686.4507
bias
-837811.4389
rankersHeldOut
center
top1
0.1294
recall2
0.2715
recall3
0.4000
recall4
0.5752
meanNormalisedRegret
0.5284
meanSpearman
0.0017
kf-cem
top1
0.3922
recall2
0.5883
recall3
0.7012
recall4
0.7800
meanNormalisedRegret
0.2939
meanSpearman
0.3281
d1@4
top1
0.5241
recall2
0.7085
recall3
0.8092
recall4
0.8749
meanNormalisedRegret
0.1977
meanSpearman
0.5305
v1
top1
0.5095
recall2
0.7022
recall3
0.8063
recall4
0.8725
meanNormalisedRegret
0.2033
meanSpearman
0.5246
v2
top1
0.6642
recall2
0.8326
recall3
0.8998
recall4
0.9392
meanNormalisedRegret
0.1020
meanSpearman
0.6952
v3
top1
0.7659
recall2
0.9095
recall3
0.9504
recall4
0.9747
meanNormalisedRegret
0.0535
meanSpearman
0.8244
lq-centred
top1
0.3080
recall2
0.4681
recall3
0.6039
recall4
0.7148
meanNormalisedRegret
0.3670
meanSpearman
0.2713
lq-raw
top1
0.3109
recall2
0.4871
recall3
0.6039
recall4
0.7002
meanNormalisedRegret
0.3396
meanSpearman
0.2767
postHoc
label
POST HOC: read after the preregistered held-out metrics; diagnostic only
d1RebuildWorstAbsDiff
0.0000
heldOutMse
affineOfD1
14263295958.9596
ridgeRaw
6991007012.7847
inSampleTop1
lqRaw
0.3429
d1
0.5291
lqClipped
0.5408
clippedRefitHeldOut
ranker
lq-clipped-posthoc
roots
2,055
top1
0.5348
recall2
0.7129
recall3
0.8092
recall4
0.8715
meanNormalisedRegret
0.1788
meanRawRegret
2451.3186
meanSpearman
0.5438
clip
-50,000
clipLambda
0.0010
prunedDepth4HeldOut
exact:4:7
top1Agreement
1
top1AgreementLb95
1
meanNormalisedRegret
0
meanNormalisedRegretUb95
0
meanRawRegret
0
workRatioOfMeans
1
priorWorkShare
0
meanPrunedNodes
0
meanWallMs
3586.8220
monotoneViolations
0
pruned:4:7:2,2:center
top1Agreement
0.6718
top1AgreementLb95
0.6530
meanNormalisedRegret
0.0895
meanNormalisedRegretUb95
0.1007
meanRawRegret
1035.1212
workRatioOfMeans
0.0987
priorWorkShare
0
meanPrunedNodes
555.9850
meanWallMs
355.2251
monotoneViolations
0
pruned:4:7:3,3:center
top1Agreement
0.6913
top1AgreementLb95
0.6711
meanNormalisedRegret
0.0724
meanNormalisedRegretUb95
0.0819
meanRawRegret
1019.4365
workRatioOfMeans
0.2009
priorWorkShare
0
meanPrunedNodes
782.3660
meanWallMs
724.0096
monotoneViolations
0
pruned:4:7:4,4:center
top1Agreement
0.7602
top1AgreementLb95
0.7460
meanNormalisedRegret
0.0506
meanNormalisedRegretUb95
0.0548
meanRawRegret
275.2358
workRatioOfMeans
0.3568
priorWorkShare
0
meanPrunedNodes
1008.4257
meanWallMs
1286.2429
monotoneViolations
0
pruned:4:7:3,2:center
top1Agreement
0.6932
top1AgreementLb95
0.6732
meanNormalisedRegret
0.0744
meanNormalisedRegretUb95
0.0834
meanRawRegret
1021.3907
workRatioOfMeans
0.1461
priorWorkShare
0
meanPrunedNodes
785.1558
meanWallMs
527.3962
monotoneViolations
0
pruned:4:7:2,2:kf=CEM
top1Agreement
0.8005
top1AgreementLb95
0.7860
meanNormalisedRegret
0.0356
meanNormalisedRegretUb95
0.0396
meanRawRegret
187.5446
workRatioOfMeans
0.1145
priorWorkShare
0
meanPrunedNodes
562.4218
meanWallMs
393.0244
monotoneViolations
0
pruned:4:7:3,3:kf=CEM
top1Agreement
0.8568
top1AgreementLb95
0.8430
meanNormalisedRegret
0.0211
meanNormalisedRegretUb95
0.0236
meanRawRegret
73.7016
workRatioOfMeans
0.2186
priorWorkShare
0
meanPrunedNodes
770.3010
meanWallMs
758.5048
monotoneViolations
0
pruned:4:7:4,4:kf=CEM
top1Agreement
0.8835
top1AgreementLb95
0.8714
meanNormalisedRegret
0.0147
meanNormalisedRegretUb95
0.0166
meanRawRegret
46.4050
workRatioOfMeans
0.3463
priorWorkShare
0
meanPrunedNodes
962.2078
meanWallMs
1211.4424
monotoneViolations
0
pruned:4:7:3,2:kf=CEM
top1Agreement
0.8403
top1AgreementLb95
0.8290
meanNormalisedRegret
0.0238
meanNormalisedRegretUb95
0.0261
meanRawRegret
86.8332
workRatioOfMeans
0.1571
priorWorkShare
0
meanPrunedNodes
772.1583
meanWallMs
542.1633
monotoneViolations
0
pruned:4:7:2,2:d1
top1Agreement
0.8927
top1AgreementLb95
0.8870
meanNormalisedRegret
0.0145
meanNormalisedRegretUb95
0.0164
meanRawRegret
39.3229
workRatioOfMeans
0.1379
priorWorkShare
0.0808
meanPrunedNodes
585.5709
meanWallMs
486.6896
monotoneViolations
0
pruned:4:7:3,3:d1
top1Agreement
0.9233
top1AgreementLb95
0.9184
meanNormalisedRegret
0.0087
meanNormalisedRegretUb95
0.0098
meanRawRegret
37.2455
workRatioOfMeans
0.2641
priorWorkShare
0.0586
meanPrunedNodes
813.3791
meanWallMs
940.7566
monotoneViolations
0
pruned:4:7:4,4:d1
top1Agreement
0.9383
top1AgreementLb95
0.9321
meanNormalisedRegret
0.0059
meanNormalisedRegretUb95
0.0071
meanRawRegret
16.4791
workRatioOfMeans
0.4089
priorWorkShare
0.0472
meanPrunedNodes
1011.6922
meanWallMs
1461.4425
monotoneViolations
0
pruned:4:7:3,2:d1
top1Agreement
0.9194
top1AgreementLb95
0.9124
meanNormalisedRegret
0.0102
meanNormalisedRegretUb95
0.0116
meanRawRegret
40.4307
workRatioOfMeans
0.1923
priorWorkShare
0.0805
meanPrunedNodes
815.1141
meanWallMs
683.0929
monotoneViolations
0
pruned:4:7:2,2:lq=LQ
top1Agreement
0.7510
top1AgreementLb95
0.7326
meanNormalisedRegret
0.0541
meanNormalisedRegretUb95
0.0596
meanRawRegret
201.2905
workRatioOfMeans
0.1311
priorWorkShare
0.0849
meanPrunedNodes
585.5102
meanWallMs
483.7916
monotoneViolations
0
pruned:4:7:3,3:lq=LQ
top1Agreement
0.7879
top1AgreementLb95
0.7728
meanNormalisedRegret
0.0370
meanNormalisedRegretUb95
0.0413
meanRawRegret
138.7630
workRatioOfMeans
0.2507
priorWorkShare
0.0616
meanPrunedNodes
811.1471
meanWallMs
921.2930
monotoneViolations
0
pruned:4:7:4,4:lq=LQ
top1Agreement
0.8175
top1AgreementLb95
0.8035
meanNormalisedRegret
0.0303
meanNormalisedRegretUb95
0.0346
meanRawRegret
100.7043
workRatioOfMeans
0.3927
priorWorkShare
0.0490
meanPrunedNodes
1008.6723
meanWallMs
1437.6486
monotoneViolations
0
pruned:4:7:3,2:lq=LQ
top1Agreement
0.7786
top1AgreementLb95
0.7630
meanNormalisedRegret
0.0426
meanNormalisedRegretUb95
0.0472
meanRawRegret
151.8085
workRatioOfMeans
0.1831
priorWorkShare
0.0844
meanPrunedNodes
812.7985
meanWallMs
675.4971
monotoneViolations
0
depth5SubPanel
roots
128
selection
every 16th held-out root in cohort order, first 128
exactD5MeanWork
190,100,000
configs
exact:4:7
top1Agreement
0.7969
meanNormalisedRegret
0.0285
meanNormalisedRegretUb95
0.0398
meanRawRegret
145.6493
workRatioOfMeans
1
meanWallMs
4951.4340
pruned:5:7:2,2,2:d1
top1Agreement
0.8047
meanNormalisedRegret
0.0174
meanNormalisedRegretUb95
0.0241
meanRawRegret
188.9523
workRatioOfMeans
1.6784
meanWallMs
8137.6059
pairedVsExactD4
meanNormalisedRegretReduction
0.0111
normalisedLb95
-0.0004
meanRawRegretReduction
-43.3030
rawLb95
-206.5648
roots
128
pruned:5:7:3,3,2:d1
top1Agreement
0.8125
meanNormalisedRegret
0.0182
meanNormalisedRegretUb95
0.0244
meanRawRegret
184.8310
workRatioOfMeans
3.2396
meanWallMs
15777.7039
pairedVsExactD4
meanNormalisedRegretReduction
0.0104
normalisedLb95
0.0002
meanRawRegretReduction
-39.1816
rawLb95
-202.9037
roots
128
pruned:5:7:3,3,3:d1
top1Agreement
0.8281
meanNormalisedRegret
0.0166
meanNormalisedRegretUb95
0.0229
meanRawRegret
179.8801
workRatioOfMeans
4.4857
meanWallMs
21765.4059
pairedVsExactD4
meanNormalisedRegretReduction
0.0119
normalisedLb95
0.0010
meanRawRegretReduction
-34.2307
rawLb95
-199.1113
roots
128
pruned:5:7:2,2,2:lq=LQ
top1Agreement
0.7188
meanNormalisedRegret
0.0486
meanNormalisedRegretUb95
0.0666
meanRawRegret
216.8668
workRatioOfMeans
1.4945
meanWallMs
7607.5829
pairedVsExactD4
meanNormalisedRegretReduction
-0.0201
normalisedLb95
-0.0432
meanRawRegretReduction
-71.2174
rawLb95
-152.5847
roots
128
pruned:5:7:3,3,2:lq=LQ
top1Agreement
0.8203
meanNormalisedRegret
0.0281
meanNormalisedRegretUb95
0.0425
meanRawRegret
118.2179
workRatioOfMeans
2.9120
meanWallMs
14687.9824
pairedVsExactD4
meanNormalisedRegretReduction
0.0004
normalisedLb95
-0.0199
meanRawRegretReduction
27.4315
rawLb95
-78.8987
roots
128
pruned:5:7:3,3,3:lq=LQ
top1Agreement
0.7969
meanNormalisedRegret
0.0338
meanNormalisedRegretUb95
0.0510
meanRawRegret
125.3047
workRatioOfMeans
4.0220
meanWallMs
20016.3636
pairedVsExactD4
meanNormalisedRegretReduction
-0.0052
normalisedLb95
-0.0277
meanRawRegretReduction
20.3447
rawLb95
-85.6609
roots
128
bootstrap
unit
game
resamples
10,000
seed
0x6b660002
Limitations
  • Pilot tier on training-role seeds: 48 games, 16 held-out games (2,060 roots) for the depth-4 metrics and 128 roots from those games for depth 5; game-clustered bounds are reported but the sub-panel's paired depth-5 differences straddle zero.
  • Root-level regret against exact search values is a decision-quality proxy, not a strength measure; the program has seen short-horizon proxies invert three times, and the preregistered gameplay pilot EX-20260902-pruned-search-gameplay-pilot-bf465b1d is the strength test.
  • The ridge lambda selected by cross-validation (0.001) is the smallest value of the preregistered grid; a smaller lambda was not tried. The post-hoc clipped refit is a diagnostic read after the held-out metrics and carries no evidential weight.
  • The fitted Q used inside the pruned search is the per-root-centred fit, as preregistered; its poor ranking makes the lq pruning rows a test of a bad prior, not of the pruning mechanism.
  • Wall budget: the panel (1,434 s) plus the stages (5,958 s, run while two pilot gameplay arms shared the 12 cores) took 7,392 s, 2.7% over the preregistered 7,200 s stop, which was not enforced by a watchdog; a first panel attempt of about 1,380 s was killed by the session harness at 47/48 games before writing anything. All stage outputs are deterministic functions of the inputs, so the overrun changes no number; wall-time figures are contended and only logical work should be compared.
  • The comparator values are the frozen-leaf search's own values; agreement with exact depth 5 measures decision reproduction under the frozen leaf, not correctness.

Open the result record

Agent contextHow to extend this record

To add a reader-facing explanation, write web/content/research/EX-20260902-oneply-q-fit-and-pruned-search-panel-46c75cdf.mdx; it renders above this record on the next request. The registered protocol itself is in the technical record above.

Record file: research/experiments/EX-20260902-oneply-q-fit-and-pruned-search-panel-46c75cdf.json, validated against research/schemas/experiment-v1.schema.json. Protocol hash: 75f3c5aa9f39a7f22f1cf86d260f94323ec65dce668cf8c478330b885fde3d49.

Registered by Claude Code / claude-fable-5-1 (claude-q-learning).