Depth-5-distilled NNUE leaf refined by paired-fitness whole-game evolution inside the depth-3 fair search (v2: operational parameters fixed in the record body)
Successor to EX-20260825-nnue-evolution-d3-bca7f330.
On this page
- Created
- Updated
No explanation has been written for this record yet.
Technical recordThe registered protocol
- Hypothesis
- Successor to EX-20260825-nnue-evolution-d3-bca7f330. The scientific protocol is carried over unchanged: candidate, comparator, information boundary, benchmark tier, leases, metrics, uncertainty method, pass criteria, failure and pass actions, GA constants and stop rules are those of v1 verbatim below. What v2 adds is the set of operational parameters v1 left open, fixed before any leased seed was read (the training lease opened 2026-09-02T04:24:10Z; these values were first recorded at 04:02Z and 04:23Z as in-place amendments of v1, which the repository rule on frozen records does not permit, hence this successor): (1) stage A plays with --move-cap 500 (teacher game lengths are heavy-tailed per the August 25 T0 pilot; labels are search values, so the cap truncates the late-game state distribution but biases no label; capped games are recorded as censored) and admits no new game after a 46,800 s wall sub-budget, in-flight games finishing and being kept, the completed whole games forming the corpus as the v1 stop rule already permits; (2) stage B uses batch 64, lr 3e-4, Huber delta 1.0 rise unit, seed 0x0e701e57, 16 epochs with the best epoch chosen by validation Huber, and 256 held-out probe roots; (3) stage C uses the v1 GA constants with a 21,600 s wall sub-budget and the elite re-selection block 0xa52e0c80 exactly as leased; (4) every stage runs with 32 threads: a seed-free scaling preflight on the already-opened probe block 0xa5277800-0xa527783f (64 complete d4s7 teacher-corpus games, 7,830 roots) gave bit-identical root and game records at 16 and 32 threads and 1.020 vs 1.577 s per root per thread, a 1.29x throughput gain from SMT on the 16 physical cores; thread count changes no decision, label, seed, gate or metric; (5) the held-out screen is opened once on the frozen candidate regardless of the theory's training-signal falsifier, which is evaluated and reported separately from the screen gate. Every stage is launched through approaches/lifetime-objective/nnue-evolution/scripts/pipeline.sh, which hard-codes the leased sub-blocks and records per-stage wall, CPU and peak-RSS usage. v1 protocol as frozen: Candidate: the stock fair expectimax search at depth 3, seven chance strata, terminal utility -1,000,000, policy seed 0xd7075eed, completion-guaranteeing work bound work_bound_for(3,7)+1, 64k-entry direct-mapped table, with an NNUE leaf (8,902-feature sparse class, 135 active, EmbeddingBag(8902,64)->ReLU->32->ReLU->1, output x 17,000 points). Weights: initialised by supervised distillation of a depth-5 seven-stratum fair teacher's sibling-complete root values (512 training games, whole-origin split, Adam on Huber loss, best epoch by validation Huber), then refined by a mutation-only GA: population 32, 60 generations, fitness = mean score of 32 complete paired games per candidate on a fresh training block per generation, 4 elites cloned, tournament size 3, per-tensor Gaussian mutation sigma = 0.05 x tensor std (floor 1e-4), initial cloud 2 sigma, evolution seed 0x0e701e58, controls fair-d3s7 and init-d3s7 playing every block. Final candidate: top 8 of the last generation re-evaluated on a fresh 128-game block, best mean frozen. The frozen candidate is then screened once on 64 never-read paired development games against the identical search with the frozen fair leaf.
- Arms
Arm Name Entry point Manifest Candidate d3s7-evolved-nnue-leaf approaches/lifetime-objective/nnue-evolution/src/bin/evolve.rs– Comparator fair-d3s7 approaches/fair-expectimax/rust-engine/src/leaf.rsresearch/benchmarks/baselines-v1.json- Classification
- algorithmic
- Information boundary
- public-policy
- Benchmark tier
- SCREEN
- Lifecycle
- completed
- Theories tested
- Primary metric
- paired mean whole-game score delta, evolved candidate minus frozen fair leaf, both at the identical d3s7 configuration, 64 held-out games
- Secondary metrics
- paired mean moves delta; numbered clears per move and cover reveals per move; mean occupied cells
- lower quartile, median, maximum score; W-T-L; first-half and second-half paired mean deltas
- ablation arm on the same seeds: the supervised-init (unevolved) NNUE at d3s7, isolating the evolutionary stage's contribution
- reference arm on the same seeds: the fair leaf at d4s7 (the program's standing reference configuration), diagnostic only
- training-curve diagnostics: per-generation best and mean fitness against the paired fair-d3s7 control on the same blocks
- supervised-init diagnostics: whole-origin validation Huber/Pearson and the deployment-faithful ordering probe (d3s7+NNUE top-1 agreement with the teacher's chosen column, teacher-value regret) on held-out corpus roots
- Statistical unit
- whole-game
- Uncertainty method
- one-sided 95% percentile bootstrap over whole games, 20,000 resamples, RNG seed 0xb0071eaf (the unchanged compare.py of the prior leaf-evolution screen), plus a one-sided 95% Student-t lower bound; detection floor 1.645*sd/sqrt(n) reported
- Data role
- public-development
- Seed leases
SL-20260825T063000Z-a52e0300SL-20260825T063000Z-a52e1300
- Whole-origin split
- yes
- Reuse disclosure
- CHECK gates and the T0 timing pilot read only the already-opened development probe block 0xa5276000-0xa5277fff (the Rust engine's gate/benchmark range); no new seed is opened for mechanics or timing. The teacher corpus, every evolution fitness block, and the elite re-selection open training-lease seeds 0xa52e0300-0xa52e0d00; the held-out screen opens 0xa52e1300-0xa52e133f exactly once, after the candidate is frozen. Training and screen ranges are disjoint by construction. The 2026-09-02 thread-scaling preflight read only the already-opened probe block 0xa5277800-0xa527783f.
- Pass criteria
- All CHECK gates passed before the first training seed was read: feature determinism and bounds, information-boundary blindness (score/level/moves_played), reflection consistency on asymmetric boards (finding-13 exclusion), fresh-searcher and 1-vs-4-worker determinism, legality and completed depth under random/zero/saturated weights, leaf finiteness fuzz, serialisation round-trip.
- Every generation artifact and the screen artifact have illegalDecisions 0, incompleteDecisions 0, and censored games reported as censored.
- Held-out screen, 64 paired games on 0xa52e1300-0xa52e133f, candidate vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0.
- Held-out screen: paired mean score delta > 0 in both halves (seeds 0-31 and 32-63).
- Held-out screen: candidate lower-quartile score >= fair-d3s7 lower-quartile score.
- The screened candidate is the exact frozen winner of the elite re-selection (candidate-weights.bin, SHA-256 recorded before the screen lease opens); no other vector is screened.
- On pass
- Freeze candidate-weights.bin with its SHA-256; register a fresh-block replication by a different runner and a fresh-development STANDARD evaluation as successors; do not open protected or final seeds.
- On fail
- Record a valid run with scientific outcome fail for this exact configuration (model class, teacher depth, optimiser, budget, deployment depth); do not adopt the vector; the ablation and reference arms are reported as diagnostics; open no further cohort for this candidate.
- Gate fixed before controlled data
- yes
- Resources
Wall seconds 86400 CPU threads 32 Max host bytes 17179869184 Max GPU bytes – GPU devices – - Stop conditions
- Stop on any rules, information-boundary, legality, determinism, or parity failure.
- The teacher corpus is checkpointed per game and stops at 512 games or the wall budget; whatever whole games completed are the corpus (the count is recorded; the supervised stage does not depend on it being exactly 512).
- Evolution stops at 60 generations, at the wall budget, or when a STOP file appears in the run directory; the candidate is then the elite re-selection over the last completed generation's population.
- A generation or screen artifact with any illegal or incomplete decision voids the run (invalid), not the candidate.
- The screen is evaluated exactly once; no re-run on the same or a different held-out block without a new experiment record.
- Stage A admits no new teacher game after 46,800 s of wall time; games in flight finish and are kept; the corpus is the set of completed whole games (count recorded).
- Stage C stops at 60 generations or 21,600 s of wall time, whichever comes first; the elite re-selection then runs on the last completed generation.
- Expected artifacts
runs/<run-id>/nnue-evolution/gates.log (CHECK)runs/<run-id>/nnue-evolution/corpus.jsonl + corpus-summary.json (stage A)runs/<run-id>/nnue-evolution/pretrain/{epoch-*.bin,init.bin,report.json,probe.json} (stage B)runs/<run-id>/nnue-evolution/evolve/{config.json,progress.jsonl,gen-*.json,population-*.bin,final-fitness.json,candidate-weights.bin,selection.json} (stage C)runs/<run-id>/nnue-evolution/screen/heldout.json + compare-*.json (stage D)runs/<run-id>/nnue-evolution/{pipeline.log,rusage.jsonl,preflight-d4-t16.log,preflight-d4-t32.log,analysis.json,analysis.md}artifacts emitted by evolve and screen carry this experiment id in their config field
- Amendments
Timestamp Before controlled data Reason 2026-09-03T02:57:51Z no lifecycle advanced running -> completed after run RUN-20260902T035644Z-c1fd8987 and result RS-20260903T025751Z-6577b33e were written; protocol content otherwise unchanged. The preregistration hash 971a39ef6bb5275a951342245a119cb3d234bed334aa28e2ecfe44eeeeea4dae is retained in the run and lease records.
Technical recordResults recorded against this protocol
Valid run of the frozen protocol EX-20260902-nnue-evolution-d3-v2-49c18bc2 on the whole workstation (32 threads), every stage completed, every artifact with zero illegal and zero incomplete decisions. Stage A: the depth-5 seven-stratum teacher played 177 complete games (21,618 sibling-complete labelled roots, 3 stopped at the 500-move cap) before the 46,800 s new-game cutoff; the teacher ran 77.5 s per root, four to seven times slower than the pilot projected, so the corpus is about a third of the 512 games the protocol allowed for. Stage B: the supervised warm start reached validation Huber 0.6873 rise units at epoch 8 of 16 on a whole-origin split (17,641/3,977 roots), and the deployment-faithful ordering probe put the depth-3 search with that leaf at top-1 agreement 0.4414 with the teacher on 256 held-out roots (mean teacher-value regret 2,395 points). Stage C: 60 generations of the mutation-only GA (population 32, 32 paired games per candidate per fresh block, fair-d3s7 and warm-start controls on every block). The population separated steadily from its warm start - ten-generation paired mean margins over the init control of +4,682, +15,853, +22,704, +32,278, +34,037, +40,431 points, best candidate +76,679 in the last block - but never approached the fair control: the population mean was above the fair leaf in 0 of the final 10 generations and the best candidate in 0 (mean margin -121,896), so the theory's training-signal falsifier fails. Elite re-selection on 128 fresh games froze candidate-29 at 202,237 (finalists spanned 187,352 to 202,237). Stage D: on the 64 never-read held-out games (0xa52e1300, opened once at 2026-09-03T02:41:29Z after the candidate's SHA-256 was recorded), the evolved candidate averaged 190,961 against the frozen fair leaf's 297,926 at the identical depth-3 configuration: paired delta -106,964 (bootstrap 95% bounds -146,580 to -69,983, Student-t lower bound -145,890, detection floor 38,357), W-T-L 14-0-50, both halves negative (-109,139 / -104,790), lower quartile 122,253 against 158,387. Every preregistered pass criterion except artifact integrity and candidate identity fails: scientific outcome fail for this exact configuration. The ablation arm shows what evolution did contribute: the unevolved warm start averaged 155,586 (-142,340 against the fair leaf), and the evolved candidate beat it on the same seeds by +35,375 (bootstrap lower bound +16,899, floor 18,949, W-T-L 36-0-28, both halves positive). Whole-game evolution with common random numbers therefore moves a 572k-weight leaf on the deployed objective, which the first (CMA-ES) leaf evolution could not show; it moved it about a quarter of the way from a warm start that plays at half the fair leaf's level. The reference arm reproduced the program's standing result: fair depth 4 over fair depth 3 +97,059 (lower bound +29,330). Read: the claim is not supported as tested; the mechanism's evolutionary leg is supported, its distillation leg is the weak link (a warm start that holds the teacher's values but not its ordering), and the budget (177 teacher games, 60 generations) was too small for evolution to cover the distance.
- ✓All CHECK gates passed before the first training seed was read — observed: 11 of 11 gates passed on the already-open probe block (gates.log), rebuilt binaries, 9/9 unit tests
- ✓Every generation artifact and the screen artifact: illegalDecisions 0, incompleteDecisions 0, censored games reported — observed: 60 generation artifacts: 0 illegal, 0 incomplete, 0 censored; screen: 0 illegal, 0 incomplete, 0 censored in all four arms
- ✕Held-out screen, candidate vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0 — observed: delta -106,964; bootstrap LB -146,580, UB -69,983; t LB -145,890; paired sd 186,538; floor 38,357; W-T-L 14-0-50
- ✕Held-out screen: paired mean score delta > 0 in both halves (seeds 0-31 and 32-63) — observed: first half -109,139, second half -104,790
- ✕Held-out screen: candidate lower-quartile score >= fair-d3s7 lower-quartile score — observed: candidate Q25 122,253 vs fair-d3s7 Q25 158,387 (delta -36,134)
- ✓The screened candidate is the exact frozen winner of the elite re-selection; SHA-256 recorded before the screen lease opened — observed: candidate-29 of generation 59, mean 202,237 over 128 fresh games; sha256 edd0d2efd181de43f35d62c4df784cb2c789db3ae296be93a1c6c59082034a9f; recorded 2026-09-03T02:41:29Z, lease opened 2026-09-03T02:41:29Z by the screen stage
- ✕Theory falsifier (training signal): population mean fitness > paired fair-d3s7 control in the majority of the final 10 generations — observed: mean above fair in 0/10, top-4 in 0/10, best in 0/10; mean margin -121,896, top-4 margin -96,788
- ✓Theory falsifier (supervised initialisation): distilled NNUE's within-root top-1 agreement with the depth-5 teacher on whole-origin held-out roots exceeds the 0.30 level materially — observed: deployment-faithful probe (d3s7 search with the NNUE leaf vs the teacher's column) top-1 0.4414 on 256 held-out roots, regret 2,395 points; for scale, the leaf-swing diagnostic put the frozen fair leaf at 0.720 and a zero leaf at 0.450 on a different 400-root sample
- –Theory falsifier (ablative): if the candidate passes but the unevolved init passes equally, evolution contributed nothing — observed: precondition not met (candidate failed); recorded anyway: candidate minus init +35,375 (bootstrap LB +16,899, t LB +16,145, floor 18,949, W-T-L 36-0-28, halves +47,004 / +23,746); init minus fair -142,340 (LB -178,492)
Technical recordRecorded metrics
- seedsStartHex
- 0xa52e1300
- games
- 64
- moveCap
- 2,000
- role
- public-development, never read before this screen (lease SL-20260825T063000Z-a52e1300, opened once)
- candidate
- games
- 64
- mean
- 190961.2188
- median
- 160107.5000
- q25
- 122253.2500
- max
- 477,758
- movesMean
- 58.5000
- clearsPerMove
- 1.7131
- revealsPerMove
- 0.9124
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 443.8840
- init-d3s7
- games
- 64
- mean
- 155585.9688
- median
- 140,282
- q25
- 107423.2500
- max
- 401,038
- movesMean
- 48.7500
- clearsPerMove
- 1.5510
- revealsPerMove
- 0.7554
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 352.7846
- fair-d3s7
- games
- 64
- mean
- 297925.5938
- median
- 254999.5000
- q25
- 158387.2500
- max
- 844,318
- movesMean
- 88.0313
- clearsPerMove
- 1.9549
- revealsPerMove
- 1.0722
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 265.4953
- fair-d4s7
- games
- 64
- mean
- 394984.3125
- median
- 302,556
- q25
- 198524.7500
- max
- 1,663,637
- movesMean
- 114.2969
- clearsPerMove
- 2.0547
- revealsPerMove
- 1.1494
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 11302.7039
- candidate-vs-fair-d3s7
- meanDelta
- -106964.3750
- bootstrapLower95
- -146579.5391
- bootstrapUpper95
- -69983.3727
- studentTLower95
- -145890.1465
- pairedSd
- 186537.5328
- detectionFloor
- 38356.7802
- wtl
- 14
- 0
- 50
- halves
- -109138.6563
- -104790.0938
- q25Delta
- -36,134
- movesDelta
- -29.5313
- init-vs-fair-d3s7
- meanDelta
- -142339.6250
- bootstrapLower95
- -178491.5461
- bootstrapUpper95
- -107770.6313
- studentTLower95
- -178510.6720
- pairedSd
- 173336.5226
- detectionFloor
- 35642.3225
- wtl
- 9
- 0
- 55
- halves
- -156142.6563
- -128536.5938
- q25Delta
- -50,964
- movesDelta
- -39.2813
- candidate-vs-init
- meanDelta
- 35375.2500
- bootstrapLower95
- 16898.9273
- bootstrapUpper95
- 54629.2750
- studentTLower95
- 16144.9914
- pairedSd
- 92153.9860
- detectionFloor
- 18949.1634
- wtl
- 36
- 0
- 28
- halves
- 47,004
- 23746.5000
- q25Delta
- 14,830
- movesDelta
- 9.7500
- fair-d4s7-vs-fair-d3s7
- meanDelta
- 97058.7188
- bootstrapLower95
- 29329.5438
- bootstrapUpper95
- 165453.2305
- studentTLower95
- 27011.3702
- pairedSd
- 335676.3163
- detectionFloor
- 69023.4425
- wtl
- 41
- 0
- 23
- halves
- 75457.7813
- 118659.6563
- q25Delta
- 40137.5000
- movesDelta
- 26.2656
- generationsCompleted
- 60
- tenGenerationBlockAverages
- generations
- 0-9
- populationMean
- 159340.4412
- top4Mean
- 176605.0742
- best
- 183544.7219
- meanMinusInit
- 4682.4068
- bestMinusInit
- 28886.6875
- meanMinusFair
- -172846.5150
- bestMinusFair
- -148642.2344
- generations
- 10-19
- populationMean
- 168372.4929
- top4Mean
- 189184.1141
- best
- 194671.8250
- meanMinusInit
- 15852.6179
- bestMinusInit
- 42151.9500
- meanMinusFair
- -154441.2915
- bestMinusFair
- -128141.9594
- generations
- 20-29
- populationMean
- 174464.3515
- top4Mean
- 197282.4398
- best
- 203909.7531
- meanMinusInit
- 22703.5952
- bestMinusInit
- 52148.9969
- meanMinusFair
- -155743.1267
- bestMinusFair
- -126297.7250
- generations
- 30-39
- populationMean
- 181219.7211
- top4Mean
- 203060.9922
- best
- 209035.9844
- meanMinusInit
- 32277.6086
- bestMinusInit
- 60093.8719
- meanMinusFair
- -125175.9352
- bestMinusFair
- -97359.6719
- generations
- 40-49
- populationMean
- 184679.3310
- top4Mean
- 206650.3125
- best
- 213143.7469
- meanMinusInit
- 34036.9685
- bestMinusInit
- 62501.3844
- meanMinusFair
- -160026.6253
- bestMinusFair
- -131562.2094
- generations
- 50-59
- populationMean
- 190907.8234
- top4Mean
- 216015.6836
- best
- 227155.7656
- meanMinusInit
- 40431.1578
- bestMinusInit
- 76679.1000
- meanMinusFair
- -121896.0891
- bestMinusFair
- -85648.1469
- trainingSignalCheck
- rule
- population mean fitness > paired fair-d3s7 control in the majority of the final 10 completed generations (theory falsifier 2)
- window
- 50
- 51
- 52
- 53
- 54
- 55
- 56
- 57
- 58
- 59
- generationsMeanAboveFair
- 0
- generationsTop4AboveFair
- 0
- generationsBestAboveFair
- 0
- generationsInitAboveFair
- 0
- meanMarginLast10
- -121896.0891
- top4MarginLast10
- -96788.2289
- passed
- false
- artifactIntegrity
- generationArtifacts
- 60
- illegalDecisions
- 0
- incompleteDecisions
- 0
- censoredGames
- 0
- config
- population
- 32
- games
- 32
- generations
- 60
- elites
- 4
- tournament
- 3
- sigmaRel
- 0.0500
- seed
- 0x0e701e58
- leaseStart
- 0xa52e0500
- moveCap
- 2,000
- wallSeconds
- 21,600
- finalGeneration
- 59
- finalists
- 12
- 1
- 6
- 29
- 30
- 23
- 13
- 4
- blockStart
- 0xa52e0c80
- selectGames
- 128
- winner
- candidate-29
- winnerMean
- 202236.7734
- games
- 177
- roots
- 21,618
- censoredGames
- 3
- score
- n
- 177
- mean
- 433131.1525
- sd
- 323856.2717
- min
- 103,368
- q25
- 229,522
- median
- 348,847
- q75
- 549,029
- max
- 1,889,608
- moves
- n
- 177
- mean
- 122.1356
- sd
- 85.7326
- min
- 34
- q25
- 70
- median
- 100
- q75
- 150
- max
- 500
- wallSecondsPerGame
- n
- 177
- mean
- 9461.8629
- sd
- 7988.6131
- min
- 1209.2638
- q25
- 4178.7932
- median
- 7476.2450
- q75
- 11333.4374
- max
- 49497.1916
- secondsPerRoot
- 77.4702
- rootsPerGame
- 122.1356
- legalColumnsPerRoot
- n
- 21,618
- mean
- 6.8572
- sd
- 0.6843
- min
- 1
- q25
- 7
- median
- 7
- q75
- 7
- max
- 7
- teacherSiblingSpreadPoints
- n
- 21,618
- mean
- 40559.6546
- sd
- 160027.6823
- min
- 0
- q25
- 2448.7576
- median
- 4187.9510
- q75
- 7480.5930
- max
- 1017575.6843
- report
- roots
- 21,618
- games
- 177
- trainRoots
- 17,641
- valRoots
- 3,977
- epochs
- 16
- batch
- 64
- lr
- 0.0003
- seed
- 0x0e701e57
- bestEpoch
- 8
- bestValHuber
- 0.6873
- probe
- probeRoots
- 256
- top1
- 0.4414
- meanTeacherValueRegret
- 2394.6708
- epochs
- epoch
- 0
- trainHuber
- 2.3016
- valHuber
- 1.8859
- valPearson
- 0.8790
- epoch
- 1
- trainHuber
- 1.2819
- valHuber
- 0.9546
- valPearson
- 0.9314
- epoch
- 2
- trainHuber
- 0.7173
- valHuber
- 0.8128
- valPearson
- 0.9449
- epoch
- 3
- trainHuber
- 0.5488
- valHuber
- 0.7449
- valPearson
- 0.9491
- epoch
- 4
- trainHuber
- 0.4514
- valHuber
- 0.7100
- valPearson
- 0.9512
- epoch
- 5
- trainHuber
- 0.3886
- valHuber
- 0.6918
- valPearson
- 0.9520
- epoch
- 6
- trainHuber
- 0.3371
- valHuber
- 0.6906
- valPearson
- 0.9519
- epoch
- 7
- trainHuber
- 0.3036
- valHuber
- 0.6919
- valPearson
- 0.9500
- epoch
- 8
- trainHuber
- 0.2706
- valHuber
- 0.6873
- valPearson
- 0.9504
- epoch
- 9
- trainHuber
- 0.2443
- valHuber
- 0.6939
- valPearson
- 0.9502
- epoch
- 10
- trainHuber
- 0.2246
- valHuber
- 0.6993
- valPearson
- 0.9479
- epoch
- 11
- trainHuber
- 0.2092
- valHuber
- 0.6955
- valPearson
- 0.9496
- epoch
- 12
- trainHuber
- 0.1915
- valHuber
- 0.7119
- valPearson
- 0.9480
- epoch
- 13
- trainHuber
- 0.1807
- valHuber
- 0.6936
- valPearson
- 0.9479
- epoch
- 14
- trainHuber
- 0.1729
- valHuber
- 0.6995
- valPearson
- 0.9485
- epoch
- 15
- trainHuber
- 0.1620
- valHuber
- 0.6969
- valPearson
- 0.9470
- label
- corpus
- command
- /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/teacher_corpus
- --seeds-start
- 0xa52e0300
- --games
- 512
- --threads
- 32
- --depth
- 5
- --move-cap
- 500
- --wall-seconds
- 46800
- --out
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/corpus
- startedAt
- 2026-09-02T04:24:10Z
- endedAt
- 2026-09-02T22:39:58Z
- wallSeconds
- 65747.5370
- userSeconds
- 1667946.0280
- systemSeconds
- 641.5270
- peakRssBytes
- 1,615,892,480
- exitCode
- 0
- label
- pretrain
- command
- /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/pretrain
- --corpus
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/corpus/corpus.jsonl
- --out
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain
- --epochs
- 16
- --batch
- 64
- --lr
- 3e-4
- --seed
- 0x0e701e57
- --probe-roots
- 256
- --threads
- 32
- startedAt
- 2026-09-02T22:42:11Z
- endedAt
- 2026-09-02T22:42:25Z
- wallSeconds
- 13.9780
- userSeconds
- 45.2200
- systemSeconds
- 0.1080
- peakRssBytes
- 265,060,352
- exitCode
- 0
- label
- evolve
- command
- /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
- --init
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain/init.bin
- --lease-start
- 0xa52e0500
- --out
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve
- --population
- 32
- --games
- 32
- --generations
- 60
- --elites
- 4
- --tournament
- 3
- --sigma-rel
- 0.05
- --seed
- 0x0e701e58
- --threads
- 32
- --wall-seconds
- 21600
- --move-cap
- 2000
- startedAt
- 2026-09-02T22:42:25Z
- endedAt
- 2026-09-03T02:37:19Z
- wallSeconds
- 14093.6190
- userSeconds
- 441673.1460
- systemSeconds
- 184.6580
- peakRssBytes
- 477,446,144
- exitCode
- 0
- label
- select
- command
- /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
- --select
- --lease-start
- 0xa52e0500
- --out
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve
- --generations
- 60
- --games
- 32
- --select-games
- 128
- --threads
- 32
- --move-cap
- 2000
- startedAt
- 2026-09-03T02:37:19Z
- endedAt
- 2026-09-03T02:41:29Z
- wallSeconds
- 250.4040
- userSeconds
- 7852.4190
- systemSeconds
- 3.9940
- peakRssBytes
- 252,526,592
- exitCode
- 0
- label
- screen
- command
- /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/screen
- --candidate
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve/candidate-weights.bin
- --init
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain/init.bin
- --seeds-start
- 0xa52e1300
- --games
- 64
- --threads
- 32
- --move-cap
- 2000
- --out
- /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/screen/heldout.json
- startedAt
- 2026-09-03T02:41:29Z
- endedAt
- 2026-09-03T02:52:22Z
- wallSeconds
- 652.9870
- userSeconds
- 12346.6780
- systemSeconds
- 3.8980
- peakRssBytes
- 1,792,266,240
- exitCode
- 0
- Single 64-game held-out screen: the paired detection floor is about 38,000 points for the primary contrast and 19,000 for the ablation; the primary result is far outside its floor, the ablation clears its own.
- The teacher corpus reached 177 of the 512 games the protocol allowed for because the depth-5 teacher ran four to seven times slower than the pilot projected; the protocol makes the completed games the corpus, so the result is valid, but it rejects this configuration at this corpus size, not the design at 512 games.
- The supervised warm start plays at about half the fair leaf's level; the deployment-faithful probe (0.441 top-1) and the leaf-swing diagnostic (zero leaf 0.450 on a different sample) suggest imitation of state values barely improved the search's ordering over no leaf at all. Evolution then had roughly 150,000 paired points to make up in 60 generations and made up about 35,000 to 40,000.
- The per-game artifact holds four arms x 64 games (256 rows); the primary contrast pairs the candidate and fair-d3s7 rows by seed.
- The fair-d4s7 arm is diagnostic only and reproduces the standing depth-4-over-depth-3 result on these seeds; it is not part of the gate.
- Two operational faults during the unattended chain (a false liveness reading, then a script replacement that crashed the corpus stage's shell after its binary had exited 0) are recorded in the run record; the 'corpus: done' marker in pipeline.log was appended by hand with a note. No artifact was affected.
- The elite re-selection block (0xa52e0c80, 128 games) and every fitness block are training-lease seeds; the last generation's 32-game leaders read 20,000 to 60,000 above their 128-game re-selection means, which is the best-of-32 bias the re-selection exists to remove.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/EX-20260902-nnue-evolution-d3-v2-49c18bc2.mdx; it renders above this record on the next request. The registered protocol itself is in the technical record above.
Record file: research/experiments/EX-20260902-nnue-evolution-d3-v2-49c18bc2.json, validated against research/schemas/experiment-v1.schema.json. Protocol hash: 66d163bb2d39cf4da157a1044be766eb124af28c158a7750e2064ad61edda942.
Registered by Claude Code / claude-fable-5-1 (claude-code-evolution-experiment).