On this page
Dates
Created
Updated
Record idEX-20260902-nnue-evolution-d3-v2-49c18bc2

No explanation has been written for this record yet.

Technical recordThe registered protocolEX-20260902-nnue-evolution-d3-v2-49c18bc2
Hypothesis
Successor to EX-20260825-nnue-evolution-d3-bca7f330. The scientific protocol is carried over unchanged: candidate, comparator, information boundary, benchmark tier, leases, metrics, uncertainty method, pass criteria, failure and pass actions, GA constants and stop rules are those of v1 verbatim below. What v2 adds is the set of operational parameters v1 left open, fixed before any leased seed was read (the training lease opened 2026-09-02T04:24:10Z; these values were first recorded at 04:02Z and 04:23Z as in-place amendments of v1, which the repository rule on frozen records does not permit, hence this successor): (1) stage A plays with --move-cap 500 (teacher game lengths are heavy-tailed per the August 25 T0 pilot; labels are search values, so the cap truncates the late-game state distribution but biases no label; capped games are recorded as censored) and admits no new game after a 46,800 s wall sub-budget, in-flight games finishing and being kept, the completed whole games forming the corpus as the v1 stop rule already permits; (2) stage B uses batch 64, lr 3e-4, Huber delta 1.0 rise unit, seed 0x0e701e57, 16 epochs with the best epoch chosen by validation Huber, and 256 held-out probe roots; (3) stage C uses the v1 GA constants with a 21,600 s wall sub-budget and the elite re-selection block 0xa52e0c80 exactly as leased; (4) every stage runs with 32 threads: a seed-free scaling preflight on the already-opened probe block 0xa5277800-0xa527783f (64 complete d4s7 teacher-corpus games, 7,830 roots) gave bit-identical root and game records at 16 and 32 threads and 1.020 vs 1.577 s per root per thread, a 1.29x throughput gain from SMT on the 16 physical cores; thread count changes no decision, label, seed, gate or metric; (5) the held-out screen is opened once on the frozen candidate regardless of the theory's training-signal falsifier, which is evaluated and reported separately from the screen gate. Every stage is launched through approaches/lifetime-objective/nnue-evolution/scripts/pipeline.sh, which hard-codes the leased sub-blocks and records per-stage wall, CPU and peak-RSS usage. v1 protocol as frozen: Candidate: the stock fair expectimax search at depth 3, seven chance strata, terminal utility -1,000,000, policy seed 0xd7075eed, completion-guaranteeing work bound work_bound_for(3,7)+1, 64k-entry direct-mapped table, with an NNUE leaf (8,902-feature sparse class, 135 active, EmbeddingBag(8902,64)->ReLU->32->ReLU->1, output x 17,000 points). Weights: initialised by supervised distillation of a depth-5 seven-stratum fair teacher's sibling-complete root values (512 training games, whole-origin split, Adam on Huber loss, best epoch by validation Huber), then refined by a mutation-only GA: population 32, 60 generations, fitness = mean score of 32 complete paired games per candidate on a fresh training block per generation, 4 elites cloned, tournament size 3, per-tensor Gaussian mutation sigma = 0.05 x tensor std (floor 1e-4), initial cloud 2 sigma, evolution seed 0x0e701e58, controls fair-d3s7 and init-d3s7 playing every block. Final candidate: top 8 of the last generation re-evaluated on a fresh 128-game block, best mean frozen. The frozen candidate is then screened once on 64 never-read paired development games against the identical search with the frozen fair leaf.
Arms
ArmNameEntry pointManifest
Candidated3s7-evolved-nnue-leafapproaches/lifetime-objective/nnue-evolution/src/bin/evolve.rs
Comparatorfair-d3s7approaches/fair-expectimax/rust-engine/src/leaf.rsresearch/benchmarks/baselines-v1.json
Classification
algorithmic
Information boundary
public-policy
Benchmark tier
SCREEN
Lifecycle
completed
Primary metric
paired mean whole-game score delta, evolved candidate minus frozen fair leaf, both at the identical d3s7 configuration, 64 held-out games
Secondary metrics
  • paired mean moves delta; numbered clears per move and cover reveals per move; mean occupied cells
  • lower quartile, median, maximum score; W-T-L; first-half and second-half paired mean deltas
  • ablation arm on the same seeds: the supervised-init (unevolved) NNUE at d3s7, isolating the evolutionary stage's contribution
  • reference arm on the same seeds: the fair leaf at d4s7 (the program's standing reference configuration), diagnostic only
  • training-curve diagnostics: per-generation best and mean fitness against the paired fair-d3s7 control on the same blocks
  • supervised-init diagnostics: whole-origin validation Huber/Pearson and the deployment-faithful ordering probe (d3s7+NNUE top-1 agreement with the teacher's chosen column, teacher-value regret) on held-out corpus roots
Statistical unit
whole-game
Uncertainty method
one-sided 95% percentile bootstrap over whole games, 20,000 resamples, RNG seed 0xb0071eaf (the unchanged compare.py of the prior leaf-evolution screen), plus a one-sided 95% Student-t lower bound; detection floor 1.645*sd/sqrt(n) reported
Data role
public-development
Seed leases
  • SL-20260825T063000Z-a52e0300
  • SL-20260825T063000Z-a52e1300
Whole-origin split
yes
Reuse disclosure
CHECK gates and the T0 timing pilot read only the already-opened development probe block 0xa5276000-0xa5277fff (the Rust engine's gate/benchmark range); no new seed is opened for mechanics or timing. The teacher corpus, every evolution fitness block, and the elite re-selection open training-lease seeds 0xa52e0300-0xa52e0d00; the held-out screen opens 0xa52e1300-0xa52e133f exactly once, after the candidate is frozen. Training and screen ranges are disjoint by construction. The 2026-09-02 thread-scaling preflight read only the already-opened probe block 0xa5277800-0xa527783f.
Pass criteria
  1. All CHECK gates passed before the first training seed was read: feature determinism and bounds, information-boundary blindness (score/level/moves_played), reflection consistency on asymmetric boards (finding-13 exclusion), fresh-searcher and 1-vs-4-worker determinism, legality and completed depth under random/zero/saturated weights, leaf finiteness fuzz, serialisation round-trip.
  2. Every generation artifact and the screen artifact have illegalDecisions 0, incompleteDecisions 0, and censored games reported as censored.
  3. Held-out screen, 64 paired games on 0xa52e1300-0xa52e133f, candidate vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0.
  4. Held-out screen: paired mean score delta > 0 in both halves (seeds 0-31 and 32-63).
  5. Held-out screen: candidate lower-quartile score >= fair-d3s7 lower-quartile score.
  6. The screened candidate is the exact frozen winner of the elite re-selection (candidate-weights.bin, SHA-256 recorded before the screen lease opens); no other vector is screened.
On pass
Freeze candidate-weights.bin with its SHA-256; register a fresh-block replication by a different runner and a fresh-development STANDARD evaluation as successors; do not open protected or final seeds.
On fail
Record a valid run with scientific outcome fail for this exact configuration (model class, teacher depth, optimiser, budget, deployment depth); do not adopt the vector; the ablation and reference arms are reported as diagnostics; open no further cohort for this candidate.
Gate fixed before controlled data
yes
Resources
Wall seconds86400
CPU threads32
Max host bytes17179869184
Max GPU bytes
GPU devices
Stop conditions
  1. Stop on any rules, information-boundary, legality, determinism, or parity failure.
  2. The teacher corpus is checkpointed per game and stops at 512 games or the wall budget; whatever whole games completed are the corpus (the count is recorded; the supervised stage does not depend on it being exactly 512).
  3. Evolution stops at 60 generations, at the wall budget, or when a STOP file appears in the run directory; the candidate is then the elite re-selection over the last completed generation's population.
  4. A generation or screen artifact with any illegal or incomplete decision voids the run (invalid), not the candidate.
  5. The screen is evaluated exactly once; no re-run on the same or a different held-out block without a new experiment record.
  6. Stage A admits no new teacher game after 46,800 s of wall time; games in flight finish and are kept; the corpus is the set of completed whole games (count recorded).
  7. Stage C stops at 60 generations or 21,600 s of wall time, whichever comes first; the elite re-selection then runs on the last completed generation.
Expected artifacts
  • runs/<run-id>/nnue-evolution/gates.log (CHECK)
  • runs/<run-id>/nnue-evolution/corpus.jsonl + corpus-summary.json (stage A)
  • runs/<run-id>/nnue-evolution/pretrain/{epoch-*.bin,init.bin,report.json,probe.json} (stage B)
  • runs/<run-id>/nnue-evolution/evolve/{config.json,progress.jsonl,gen-*.json,population-*.bin,final-fitness.json,candidate-weights.bin,selection.json} (stage C)
  • runs/<run-id>/nnue-evolution/screen/heldout.json + compare-*.json (stage D)
  • runs/<run-id>/nnue-evolution/{pipeline.log,rusage.jsonl,preflight-d4-t16.log,preflight-d4-t32.log,analysis.json,analysis.md}
  • artifacts emitted by evolve and screen carry this experiment id in their config field
Amendments
TimestampBefore controlled dataReason
2026-09-03T02:57:51Znolifecycle advanced running -> completed after run RUN-20260902T035644Z-c1fd8987 and result RS-20260903T025751Z-6577b33e were written; protocol content otherwise unchanged. The preregistration hash 971a39ef6bb5275a951342245a119cb3d234bed334aa28e2ecfe44eeeeea4dae is retained in the run and lease records.
Technical recordResults recorded against this protocol1 record
valid runoutcome: failnot-supported-as-testedtier: public-developmentRS-20260903T025751Z-6577b33e

Valid run of the frozen protocol EX-20260902-nnue-evolution-d3-v2-49c18bc2 on the whole workstation (32 threads), every stage completed, every artifact with zero illegal and zero incomplete decisions. Stage A: the depth-5 seven-stratum teacher played 177 complete games (21,618 sibling-complete labelled roots, 3 stopped at the 500-move cap) before the 46,800 s new-game cutoff; the teacher ran 77.5 s per root, four to seven times slower than the pilot projected, so the corpus is about a third of the 512 games the protocol allowed for. Stage B: the supervised warm start reached validation Huber 0.6873 rise units at epoch 8 of 16 on a whole-origin split (17,641/3,977 roots), and the deployment-faithful ordering probe put the depth-3 search with that leaf at top-1 agreement 0.4414 with the teacher on 256 held-out roots (mean teacher-value regret 2,395 points). Stage C: 60 generations of the mutation-only GA (population 32, 32 paired games per candidate per fresh block, fair-d3s7 and warm-start controls on every block). The population separated steadily from its warm start - ten-generation paired mean margins over the init control of +4,682, +15,853, +22,704, +32,278, +34,037, +40,431 points, best candidate +76,679 in the last block - but never approached the fair control: the population mean was above the fair leaf in 0 of the final 10 generations and the best candidate in 0 (mean margin -121,896), so the theory's training-signal falsifier fails. Elite re-selection on 128 fresh games froze candidate-29 at 202,237 (finalists spanned 187,352 to 202,237). Stage D: on the 64 never-read held-out games (0xa52e1300, opened once at 2026-09-03T02:41:29Z after the candidate's SHA-256 was recorded), the evolved candidate averaged 190,961 against the frozen fair leaf's 297,926 at the identical depth-3 configuration: paired delta -106,964 (bootstrap 95% bounds -146,580 to -69,983, Student-t lower bound -145,890, detection floor 38,357), W-T-L 14-0-50, both halves negative (-109,139 / -104,790), lower quartile 122,253 against 158,387. Every preregistered pass criterion except artifact integrity and candidate identity fails: scientific outcome fail for this exact configuration. The ablation arm shows what evolution did contribute: the unevolved warm start averaged 155,586 (-142,340 against the fair leaf), and the evolved candidate beat it on the same seeds by +35,375 (bootstrap lower bound +16,899, floor 18,949, W-T-L 36-0-28, both halves positive). Whole-game evolution with common random numbers therefore moves a 572k-weight leaf on the deployed objective, which the first (CMA-ES) leaf evolution could not show; it moved it about a quarter of the way from a warm start that plays at half the fair leaf's level. The reference arm reproduced the program's standing result: fair depth 4 over fair depth 3 +97,059 (lower bound +29,330). Read: the claim is not supported as tested; the mechanism's evolutionary leg is supported, its distillation leg is the weak link (a warm start that holds the teacher's values but not its ordering), and the budget (177 teacher games, 60 generations) was too small for evolution to cover the distance.

What it had to pass
  • All CHECK gates passed before the first training seed was read — observed: 11 of 11 gates passed on the already-open probe block (gates.log), rebuilt binaries, 9/9 unit tests
  • Every generation artifact and the screen artifact: illegalDecisions 0, incompleteDecisions 0, censored games reported — observed: 60 generation artifacts: 0 illegal, 0 incomplete, 0 censored; screen: 0 illegal, 0 incomplete, 0 censored in all four arms
  • Held-out screen, candidate vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0 — observed: delta -106,964; bootstrap LB -146,580, UB -69,983; t LB -145,890; paired sd 186,538; floor 38,357; W-T-L 14-0-50
  • Held-out screen: paired mean score delta > 0 in both halves (seeds 0-31 and 32-63) — observed: first half -109,139, second half -104,790
  • Held-out screen: candidate lower-quartile score >= fair-d3s7 lower-quartile score — observed: candidate Q25 122,253 vs fair-d3s7 Q25 158,387 (delta -36,134)
  • The screened candidate is the exact frozen winner of the elite re-selection; SHA-256 recorded before the screen lease opened — observed: candidate-29 of generation 59, mean 202,237 over 128 fresh games; sha256 edd0d2efd181de43f35d62c4df784cb2c789db3ae296be93a1c6c59082034a9f; recorded 2026-09-03T02:41:29Z, lease opened 2026-09-03T02:41:29Z by the screen stage
  • Theory falsifier (training signal): population mean fitness > paired fair-d3s7 control in the majority of the final 10 generations — observed: mean above fair in 0/10, top-4 in 0/10, best in 0/10; mean margin -121,896, top-4 margin -96,788
  • Theory falsifier (supervised initialisation): distilled NNUE's within-root top-1 agreement with the depth-5 teacher on whole-origin held-out roots exceeds the 0.30 level materially — observed: deployment-faithful probe (d3s7 search with the NNUE leaf vs the teacher's column) top-1 0.4414 on 256 held-out roots, regret 2,395 points; for scale, the leaf-swing diagnostic put the frozen fair leaf at 0.720 and a zero leaf at 0.450 on a different 400-root sample
  • Theory falsifier (ablative): if the candidate passes but the unevolved init passes equally, evolution contributed nothing — observed: precondition not met (candidate failed); recorded anyway: candidate minus init +35,375 (bootstrap LB +16,899, t LB +16,145, floor 18,949, W-T-L 36-0-28, halves +47,004 / +23,746); init minus fair -142,340 (LB -178,492)
Technical recordRecorded metricsRS-20260903T025751Z-6577b33e
cohort
seedsStartHex
0xa52e1300
games
64
moveCap
2,000
role
public-development, never read before this screen (lease SL-20260825T063000Z-a52e1300, opened once)
arms
candidate
games
64
mean
190961.2188
median
160107.5000
q25
122253.2500
max
477,758
movesMean
58.5000
clearsPerMove
1.7131
revealsPerMove
0.9124
censored
0
illegal
0
incomplete
0
wallSeconds
443.8840
init-d3s7
games
64
mean
155585.9688
median
140,282
q25
107423.2500
max
401,038
movesMean
48.7500
clearsPerMove
1.5510
revealsPerMove
0.7554
censored
0
illegal
0
incomplete
0
wallSeconds
352.7846
fair-d3s7
games
64
mean
297925.5938
median
254999.5000
q25
158387.2500
max
844,318
movesMean
88.0313
clearsPerMove
1.9549
revealsPerMove
1.0722
censored
0
illegal
0
incomplete
0
wallSeconds
265.4953
fair-d4s7
games
64
mean
394984.3125
median
302,556
q25
198524.7500
max
1,663,637
movesMean
114.2969
clearsPerMove
2.0547
revealsPerMove
1.1494
censored
0
illegal
0
incomplete
0
wallSeconds
11302.7039
paired
candidate-vs-fair-d3s7
meanDelta
-106964.3750
bootstrapLower95
-146579.5391
bootstrapUpper95
-69983.3727
studentTLower95
-145890.1465
pairedSd
186537.5328
detectionFloor
38356.7802
wtl
  1. 14
  2. 0
  3. 50
halves
  1. -109138.6563
  2. -104790.0938
q25Delta
-36,134
movesDelta
-29.5313
init-vs-fair-d3s7
meanDelta
-142339.6250
bootstrapLower95
-178491.5461
bootstrapUpper95
-107770.6313
studentTLower95
-178510.6720
pairedSd
173336.5226
detectionFloor
35642.3225
wtl
  1. 9
  2. 0
  3. 55
halves
  1. -156142.6563
  2. -128536.5938
q25Delta
-50,964
movesDelta
-39.2813
candidate-vs-init
meanDelta
35375.2500
bootstrapLower95
16898.9273
bootstrapUpper95
54629.2750
studentTLower95
16144.9914
pairedSd
92153.9860
detectionFloor
18949.1634
wtl
  1. 36
  2. 0
  3. 28
halves
  1. 47,004
  2. 23746.5000
q25Delta
14,830
movesDelta
9.7500
fair-d4s7-vs-fair-d3s7
meanDelta
97058.7188
bootstrapLower95
29329.5438
bootstrapUpper95
165453.2305
studentTLower95
27011.3702
pairedSd
335676.3163
detectionFloor
69023.4425
wtl
  1. 41
  2. 0
  3. 23
halves
  1. 75457.7813
  2. 118659.6563
q25Delta
40137.5000
movesDelta
26.2656
training
generationsCompleted
60
tenGenerationBlockAverages
  1. generations
    0-9
    populationMean
    159340.4412
    top4Mean
    176605.0742
    best
    183544.7219
    meanMinusInit
    4682.4068
    bestMinusInit
    28886.6875
    meanMinusFair
    -172846.5150
    bestMinusFair
    -148642.2344
  2. generations
    10-19
    populationMean
    168372.4929
    top4Mean
    189184.1141
    best
    194671.8250
    meanMinusInit
    15852.6179
    bestMinusInit
    42151.9500
    meanMinusFair
    -154441.2915
    bestMinusFair
    -128141.9594
  3. generations
    20-29
    populationMean
    174464.3515
    top4Mean
    197282.4398
    best
    203909.7531
    meanMinusInit
    22703.5952
    bestMinusInit
    52148.9969
    meanMinusFair
    -155743.1267
    bestMinusFair
    -126297.7250
  4. generations
    30-39
    populationMean
    181219.7211
    top4Mean
    203060.9922
    best
    209035.9844
    meanMinusInit
    32277.6086
    bestMinusInit
    60093.8719
    meanMinusFair
    -125175.9352
    bestMinusFair
    -97359.6719
  5. generations
    40-49
    populationMean
    184679.3310
    top4Mean
    206650.3125
    best
    213143.7469
    meanMinusInit
    34036.9685
    bestMinusInit
    62501.3844
    meanMinusFair
    -160026.6253
    bestMinusFair
    -131562.2094
  6. generations
    50-59
    populationMean
    190907.8234
    top4Mean
    216015.6836
    best
    227155.7656
    meanMinusInit
    40431.1578
    bestMinusInit
    76679.1000
    meanMinusFair
    -121896.0891
    bestMinusFair
    -85648.1469
trainingSignalCheck
rule
population mean fitness > paired fair-d3s7 control in the majority of the final 10 completed generations (theory falsifier 2)
window
  1. 50
  2. 51
  3. 52
  4. 53
  5. 54
  6. 55
  7. 56
  8. 57
  9. 58
  10. 59
generationsMeanAboveFair
0
generationsTop4AboveFair
0
generationsBestAboveFair
0
generationsInitAboveFair
0
meanMarginLast10
-121896.0891
top4MarginLast10
-96788.2289
passed
false
artifactIntegrity
generationArtifacts
60
illegalDecisions
0
incompleteDecisions
0
censoredGames
0
config
population
32
games
32
generations
60
elites
4
tournament
3
sigmaRel
0.0500
seed
0x0e701e58
leaseStart
0xa52e0500
moveCap
2,000
wallSeconds
21,600
selection
finalGeneration
59
finalists
  1. 12
  2. 1
  3. 6
  4. 29
  5. 30
  6. 23
  7. 13
  8. 4
blockStart
0xa52e0c80
selectGames
128
winner
candidate-29
winnerMean
202236.7734
corpus
games
177
roots
21,618
censoredGames
3
score
n
177
mean
433131.1525
sd
323856.2717
min
103,368
q25
229,522
median
348,847
q75
549,029
max
1,889,608
moves
n
177
mean
122.1356
sd
85.7326
min
34
q25
70
median
100
q75
150
max
500
wallSecondsPerGame
n
177
mean
9461.8629
sd
7988.6131
min
1209.2638
q25
4178.7932
median
7476.2450
q75
11333.4374
max
49497.1916
secondsPerRoot
77.4702
rootsPerGame
122.1356
legalColumnsPerRoot
n
21,618
mean
6.8572
sd
0.6843
min
1
q25
7
median
7
q75
7
max
7
teacherSiblingSpreadPoints
n
21,618
mean
40559.6546
sd
160027.6823
min
0
q25
2448.7576
median
4187.9510
q75
7480.5930
max
1017575.6843
pretrain
report
roots
21,618
games
177
trainRoots
17,641
valRoots
3,977
epochs
16
batch
64
lr
0.0003
seed
0x0e701e57
bestEpoch
8
bestValHuber
0.6873
probe
probeRoots
256
top1
0.4414
meanTeacherValueRegret
2394.6708
epochs
  1. epoch
    0
    trainHuber
    2.3016
    valHuber
    1.8859
    valPearson
    0.8790
  2. epoch
    1
    trainHuber
    1.2819
    valHuber
    0.9546
    valPearson
    0.9314
  3. epoch
    2
    trainHuber
    0.7173
    valHuber
    0.8128
    valPearson
    0.9449
  4. epoch
    3
    trainHuber
    0.5488
    valHuber
    0.7449
    valPearson
    0.9491
  5. epoch
    4
    trainHuber
    0.4514
    valHuber
    0.7100
    valPearson
    0.9512
  6. epoch
    5
    trainHuber
    0.3886
    valHuber
    0.6918
    valPearson
    0.9520
  7. epoch
    6
    trainHuber
    0.3371
    valHuber
    0.6906
    valPearson
    0.9519
  8. epoch
    7
    trainHuber
    0.3036
    valHuber
    0.6919
    valPearson
    0.9500
  9. epoch
    8
    trainHuber
    0.2706
    valHuber
    0.6873
    valPearson
    0.9504
  10. epoch
    9
    trainHuber
    0.2443
    valHuber
    0.6939
    valPearson
    0.9502
  11. epoch
    10
    trainHuber
    0.2246
    valHuber
    0.6993
    valPearson
    0.9479
  12. epoch
    11
    trainHuber
    0.2092
    valHuber
    0.6955
    valPearson
    0.9496
  13. epoch
    12
    trainHuber
    0.1915
    valHuber
    0.7119
    valPearson
    0.9480
  14. epoch
    13
    trainHuber
    0.1807
    valHuber
    0.6936
    valPearson
    0.9479
  15. epoch
    14
    trainHuber
    0.1729
    valHuber
    0.6995
    valPearson
    0.9485
  16. epoch
    15
    trainHuber
    0.1620
    valHuber
    0.6969
    valPearson
    0.9470
resources
  1. label
    corpus
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/teacher_corpus
    2. --seeds-start
    3. 0xa52e0300
    4. --games
    5. 512
    6. --threads
    7. 32
    8. --depth
    9. 5
    10. --move-cap
    11. 500
    12. --wall-seconds
    13. 46800
    14. --out
    15. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/corpus
    startedAt
    2026-09-02T04:24:10Z
    endedAt
    2026-09-02T22:39:58Z
    wallSeconds
    65747.5370
    userSeconds
    1667946.0280
    systemSeconds
    641.5270
    peakRssBytes
    1,615,892,480
    exitCode
    0
  2. label
    pretrain
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/pretrain
    2. --corpus
    3. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/corpus/corpus.jsonl
    4. --out
    5. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain
    6. --epochs
    7. 16
    8. --batch
    9. 64
    10. --lr
    11. 3e-4
    12. --seed
    13. 0x0e701e57
    14. --probe-roots
    15. 256
    16. --threads
    17. 32
    startedAt
    2026-09-02T22:42:11Z
    endedAt
    2026-09-02T22:42:25Z
    wallSeconds
    13.9780
    userSeconds
    45.2200
    systemSeconds
    0.1080
    peakRssBytes
    265,060,352
    exitCode
    0
  3. label
    evolve
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
    2. --init
    3. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain/init.bin
    4. --lease-start
    5. 0xa52e0500
    6. --out
    7. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve
    8. --population
    9. 32
    10. --games
    11. 32
    12. --generations
    13. 60
    14. --elites
    15. 4
    16. --tournament
    17. 3
    18. --sigma-rel
    19. 0.05
    20. --seed
    21. 0x0e701e58
    22. --threads
    23. 32
    24. --wall-seconds
    25. 21600
    26. --move-cap
    27. 2000
    startedAt
    2026-09-02T22:42:25Z
    endedAt
    2026-09-03T02:37:19Z
    wallSeconds
    14093.6190
    userSeconds
    441673.1460
    systemSeconds
    184.6580
    peakRssBytes
    477,446,144
    exitCode
    0
  4. label
    select
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
    2. --select
    3. --lease-start
    4. 0xa52e0500
    5. --out
    6. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve
    7. --generations
    8. 60
    9. --games
    10. 32
    11. --select-games
    12. 128
    13. --threads
    14. 32
    15. --move-cap
    16. 2000
    startedAt
    2026-09-03T02:37:19Z
    endedAt
    2026-09-03T02:41:29Z
    wallSeconds
    250.4040
    userSeconds
    7852.4190
    systemSeconds
    3.9940
    peakRssBytes
    252,526,592
    exitCode
    0
  5. label
    screen
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/screen
    2. --candidate
    3. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve/candidate-weights.bin
    4. --init
    5. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain/init.bin
    6. --seeds-start
    7. 0xa52e1300
    8. --games
    9. 64
    10. --threads
    11. 32
    12. --move-cap
    13. 2000
    14. --out
    15. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/screen/heldout.json
    startedAt
    2026-09-03T02:41:29Z
    endedAt
    2026-09-03T02:52:22Z
    wallSeconds
    652.9870
    userSeconds
    12346.6780
    systemSeconds
    3.8980
    peakRssBytes
    1,792,266,240
    exitCode
    0
Limitations
  • Single 64-game held-out screen: the paired detection floor is about 38,000 points for the primary contrast and 19,000 for the ablation; the primary result is far outside its floor, the ablation clears its own.
  • The teacher corpus reached 177 of the 512 games the protocol allowed for because the depth-5 teacher ran four to seven times slower than the pilot projected; the protocol makes the completed games the corpus, so the result is valid, but it rejects this configuration at this corpus size, not the design at 512 games.
  • The supervised warm start plays at about half the fair leaf's level; the deployment-faithful probe (0.441 top-1) and the leaf-swing diagnostic (zero leaf 0.450 on a different sample) suggest imitation of state values barely improved the search's ordering over no leaf at all. Evolution then had roughly 150,000 paired points to make up in 60 generations and made up about 35,000 to 40,000.
  • The per-game artifact holds four arms x 64 games (256 rows); the primary contrast pairs the candidate and fair-d3s7 rows by seed.
  • The fair-d4s7 arm is diagnostic only and reproduces the standing depth-4-over-depth-3 result on these seeds; it is not part of the gate.
  • Two operational faults during the unattended chain (a false liveness reading, then a script replacement that crashed the corpus stage's shell after its binary had exited 0) are recorded in the run record; the 'corpus: done' marker in pipeline.log was appended by hand with a note. No artifact was affected.
  • The elite re-selection block (0xa52e0c80, 128 games) and every fitness block are training-lease seeds; the last generation's 32-game leaders read 20,000 to 60,000 above their 128-game re-selection means, which is the best-of-32 bias the re-selection exists to remove.

Open the result record

Agent contextHow to extend this record

To add a reader-facing explanation, write web/content/research/EX-20260902-nnue-evolution-d3-v2-49c18bc2.mdx; it renders above this record on the next request. The registered protocol itself is in the technical record above.

Record file: research/experiments/EX-20260902-nnue-evolution-d3-v2-49c18bc2.json, validated against research/schemas/experiment-v1.schema.json. Protocol hash: 66d163bb2d39cf4da157a1044be766eb124af28c158a7750e2064ad61edda942.

Registered by Claude Code / claude-fable-5-1 (claude-code-evolution-experiment).