Second continuation of the depth-5-distilled NNUE leaf evolution: resume from the plateau population with a slower mutation-size decay (up to 1,000 further generations)
Second successor experiment to EX-20260903-nnue-evolution-continuation-d3-f8ce9181 (result RS-20260903T163321Z-733076b5: 150 generations from the first run's generation-60 population, stopped by the preregistered plateau rule after generation 149 — the fitted slope of the paired margin over the fair control across generations 50-149 was +37 points/generation, one-sided 95% lower bound -159, no detectable improvement — after rising +239/generation, lower bound +30, across generations 0-99.
On this page
- Created
- Updated
No explanation has been written for this record yet.
Technical recordThe registered protocol
- Hypothesis
- Second successor experiment to EX-20260903-nnue-evolution-continuation-d3-f8ce9181 (result RS-20260903T163321Z-733076b5: 150 generations from the first run's generation-60 population, stopped by the preregistered plateau rule after generation 149 — the fitted slope of the paired margin over the fair control across generations 50-149 was +37 points/generation, one-sided 95% lower bound -159, no detectable improvement — after rising +239/generation, lower bound +30, across generations 0-99. The frozen candidate lost its 64-game held-out screen to the fair leaf by -68,441 paired points but beat the prior run's candidate by +36,278 out of sample). At the moment the rule fired, sigma_rel had annealed from 0.05 to 0.0345 (time constant 400 generations, floor 0.01), so the design could not separate two explanations: the leaf class had exhausted the improvement available to this evolutionary process, or the shrinking step size had simply become too small to find further gains before the rule's 100-generation window could detect them. This experiment tests the second explanation directly by resuming the identical process from the identical population with a much slower decay, so sigma stays close to its starting value for far longer: if the population still cannot separate from the fair control at the earlier rate under room to explore, the plateau is a property of the leaf, not the schedule. Candidate: the identical deployed configuration (stock fair expectimax, depth 3, seven chance strata, terminal utility -1,000,000, policy seed 0xd7075eed, completion-guaranteeing work bound, 64k-entry table, NNUE leaf 8,902-feature sparse class, 135 active, EmbeddingBag(8902,64)->ReLU->32->ReLU->1, x17,000 points). Starting population: the 32 vectors of runs/RUN-20260903T032832Z-a76a6cf7/nnue-evolution/evolve/population-150.bin, the population the second experiment's generation-149 selection produced and never evaluated (SHA-256 33dcc25b88ab56a0a80329bf11983b539883aa922900b03de3cea71e1da3529d); generation 0 of this run plays it exactly as saved. GA per generation, unchanged from both prior runs except where stated: population 32; every candidate plays the same fresh block of 32 training seeds, generation g playing [0xa52ea100 + 32 g, +32); fitness = mean final score of the 32 complete games (2,000-move cap); 4 elites cloned; the other 28 slots filled by tournaments of 3 whose winner is copied and mutated per tensor with Gaussian noise; evolution seed 0x0e701e5a (fresh, so no mutation realisation of either prior run repeats). Slower-decaying mutation (the change under test): sigma_rel(g) = max(0.01, 0.05 x exp(-g / 1500)), i.e. the same starting value and floor as the first continuation but a time constant 3.75x longer: 0.0500 at g=0 (identical to both prior runs' start), 0.0452 at g=150 (against 0.0345 at the equivalent point last time), 0.0358 at g=500, 0.0257 at g=1,000, and the 0.01 floor not reached until g~2,414, i.e. never within this experiment's 1,000-generation cap. Three paired controls play every block and are never selected: the frozen fair leaf (control-fair-d3s7), the first run's supervised warm start (7d6fa5563de6...; control-init-d3s7) and the SECOND experiment's frozen candidate (artifacts/results/EX-20260903-nnue-evolution-continuation-d3-f8ce9181/RUN-20260903T032832Z-a76a6cf7/evolve/candidate-weights.bin, SHA-256 759084fab97599818d01a5da6d18b9cc77172249f09a68d5104361bca385e09b; control-baseline-d3s7) — the immediately-prior best, so the paired margin over it is the direct measure of whether this run gets past where the last one stopped. Preregistered plateau stop, identical rule and parameters to the second experiment for direct comparability: after every 50th completed generation from the 100th, OLS regresses the paired margin (population mean fitness minus the fair control) over generation index across the last 100 generations; the run stops when the slope's one-sided 95% lower bound (slope - 1.645 x standard error, residual-based, n-2 degrees of freedom) is not greater than zero. Sensitivity, from the second experiment's own artifacts: residual scatter of the tracked margin was 34,221 points, at which a 100-generation window detects slopes of about 195 points per generation (19,500 per 100 generations); this experiment's own last-100-generation slope before its stop will be reported for direct comparison against the prior +37 and +239 readings. Other stops: 1,000 generations; a 259,200 s (72 h) evolve wall budget (pinned once at first launch per the current evolve.rs, cumulative across any resume); a STOP file. Final candidate: the top 8 of the last completed generation each replay the 128 fresh training seeds immediately after the last fitness block played; the highest mean is frozen as candidate-weights.bin and its SHA-256 recorded before the screen lease opens. Held-out screen, once: 64 never-read public-development seeds 0xa52f2100-0xa52f213f, five arms on identical seeds: candidate, baseline-run2 (the second experiment's frozen candidate, i.e. this run's own baseline control), init-d3s7 (the first run's warm start), fair-d3s7 (the primary comparator) and fair-d4s7 (the program's standing reference, diagnostic only). Primary contrast candidate minus fair-d3s7 as in both prior screens; preregistered secondary contrast candidate minus baseline-run2, the out-of-sample measure of whether the slower decay found anything the faster one could not. Operational: 32 threads, every stage through scripts/pipeline.sh (env: EXPERIMENT_ID, LEASE_START 0xa52ea100, SCREEN_START 0xa52f2100, GENERATIONS 1000, EVOLVE_SEED 0x0e701e5a, RESUME_POP, BASELINE, INIT, SIGMA_TAU 1500, SIGMA_FLOOR 0.01, PLATEAU_WINDOW 100, PLATEAU_EVERY 50, PLATEAU_MIN 100, EVOLVE_WALL 259200) with per-stage rusage; binaries rebuilt from the current branch, which includes the owner's 2026-09-03 fix (wall-budget.json pinning and progress/plateau recovery from committed generation artifacts across a resume) — this run is not expected to need that recovery path since it starts fresh, but inherits it for safety.
- Arms
Arm Name Entry point Manifest Candidate d3s7-evolved-nnue-leaf-continued2 approaches/lifetime-objective/nnue-evolution/src/bin/evolve.rs– Comparator fair-d3s7 approaches/fair-expectimax/rust-engine/src/leaf.rsresearch/benchmarks/baselines-v1.json- Classification
- algorithmic
- Information boundary
- public-policy
- Benchmark tier
- SCREEN
- Lifecycle
- completed
- Theories tested
- Primary metric
- paired mean whole-game score delta, second-continuation candidate minus frozen fair leaf, both at the identical d3s7 configuration, 64 held-out games
- Secondary metrics
- paired mean score delta, second-continuation candidate minus baseline-run2 (the second experiment's frozen candidate), same 64 games: whether the slower decay found improvement the faster one could not
- paired mean moves delta; numbered clears per move and cover reveals per move; lower quartile, median, maximum; W-T-L; first-half and second-half deltas, for every arm pair
- training curve: per-generation population mean, top-4 mean, best, and the three paired controls; the plateau statistic and every plateau check (slope, standard error, lower bound); the generation at which the run stopped and why, compared against the second experiment's own checks (+239/gen then +37/gen)
- fifty- and ten-generation block averages of population mean minus each control, continuing both prior runs' tables
- reference arm fair-d4s7 vs fair-d3s7, diagnostic only
- Statistical unit
- whole-game
- Uncertainty method
- one-sided 95% percentile bootstrap over whole games, 20,000 resamples, RNG seed 0xb0071eaf (the unchanged compare.py), plus a one-sided 95% Student-t lower bound; detection floor 1.645*sd/sqrt(n) reported; plateau rule as stated in the hypothesis
- Data role
- public-development
- Seed leases
SL-20260903T190000Z-a52ea100SL-20260903T190000Z-a52f2100
- Whole-origin split
- yes
- Reuse disclosure
- The starting population and both controls are products of EX-20260902-nnue-evolution-d3-v2-49c18bc2's and EX-20260903-nnue-evolution-continuation-d3-f8ce9181's training data, already read; no seed of either prior lease is read again. CHECK gates for the rebuilt binaries read only the already-opened development probe block 0xa5277000-0xa5277107. Fitness blocks and the re-selection open training-lease seeds 0xa52ea100-0xa52f2100 (1,000 blocks of 32 = 32,000 seeds plus 128 re-selection seeds; the remainder above 32,128 stays unopened); the screen opens 0xa52f2100-0xa52f213f exactly once, after the candidate is frozen. Neither range contains any 8-hex-digit constant present in docs/, approaches/, src/, research/, artifacts/ or web/content/ (checked 2026-09-03), and both are disjoint from every existing lease.
- Pass criteria
- All 11 CHECK gates passed on the rebuilt binaries before the first leased seed was read (gates.log in the run directory).
- Every generation artifact and the screen artifact have illegalDecisions 0, incompleteDecisions 0, and censored games reported as censored.
- Held-out screen, 64 paired games on 0xa52f2100-0xa52f213f, candidate vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0.
- Held-out screen: paired mean score delta > 0 in both halves (seeds 0-31 and 32-63).
- Held-out screen: candidate lower-quartile score >= fair-d3s7 lower-quartile score.
- The screened candidate is the exact frozen winner of the elite re-selection (candidate-weights.bin, SHA-256 recorded before the screen lease opens); no other vector is screened.
- Reported alongside the gate, not part of it: candidate vs baseline-run2 bootstrap 95% lower bound > 0 establishes that the slower decay found improvement the faster one had not; the generation of the stop and its plateau slope, compared against the prior +239/gen and +37/gen readings, establish whether more room to explore changed the outcome.
- On pass
- Freeze candidate-weights.bin with its SHA-256; register a fresh-block replication by a different runner and a fresh-development STANDARD evaluation as successors; do not open protected or final seeds.
- On fail
- Record a valid run with scientific outcome fail for this configuration (this model class, this population, this slower annealing schedule, deployment depth 3), reporting separately whether the slower decay improved on the second experiment's candidate and at what generation (if any) this run's own curve plateaued; do not adopt the vector; open no further cohort for it without a new theory-level decision.
- Gate fixed before controlled data
- yes
- Resources
Wall seconds 275000 CPU threads 32 Max host bytes 17179869184 Max GPU bytes – GPU devices – - Stop conditions
- Stop on any rules, information-boundary, legality, determinism, or parity failure.
- Evolution stops at the preregistered plateau rule (one-sided 95% lower bound of the OLS slope of the paired margin over the fair control across the last 100 generations <= 0, checked after every 50th completed generation from the 100th), at 1,000 generations, at 259,200 s of evolve wall time, or when a STOP file appears; the candidate is then the elite re-selection over the last completed generation's population on the 128 seeds following the last fitness block.
- A generation or screen artifact with any illegal or incomplete decision voids the run (invalid), not the candidate.
- The screen is evaluated exactly once; no re-run on the same or a different held-out block without a new experiment record.
- The run may be killed and resumed at a generation boundary (checkpoint contract, with the current evolve.rs's progress/plateau recovery and pinned wall budget); a resume replays no seed and is recorded in the run record.
- Expected artifacts
runs/<run-id>/nnue-evolution/gates.log (CHECK, rebuilt binaries)runs/<run-id>/nnue-evolution/evolve/{config.json (with resumePopulationSha256 and baselineSha256),progress.jsonl,plateau.jsonl,wall-budget.json,PLATEAU (if the rule stopped the run),gen-*.json,population-*.bin,final-fitness.json,candidate-weights.bin,candidate-weights.sha256,selection.json}runs/<run-id>/nnue-evolution/screen/heldout.json (five arms) + compare-{candidate-vs-fair-d3s7,candidate-vs-baseline-run2,baseline-run2-vs-fair-d3s7,init-vs-fair-d3s7,candidate-vs-init,fair-d4s7-vs-fair-d3s7}.jsonruns/<run-id>/nnue-evolution/{pipeline.log,rusage.jsonl,evolve.log,select.log,screen.log,analysis.json,analysis.md}artifacts emitted by evolve and screen carry this experiment id in their config field
- Amendments
Timestamp Before controlled data Reason 2026-09-04T09:06:20Z no lifecycle advanced preregistered -> completed after run RUN-20260903T190433Z-a87fd7fc and result RS-20260904T090620Z-e5731bf0 were written; protocol content otherwise unchanged. The preregistration hash d1a8f687327b382459096fd5db0bf45c201f7147b8b4b30a707ebd00cfa1889c is retained in the run and lease records.
Technical recordResults recorded against this protocol
Valid run of the frozen protocol EX-20260903-nnue-evolution-continuation2-d3-80eebad3 (second successor to the theory, third experiment in the series), every stage completed, every artifact with zero illegal and zero incomplete decisions. Stage C resumed from the second experiment's checkpointed generation-150 population (SHA-256 33dcc25b88ab56a0a80329bf11983b539883aa922900b03de3cea71e1da3529d) with that experiment's frozen candidate (SHA-256 759084fab97599818d01a5da6d18b9cc77172249f09a68d5104361bca385e09b) as the third paired control, testing whether a 3.75x slower mutation-size decay (time constant 1,500 generations vs. 400 before; sigma stayed at 0.0500 at generation 0, 0.0453 at the stop, versus 0.0345 at the equivalent generation last time) would let evolution get further past the point where the prior run plateaued. The identical preregistered plateau rule stopped this run at the identical generation as before, 149: the check after generation 99 found the paired margin over the fair control rising at +201.7 points per generation (lower bound +11.7, versus +239 and +30 in the second experiment) and continued; the check after generation 149 found +164.0 per generation (lower bound -28.5, versus +37 and -159 before) and stopped. The improvement rate at the stopping point was more than four times the prior run's (164 against 37 points/generation), and the best candidate beat the fair control on 26 of 150 blocks against 9 in the second experiment, but the lower bound of the fitted slope still crossed zero at the same 100-generation checkpoint, so the rule fired regardless. Fifty-generation population-mean averages: 235,353, 238,797, 244,800 (rising monotonically, unlike the second experiment's flat third block); paired margin over the immediately-prior candidate +4,301, -949, +13,589; over the fair control -99,479, -88,006, -77,654. Training-signal falsifier fails again (population mean above fair in 0 of the last 10 generations, margin -70,199; best above fair in 4 of the last 10). Elite re-selection on 128 fresh games (0xa52eb3c0) froze candidate-28 at 257,314 (finalists 232,499 to 257,314, a tighter spread than the second experiment's). Stage D, 64 never-read games (0xa52f2100, opened once at 2026-09-04T08:56:37Z after the candidate's SHA-256 was recorded), five arms: this candidate averaged 249,757 against the frozen fair leaf's 335,266, paired -85,509 (bootstrap 95% bounds -128,482 to -43,591, Student-t lower bound -129,423, floor 43,272), W-T-L 22-0-42, both halves negative (-59,644 / -111,374), lower quartile 156,534 against 176,108: every screen criterion fails, scientific outcome fail for this configuration, consistent with both prior screens. The preregistered secondary contrast is the key negative finding of this run: the candidate beat the immediately-prior (second experiment's) frozen candidate by only +13,572 on the same seeds, with a bootstrap 95% lower bound of -30,165 and a Student-t lower bound of -30,869 -- both crossing zero (detection floor 43,791), W-T-L 31-0-33. Unlike the second experiment's clearly positive +36,278 (lower bound +9,085) over the first experiment's candidate, this third experiment's gain over the second is NOT statistically distinguishable from zero at this screen size: three months of relatively larger mutations bought a point estimate about a third the size of the previous continuation's out-of-sample gain, and it is not confidently positive. The second experiment's candidate itself scored -99,081 against the fair leaf on this fresh block (lower bound -154,094), broadly consistent with its own screen result of -68,441 (lower bound -112,090), a second informal replication. Over the warm start the candidate is +106,846 (lower bound +76,970). The reference arm gave fair depth 4 minus fair depth 3 of -1,409 (bounds -62,290 to +58,933, crossing zero on this cohort, a diagnostic-only reading and not comparable across screens with different fresh seeds). Read together with the second experiment: slowing the mutation-size decay produced a visibly healthier training curve (a still-rising population mean through all three 50-generation blocks, more generations where the best candidate beat the fair control) but did NOT produce a statistically confirmed improvement over the immediately-prior candidate on held-out games, and the plateau rule still stopped the run at the same generation. The most defensible reading is that the annealing schedule was not the dominant cause of the earlier plateau -- something else (population size, games per candidate, or the leaf class's genuine ceiling under this search depth) is the binding constraint, though a schedule 3.75x slower is not proof that no schedule would help; a much slower schedule, or removing the anneal-driven exploration decay entirely in favour of a fixed sigma with more games per candidate, remains untested.
- ✓All 11 CHECK gates passed on the rebuilt binaries before the first leased seed was read — observed: ALL GATES PASSED (gates.log), 13/13 unit tests, smoke test of resume/tau/baseline/screen-arm flags on the already-open probe block
- ✓Every generation artifact and the screen artifact: illegalDecisions 0, incompleteDecisions 0, censored games reported — observed: 150 generation artifacts: 0 illegal, 0 incomplete, 0 censored; screen: 0 illegal, 0 incomplete, 0 censored in all five arms
- ✕Held-out screen, candidate vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0 — observed: delta -85,509; bootstrap LB -128,482, UB -43,591; t LB -129,423; paired sd 210,440; floor 43,272; W-T-L 22-0-42
- ✕Held-out screen: paired mean score delta > 0 in both halves — observed: first half -59,644, second half -111,374
- ✕Held-out screen: candidate lower-quartile score >= fair-d3s7 lower-quartile score — observed: candidate Q25 156,534 vs fair-d3s7 Q25 176,108 (delta -19,573)
- ✓The screened candidate is the exact frozen winner of the elite re-selection; SHA-256 recorded before the screen lease opened — observed: candidate-28 of generation 149, mean 257,314 over 128 fresh games at 0xa52eb3c0; sha256 5318e70e8e1c2bc5a872d4b61a472d657b96d0a6bf50875a3a325403de682f2a; lease opened 2026-09-04T08:56:37Z by the screen stage
- ✕Reported alongside the gate: candidate vs baseline-run1 (the immediately-prior candidate) bootstrap 95% lower bound > 0 (the slower decay found improvement the faster one had not) — observed: delta +13,572; bootstrap LB -30,165, UB +57,575; t LB -30,869; floor 43,791; W-T-L 31-0-33; NOT confidently positive, unlike the prior continuation's +36,278 (LB +9,085) over its own predecessor
- –Reported alongside the gate: the generation of the stop and the plateau statistics, compared against the second experiment's own checks — observed: stopped by the plateau rule after generation 149 (150 completed), the SAME generation as the second experiment despite the 3.75x slower decay; check after generation 99: slope +201.7/gen (second experiment: +239) lower bound +11.7 (continue); after generation 149: slope +164.0/gen (second experiment: +37) lower bound -28.5 (stop); best-above-fair on 26/150 blocks (second experiment: 9/150)
Technical recordRecorded metrics
- seedsStartHex
- 0xa52f2100
- games
- 64
- moveCap
- 2,000
- role
- public-development, never read before this screen (lease SL-20260903T190000Z-a52f2100, opened once)
- candidate
- games
- 64
- mean
- 249756.6875
- median
- 192,147
- q25
- 156534.5000
- max
- 710,297
- movesMean
- 75.0938
- clearsPerMove
- 1.8658
- revealsPerMove
- 1.0121
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 566.6732
- baseline-run1
- games
- 64
- mean
- 236184.1875
- median
- 209822.5000
- q25
- 141397.7500
- max
- 660,003
- movesMean
- 70.7813
- clearsPerMove
- 1.8567
- revealsPerMove
- 1.0188
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 525.1998
- init-d3s7
- games
- 64
- mean
- 142910.7656
- median
- 124023.5000
- q25
- 106530.2500
- max
- 269,984
- movesMean
- 45.2500
- clearsPerMove
- 1.4914
- revealsPerMove
- 0.7117
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 314.6850
- fair-d3s7
- games
- 64
- mean
- 335265.6406
- median
- 272317.5000
- q25
- 176107.5000
- max
- 1,080,133
- movesMean
- 98.2188
- clearsPerMove
- 2.0051
- revealsPerMove
- 1.1195
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 273.8052
- fair-d4s7
- games
- 64
- mean
- 333857.0625
- median
- 285,125
- q25
- 194039.2500
- max
- 800,130
- movesMean
- 97.2344
- clearsPerMove
- 1.9994
- revealsPerMove
- 1.1093
- censored
- 0
- illegal
- 0
- incomplete
- 0
- wallSeconds
- 9576.0764
- candidate-vs-fair-d3s7
- meanDelta
- -85508.9531
- bootstrapLower95
- -128482.1328
- bootstrapUpper95
- -43591.2375
- studentTLower95
- -129422.5221
- pairedSd
- 210439.7295
- detectionFloor
- 43271.6694
- wtl
- 22
- 0
- 42
- halves
- -59643.7500
- -111374.1563
- q25Delta
- -19,573
- movesDelta
- -23.1250
- init-vs-fair-d3s7
- meanDelta
- -192354.8750
- bootstrapLower95
- -234603.4492
- bootstrapUpper95
- -152044.4570
- studentTLower95
- -234770.0951
- pairedSd
- 203259.4400
- detectionFloor
- 41795.2224
- wtl
- 6
- 0
- 58
- halves
- -171614.0313
- -213095.7188
- q25Delta
- -69577.2500
- movesDelta
- -52.9688
- candidate-vs-init
- meanDelta
- 106845.9219
- bootstrapLower95
- 76970.4852
- bootstrapUpper95
- 138415.2789
- studentTLower95
- 75526.2930
- pairedSd
- 150087.8742
- detectionFloor
- 30861.8191
- wtl
- 53
- 0
- 11
- halves
- 111970.2813
- 101721.5625
- q25Delta
- 50004.2500
- movesDelta
- 29.8438
- fair-d4s7-vs-fair-d3s7
- meanDelta
- -1408.5781
- bootstrapLower95
- -62290.3523
- bootstrapUpper95
- 58933.4102
- studentTLower95
- -63294.7565
- pairedSd
- 296566.8913
- detectionFloor
- 60981.5670
- wtl
- 33
- 0
- 31
- halves
- 32695.7813
- -35512.9375
- q25Delta
- 17931.7500
- movesDelta
- -0.9844
- candidate-vs-baseline-run1
- meanDelta
- 13572.5000
- bootstrapLower95
- -30164.9336
- bootstrapUpper95
- 57575.2000
- studentTLower95
- -30868.5694
- pairedSd
- 212967.5824
- detectionFloor
- 43791.4591
- wtl
- 31
- 0
- 33
- halves
- 2355.4688
- 24789.5313
- q25Delta
- 15136.7500
- movesDelta
- 4.3125
- baseline-run1-vs-fair-d3s7
- meanDelta
- -99081.4531
- bootstrapLower95
- -154094.1469
- bootstrapUpper95
- -44893.0578
- studentTLower95
- -154662.2389
- pairedSd
- 266350.6017
- detectionFloor
- 54768.3425
- wtl
- 24
- 0
- 40
- halves
- -61999.2188
- -136163.6875
- q25Delta
- -34709.7500
- movesDelta
- -27.4375
- generationsCompleted
- 150
- stoppedOnPlateau
- true
- plateauChecks
- generation
- 99
- window
- 100
- slopePerGeneration
- 201.7241
- standardError
- 115.5014
- lowerBound95
- 11.7413
- windowMeanFirstHalf
- -99479.2672
- windowMeanSecondHalf
- -88005.6168
- stop
- false
- generation
- 149
- window
- 100
- slopePerGeneration
- 164.0170
- standardError
- 117.0542
- lowerBound95
- -28.5201
- windowMeanFirstHalf
- -88005.6168
- windowMeanSecondHalf
- -77654.0195
- stop
- true
- fiftyGenerationBlockAverages
- generations
- 0-49
- populationMean
- 235352.7134
- meanMinusFair
- -99479.2672
- meanMinusBaseline
- 4300.7409
- sigmaAtEnd
- 0.0484
- generations
- 50-99
- populationMean
- 238796.7395
- meanMinusFair
- -88005.6168
- meanMinusBaseline
- -948.5393
- sigmaAtEnd
- 0.0468
- generations
- 100-149
- populationMean
- 244799.5112
- meanMinusFair
- -77654.0195
- meanMinusBaseline
- 13589.1018
- sigmaAtEnd
- 0.0453
- bestAboveFairGenerations
- 1
- 20
- 34
- 38
- 40
- 45
- 52
- 53
- 71
- 74
- 77
- 88
- 91
- 97
- 100
- 103
- 107
- 108
- 113
- 121
- 124
- 130
- 140
- 142
- 146
- 147
- trainingSignalCheck
- rule
- population mean fitness > paired fair-d3s7 control in the majority of the final 10 completed generations (theory falsifier 2)
- window
- 140
- 141
- 142
- 143
- 144
- 145
- 146
- 147
- 148
- 149
- generationsMeanAboveFair
- 0
- generationsTop4AboveFair
- 2
- generationsBestAboveFair
- 4
- generationsInitAboveFair
- 0
- meanMarginLast10
- -70199.2831
- top4MarginLast10
- -28764.4430
- passed
- false
- artifactIntegrity
- generationArtifacts
- 150
- illegalDecisions
- 0
- incompleteDecisions
- 0
- censoredGames
- 0
- config
- experiment
- EX-20260903-nnue-evolution-continuation2-d3-80eebad3
- population
- 32
- games
- 32
- generations
- 1,000
- elites
- 4
- tournament
- 3
- sigmaRel
- 0.0500
- sigmaDecayTau
- 1,500
- sigmaFloor
- 0.0100
- seed
- 0x0e701e5a
- leaseStart
- 0xa52ea100
- moveCap
- 2,000
- wallSeconds
- 259,200
- resumePopulation
- /home/keshav/Developer/drop7-bench/runs/RUN-20260903T032832Z-a76a6cf7/nnue-evolution/evolve/population-150.bin
- resumePopulationSha256
- 33dcc25b88ab56a0a80329bf11983b539883aa922900b03de3cea71e1da3529d
- baseline
- /home/keshav/Developer/drop7-bench/artifacts/results/EX-20260903-nnue-evolution-continuation-d3-f8ce9181/RUN-20260903T032832Z-a76a6cf7/evolve/candidate-weights.bin
- baselineSha256
- 759084fab97599818d01a5da6d18b9cc77172249f09a68d5104361bca385e09b
- plateauWindow
- 100
- plateauCheckEvery
- 50
- plateauMinGenerations
- 100
- finalGeneration
- 149
- finalists
- 23
- 18
- 29
- 11
- 1
- 16
- 28
- 24
- blockStart
- 0xa52eb3c0
- selectGames
- 128
- winner
- candidate-28
- winnerMean
- 257313.8203
- label
- evolve
- command
- /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
- --init
- /home/keshav/Developer/drop7-bench/artifacts/results/EX-20260902-nnue-evolution-d3-v2-49c18bc2/RUN-20260902T035644Z-c1fd8987/pretrain/init.bin
- --lease-start
- 0xa52ea100
- --out
- /home/keshav/Developer/drop7-bench/runs/RUN-20260903T190433Z-a87fd7fc/nnue-evolution/evolve
- --population
- 32
- --games
- 32
- --generations
- 1000
- --elites
- 4
- --tournament
- 3
- --sigma-rel
- 0.05
- --seed
- 0x0e701e5a
- --threads
- 32
- --wall-seconds
- 259200
- --move-cap
- 2000
- --experiment-id
- EX-20260903-nnue-evolution-continuation2-d3-80eebad3
- --resume-population
- /home/keshav/Developer/drop7-bench/runs/RUN-20260903T032832Z-a76a6cf7/nnue-evolution/evolve/population-150.bin
- --baseline
- /home/keshav/Developer/drop7-bench/artifacts/results/EX-20260903-nnue-evolution-continuation-d3-f8ce9181/RUN-20260903T032832Z-a76a6cf7/evolve/candidate-weights.bin
- --sigma-decay-tau
- 1500
- --sigma-floor
- 0.01
- --plateau-window
- 100
- --plateau-check-every
- 50
- --plateau-min-generations
- 100
- startedAt
- 2026-09-03T19:10:50Z
- endedAt
- 2026-09-04T08:51:15Z
- wallSeconds
- 49225.9710
- userSeconds
- 1531743.1750
- systemSeconds
- 682.9510
- peakRssBytes
- 479,649,792
- exitCode
- 0
- label
- select
- command
- /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
- --select
- --lease-start
- 0xa52ea100
- --out
- /home/keshav/Developer/drop7-bench/runs/RUN-20260903T190433Z-a87fd7fc/nnue-evolution/evolve
- --generations
- 1000
- --games
- 32
- --select-games
- 128
- --threads
- 32
- --move-cap
- 2000
- startedAt
- 2026-09-04T08:51:16Z
- endedAt
- 2026-09-04T08:56:37Z
- wallSeconds
- 321.4130
- userSeconds
- 10063.2610
- systemSeconds
- 3.3730
- peakRssBytes
- 251,932,672
- exitCode
- 0
- label
- screen
- command
- /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/screen
- --candidate
- /home/keshav/Developer/drop7-bench/runs/RUN-20260903T190433Z-a87fd7fc/nnue-evolution/evolve/candidate-weights.bin
- --init
- /home/keshav/Developer/drop7-bench/artifacts/results/EX-20260902-nnue-evolution-d3-v2-49c18bc2/RUN-20260902T035644Z-c1fd8987/pretrain/init.bin
- --seeds-start
- 0xa52f2100
- --games
- 64
- --threads
- 32
- --move-cap
- 2000
- --out
- /home/keshav/Developer/drop7-bench/runs/RUN-20260903T190433Z-a87fd7fc/nnue-evolution/screen/heldout.json
- --experiment-id
- EX-20260903-nnue-evolution-continuation2-d3-80eebad3
- --arm
- baseline-run1=/home/keshav/Developer/drop7-bench/artifacts/results/EX-20260903-nnue-evolution-continuation-d3-f8ce9181/RUN-20260903T032832Z-a76a6cf7/evolve/candidate-weights.bin
- startedAt
- 2026-09-04T08:56:37Z
- endedAt
- 2026-09-04T09:04:10Z
- wallSeconds
- 453.4720
- userSeconds
- 11238.0470
- systemSeconds
- 2.3390
- peakRssBytes
- 1,793,867,776
- exitCode
- 0
- secondExperimentResultId
- RS-20260903T163321Z-733076b5
- secondExperimentRunId
- RUN-20260903T032832Z-a76a6cf7
- secondExperimentScreenCandidateVsFair
- -68,441
- secondExperimentScreenCandidateVsBaseline
- 36,278
- secondExperimentPlateauSlopeAt99
- 238.6000
- secondExperimentPlateauSlopeAt149
- 37.3000
- Single 64-game held-out screen: paired detection floors of about 43,000 points for the primary contrast and 44,000 for the baseline contrast; the baseline-run1 (secondary) contrast's point estimate of +13,572 sits well inside its own floor, so 'no confirmed improvement over the prior candidate' is the correct reading, not 'no improvement occurred' -- a true effect below about 44,000 points could not be distinguished from zero at this sample size.
- The plateau rule fired at the identical generation (149) as the second experiment despite a 3.75x slower decay constant; this is evidence against the annealing schedule being the dominant cause of the second experiment's plateau, but it is a single comparison at one alternative time constant, not a sweep, and cannot rule out that some other (e.g. much slower, or non-exponential) schedule would behave differently.
- The per-game artifact holds five arms x 64 games (320 rows); each contrast pairs two arms by seed.
- The starting population and both controls are products of the first and second experiments' training data; nothing in this run re-read either prior lease.
- pipeline.sh's screen and compare stages hardcode the third control's arm name as 'baseline-run1' regardless of which run's candidate is actually supplied via $BASELINE; in this run that arm holds the SECOND experiment's frozen candidate, not the first. The mapping is recorded accurately in evolve/config.json's baselineSha256 field (matches the second experiment's candidate hash) and in this record's metrics.priorRuns; the label itself is cosmetic and should be parameterised in a future pipeline.sh edit made only while no stage is executing.
- The fair-d4s7 arm is diagnostic only and its contrast against fair-d3s7 crossed zero on this cohort's fresh seeds (-1,409, bounds -62,290 to +58,933); this is expected cohort-to-cohort variation on a small paired sample and is not comparable across the three screens, which drew different seed blocks.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/EX-20260903-nnue-evolution-continuation2-d3-80eebad3.mdx; it renders above this record on the next request. The registered protocol itself is in the technical record above.
Record file: research/experiments/EX-20260903-nnue-evolution-continuation2-d3-80eebad3.json, validated against research/schemas/experiment-v1.schema.json. Protocol hash: c7733786132d53b7d65b1f486afb6aef7456d8e1196af943d0e18eb3f2190d90.
Registered by Claude Code / claude-sonnet-5 (claude-code-evolution-experiment).