On this page
Dates
Recorded
Record idRS-20260904T090620Z-e5731bf0

No explanation has been written for this record yet.

Technical recordMetrics, gate checks and limitationsRS-20260904T090620Z-e5731bf0
valid runoutcome: failnot-supported-as-testedtier: public-developmentRS-20260904T090620Z-e5731bf0

Valid run of the frozen protocol EX-20260903-nnue-evolution-continuation2-d3-80eebad3 (second successor to the theory, third experiment in the series), every stage completed, every artifact with zero illegal and zero incomplete decisions. Stage C resumed from the second experiment's checkpointed generation-150 population (SHA-256 33dcc25b88ab56a0a80329bf11983b539883aa922900b03de3cea71e1da3529d) with that experiment's frozen candidate (SHA-256 759084fab97599818d01a5da6d18b9cc77172249f09a68d5104361bca385e09b) as the third paired control, testing whether a 3.75x slower mutation-size decay (time constant 1,500 generations vs. 400 before; sigma stayed at 0.0500 at generation 0, 0.0453 at the stop, versus 0.0345 at the equivalent generation last time) would let evolution get further past the point where the prior run plateaued. The identical preregistered plateau rule stopped this run at the identical generation as before, 149: the check after generation 99 found the paired margin over the fair control rising at +201.7 points per generation (lower bound +11.7, versus +239 and +30 in the second experiment) and continued; the check after generation 149 found +164.0 per generation (lower bound -28.5, versus +37 and -159 before) and stopped. The improvement rate at the stopping point was more than four times the prior run's (164 against 37 points/generation), and the best candidate beat the fair control on 26 of 150 blocks against 9 in the second experiment, but the lower bound of the fitted slope still crossed zero at the same 100-generation checkpoint, so the rule fired regardless. Fifty-generation population-mean averages: 235,353, 238,797, 244,800 (rising monotonically, unlike the second experiment's flat third block); paired margin over the immediately-prior candidate +4,301, -949, +13,589; over the fair control -99,479, -88,006, -77,654. Training-signal falsifier fails again (population mean above fair in 0 of the last 10 generations, margin -70,199; best above fair in 4 of the last 10). Elite re-selection on 128 fresh games (0xa52eb3c0) froze candidate-28 at 257,314 (finalists 232,499 to 257,314, a tighter spread than the second experiment's). Stage D, 64 never-read games (0xa52f2100, opened once at 2026-09-04T08:56:37Z after the candidate's SHA-256 was recorded), five arms: this candidate averaged 249,757 against the frozen fair leaf's 335,266, paired -85,509 (bootstrap 95% bounds -128,482 to -43,591, Student-t lower bound -129,423, floor 43,272), W-T-L 22-0-42, both halves negative (-59,644 / -111,374), lower quartile 156,534 against 176,108: every screen criterion fails, scientific outcome fail for this configuration, consistent with both prior screens. The preregistered secondary contrast is the key negative finding of this run: the candidate beat the immediately-prior (second experiment's) frozen candidate by only +13,572 on the same seeds, with a bootstrap 95% lower bound of -30,165 and a Student-t lower bound of -30,869 -- both crossing zero (detection floor 43,791), W-T-L 31-0-33. Unlike the second experiment's clearly positive +36,278 (lower bound +9,085) over the first experiment's candidate, this third experiment's gain over the second is NOT statistically distinguishable from zero at this screen size: three months of relatively larger mutations bought a point estimate about a third the size of the previous continuation's out-of-sample gain, and it is not confidently positive. The second experiment's candidate itself scored -99,081 against the fair leaf on this fresh block (lower bound -154,094), broadly consistent with its own screen result of -68,441 (lower bound -112,090), a second informal replication. Over the warm start the candidate is +106,846 (lower bound +76,970). The reference arm gave fair depth 4 minus fair depth 3 of -1,409 (bounds -62,290 to +58,933, crossing zero on this cohort, a diagnostic-only reading and not comparable across screens with different fresh seeds). Read together with the second experiment: slowing the mutation-size decay produced a visibly healthier training curve (a still-rising population mean through all three 50-generation blocks, more generations where the best candidate beat the fair control) but did NOT produce a statistically confirmed improvement over the immediately-prior candidate on held-out games, and the plateau rule still stopped the run at the same generation. The most defensible reading is that the annealing schedule was not the dominant cause of the earlier plateau -- something else (population size, games per candidate, or the leaf class's genuine ceiling under this search depth) is the binding constraint, though a schedule 3.75x slower is not proof that no schedule would help; a much slower schedule, or removing the anneal-driven exploration decay entirely in favour of a fixed sigma with more games per candidate, remains untested.

What it had to pass
  • All 11 CHECK gates passed on the rebuilt binaries before the first leased seed was read — observed: ALL GATES PASSED (gates.log), 13/13 unit tests, smoke test of resume/tau/baseline/screen-arm flags on the already-open probe block
  • Every generation artifact and the screen artifact: illegalDecisions 0, incompleteDecisions 0, censored games reported — observed: 150 generation artifacts: 0 illegal, 0 incomplete, 0 censored; screen: 0 illegal, 0 incomplete, 0 censored in all five arms
  • Held-out screen, candidate vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0 — observed: delta -85,509; bootstrap LB -128,482, UB -43,591; t LB -129,423; paired sd 210,440; floor 43,272; W-T-L 22-0-42
  • Held-out screen: paired mean score delta > 0 in both halves — observed: first half -59,644, second half -111,374
  • Held-out screen: candidate lower-quartile score >= fair-d3s7 lower-quartile score — observed: candidate Q25 156,534 vs fair-d3s7 Q25 176,108 (delta -19,573)
  • The screened candidate is the exact frozen winner of the elite re-selection; SHA-256 recorded before the screen lease opened — observed: candidate-28 of generation 149, mean 257,314 over 128 fresh games at 0xa52eb3c0; sha256 5318e70e8e1c2bc5a872d4b61a472d657b96d0a6bf50875a3a325403de682f2a; lease opened 2026-09-04T08:56:37Z by the screen stage
  • Reported alongside the gate: candidate vs baseline-run1 (the immediately-prior candidate) bootstrap 95% lower bound > 0 (the slower decay found improvement the faster one had not) — observed: delta +13,572; bootstrap LB -30,165, UB +57,575; t LB -30,869; floor 43,791; W-T-L 31-0-33; NOT confidently positive, unlike the prior continuation's +36,278 (LB +9,085) over its own predecessor
  • Reported alongside the gate: the generation of the stop and the plateau statistics, compared against the second experiment's own checks — observed: stopped by the plateau rule after generation 149 (150 completed), the SAME generation as the second experiment despite the 3.75x slower decay; check after generation 99: slope +201.7/gen (second experiment: +239) lower bound +11.7 (continue); after generation 149: slope +164.0/gen (second experiment: +37) lower bound -28.5 (stop); best-above-fair on 26/150 blocks (second experiment: 9/150)
Technical recordRecorded metricsRS-20260904T090620Z-e5731bf0
cohort
seedsStartHex
0xa52f2100
games
64
moveCap
2,000
role
public-development, never read before this screen (lease SL-20260903T190000Z-a52f2100, opened once)
arms
candidate
games
64
mean
249756.6875
median
192,147
q25
156534.5000
max
710,297
movesMean
75.0938
clearsPerMove
1.8658
revealsPerMove
1.0121
censored
0
illegal
0
incomplete
0
wallSeconds
566.6732
baseline-run1
games
64
mean
236184.1875
median
209822.5000
q25
141397.7500
max
660,003
movesMean
70.7813
clearsPerMove
1.8567
revealsPerMove
1.0188
censored
0
illegal
0
incomplete
0
wallSeconds
525.1998
init-d3s7
games
64
mean
142910.7656
median
124023.5000
q25
106530.2500
max
269,984
movesMean
45.2500
clearsPerMove
1.4914
revealsPerMove
0.7117
censored
0
illegal
0
incomplete
0
wallSeconds
314.6850
fair-d3s7
games
64
mean
335265.6406
median
272317.5000
q25
176107.5000
max
1,080,133
movesMean
98.2188
clearsPerMove
2.0051
revealsPerMove
1.1195
censored
0
illegal
0
incomplete
0
wallSeconds
273.8052
fair-d4s7
games
64
mean
333857.0625
median
285,125
q25
194039.2500
max
800,130
movesMean
97.2344
clearsPerMove
1.9994
revealsPerMove
1.1093
censored
0
illegal
0
incomplete
0
wallSeconds
9576.0764
paired
candidate-vs-fair-d3s7
meanDelta
-85508.9531
bootstrapLower95
-128482.1328
bootstrapUpper95
-43591.2375
studentTLower95
-129422.5221
pairedSd
210439.7295
detectionFloor
43271.6694
wtl
  1. 22
  2. 0
  3. 42
halves
  1. -59643.7500
  2. -111374.1563
q25Delta
-19,573
movesDelta
-23.1250
init-vs-fair-d3s7
meanDelta
-192354.8750
bootstrapLower95
-234603.4492
bootstrapUpper95
-152044.4570
studentTLower95
-234770.0951
pairedSd
203259.4400
detectionFloor
41795.2224
wtl
  1. 6
  2. 0
  3. 58
halves
  1. -171614.0313
  2. -213095.7188
q25Delta
-69577.2500
movesDelta
-52.9688
candidate-vs-init
meanDelta
106845.9219
bootstrapLower95
76970.4852
bootstrapUpper95
138415.2789
studentTLower95
75526.2930
pairedSd
150087.8742
detectionFloor
30861.8191
wtl
  1. 53
  2. 0
  3. 11
halves
  1. 111970.2813
  2. 101721.5625
q25Delta
50004.2500
movesDelta
29.8438
fair-d4s7-vs-fair-d3s7
meanDelta
-1408.5781
bootstrapLower95
-62290.3523
bootstrapUpper95
58933.4102
studentTLower95
-63294.7565
pairedSd
296566.8913
detectionFloor
60981.5670
wtl
  1. 33
  2. 0
  3. 31
halves
  1. 32695.7813
  2. -35512.9375
q25Delta
17931.7500
movesDelta
-0.9844
candidate-vs-baseline-run1
meanDelta
13572.5000
bootstrapLower95
-30164.9336
bootstrapUpper95
57575.2000
studentTLower95
-30868.5694
pairedSd
212967.5824
detectionFloor
43791.4591
wtl
  1. 31
  2. 0
  3. 33
halves
  1. 2355.4688
  2. 24789.5313
q25Delta
15136.7500
movesDelta
4.3125
baseline-run1-vs-fair-d3s7
meanDelta
-99081.4531
bootstrapLower95
-154094.1469
bootstrapUpper95
-44893.0578
studentTLower95
-154662.2389
pairedSd
266350.6017
detectionFloor
54768.3425
wtl
  1. 24
  2. 0
  3. 40
halves
  1. -61999.2188
  2. -136163.6875
q25Delta
-34709.7500
movesDelta
-27.4375
training
generationsCompleted
150
stoppedOnPlateau
true
plateauChecks
  1. generation
    99
    window
    100
    slopePerGeneration
    201.7241
    standardError
    115.5014
    lowerBound95
    11.7413
    windowMeanFirstHalf
    -99479.2672
    windowMeanSecondHalf
    -88005.6168
    stop
    false
  2. generation
    149
    window
    100
    slopePerGeneration
    164.0170
    standardError
    117.0542
    lowerBound95
    -28.5201
    windowMeanFirstHalf
    -88005.6168
    windowMeanSecondHalf
    -77654.0195
    stop
    true
fiftyGenerationBlockAverages
  1. generations
    0-49
    populationMean
    235352.7134
    meanMinusFair
    -99479.2672
    meanMinusBaseline
    4300.7409
    sigmaAtEnd
    0.0484
  2. generations
    50-99
    populationMean
    238796.7395
    meanMinusFair
    -88005.6168
    meanMinusBaseline
    -948.5393
    sigmaAtEnd
    0.0468
  3. generations
    100-149
    populationMean
    244799.5112
    meanMinusFair
    -77654.0195
    meanMinusBaseline
    13589.1018
    sigmaAtEnd
    0.0453
bestAboveFairGenerations
  1. 1
  2. 20
  3. 34
  4. 38
  5. 40
  6. 45
  7. 52
  8. 53
  9. 71
  10. 74
  11. 77
  12. 88
  13. 91
  14. 97
  15. 100
  16. 103
  17. 107
  18. 108
  19. 113
  20. 121
  21. 124
  22. 130
  23. 140
  24. 142
  25. 146
  26. 147
trainingSignalCheck
rule
population mean fitness > paired fair-d3s7 control in the majority of the final 10 completed generations (theory falsifier 2)
window
  1. 140
  2. 141
  3. 142
  4. 143
  5. 144
  6. 145
  7. 146
  8. 147
  9. 148
  10. 149
generationsMeanAboveFair
0
generationsTop4AboveFair
2
generationsBestAboveFair
4
generationsInitAboveFair
0
meanMarginLast10
-70199.2831
top4MarginLast10
-28764.4430
passed
false
artifactIntegrity
generationArtifacts
150
illegalDecisions
0
incompleteDecisions
0
censoredGames
0
config
experiment
EX-20260903-nnue-evolution-continuation2-d3-80eebad3
population
32
games
32
generations
1,000
elites
4
tournament
3
sigmaRel
0.0500
sigmaDecayTau
1,500
sigmaFloor
0.0100
seed
0x0e701e5a
leaseStart
0xa52ea100
moveCap
2,000
wallSeconds
259,200
resumePopulation
/home/keshav/Developer/drop7-bench/runs/RUN-20260903T032832Z-a76a6cf7/nnue-evolution/evolve/population-150.bin
resumePopulationSha256
33dcc25b88ab56a0a80329bf11983b539883aa922900b03de3cea71e1da3529d
baseline
/home/keshav/Developer/drop7-bench/artifacts/results/EX-20260903-nnue-evolution-continuation-d3-f8ce9181/RUN-20260903T032832Z-a76a6cf7/evolve/candidate-weights.bin
baselineSha256
759084fab97599818d01a5da6d18b9cc77172249f09a68d5104361bca385e09b
plateauWindow
100
plateauCheckEvery
50
plateauMinGenerations
100
selection
finalGeneration
149
finalists
  1. 23
  2. 18
  3. 29
  4. 11
  5. 1
  6. 16
  7. 28
  8. 24
blockStart
0xa52eb3c0
selectGames
128
winner
candidate-28
winnerMean
257313.8203
resources
  1. label
    evolve
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
    2. --init
    3. /home/keshav/Developer/drop7-bench/artifacts/results/EX-20260902-nnue-evolution-d3-v2-49c18bc2/RUN-20260902T035644Z-c1fd8987/pretrain/init.bin
    4. --lease-start
    5. 0xa52ea100
    6. --out
    7. /home/keshav/Developer/drop7-bench/runs/RUN-20260903T190433Z-a87fd7fc/nnue-evolution/evolve
    8. --population
    9. 32
    10. --games
    11. 32
    12. --generations
    13. 1000
    14. --elites
    15. 4
    16. --tournament
    17. 3
    18. --sigma-rel
    19. 0.05
    20. --seed
    21. 0x0e701e5a
    22. --threads
    23. 32
    24. --wall-seconds
    25. 259200
    26. --move-cap
    27. 2000
    28. --experiment-id
    29. EX-20260903-nnue-evolution-continuation2-d3-80eebad3
    30. --resume-population
    31. /home/keshav/Developer/drop7-bench/runs/RUN-20260903T032832Z-a76a6cf7/nnue-evolution/evolve/population-150.bin
    32. --baseline
    33. /home/keshav/Developer/drop7-bench/artifacts/results/EX-20260903-nnue-evolution-continuation-d3-f8ce9181/RUN-20260903T032832Z-a76a6cf7/evolve/candidate-weights.bin
    34. --sigma-decay-tau
    35. 1500
    36. --sigma-floor
    37. 0.01
    38. --plateau-window
    39. 100
    40. --plateau-check-every
    41. 50
    42. --plateau-min-generations
    43. 100
    startedAt
    2026-09-03T19:10:50Z
    endedAt
    2026-09-04T08:51:15Z
    wallSeconds
    49225.9710
    userSeconds
    1531743.1750
    systemSeconds
    682.9510
    peakRssBytes
    479,649,792
    exitCode
    0
  2. label
    select
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
    2. --select
    3. --lease-start
    4. 0xa52ea100
    5. --out
    6. /home/keshav/Developer/drop7-bench/runs/RUN-20260903T190433Z-a87fd7fc/nnue-evolution/evolve
    7. --generations
    8. 1000
    9. --games
    10. 32
    11. --select-games
    12. 128
    13. --threads
    14. 32
    15. --move-cap
    16. 2000
    startedAt
    2026-09-04T08:51:16Z
    endedAt
    2026-09-04T08:56:37Z
    wallSeconds
    321.4130
    userSeconds
    10063.2610
    systemSeconds
    3.3730
    peakRssBytes
    251,932,672
    exitCode
    0
  3. label
    screen
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/screen
    2. --candidate
    3. /home/keshav/Developer/drop7-bench/runs/RUN-20260903T190433Z-a87fd7fc/nnue-evolution/evolve/candidate-weights.bin
    4. --init
    5. /home/keshav/Developer/drop7-bench/artifacts/results/EX-20260902-nnue-evolution-d3-v2-49c18bc2/RUN-20260902T035644Z-c1fd8987/pretrain/init.bin
    6. --seeds-start
    7. 0xa52f2100
    8. --games
    9. 64
    10. --threads
    11. 32
    12. --move-cap
    13. 2000
    14. --out
    15. /home/keshav/Developer/drop7-bench/runs/RUN-20260903T190433Z-a87fd7fc/nnue-evolution/screen/heldout.json
    16. --experiment-id
    17. EX-20260903-nnue-evolution-continuation2-d3-80eebad3
    18. --arm
    19. baseline-run1=/home/keshav/Developer/drop7-bench/artifacts/results/EX-20260903-nnue-evolution-continuation-d3-f8ce9181/RUN-20260903T032832Z-a76a6cf7/evolve/candidate-weights.bin
    startedAt
    2026-09-04T08:56:37Z
    endedAt
    2026-09-04T09:04:10Z
    wallSeconds
    453.4720
    userSeconds
    11238.0470
    systemSeconds
    2.3390
    peakRssBytes
    1,793,867,776
    exitCode
    0
priorRuns
secondExperimentResultId
RS-20260903T163321Z-733076b5
secondExperimentRunId
RUN-20260903T032832Z-a76a6cf7
secondExperimentScreenCandidateVsFair
-68,441
secondExperimentScreenCandidateVsBaseline
36,278
secondExperimentPlateauSlopeAt99
238.6000
secondExperimentPlateauSlopeAt149
37.3000
Limitations
  • Single 64-game held-out screen: paired detection floors of about 43,000 points for the primary contrast and 44,000 for the baseline contrast; the baseline-run1 (secondary) contrast's point estimate of +13,572 sits well inside its own floor, so 'no confirmed improvement over the prior candidate' is the correct reading, not 'no improvement occurred' -- a true effect below about 44,000 points could not be distinguished from zero at this sample size.
  • The plateau rule fired at the identical generation (149) as the second experiment despite a 3.75x slower decay constant; this is evidence against the annealing schedule being the dominant cause of the second experiment's plateau, but it is a single comparison at one alternative time constant, not a sweep, and cannot rule out that some other (e.g. much slower, or non-exponential) schedule would behave differently.
  • The per-game artifact holds five arms x 64 games (320 rows); each contrast pairs two arms by seed.
  • The starting population and both controls are products of the first and second experiments' training data; nothing in this run re-read either prior lease.
  • pipeline.sh's screen and compare stages hardcode the third control's arm name as 'baseline-run1' regardless of which run's candidate is actually supplied via $BASELINE; in this run that arm holds the SECOND experiment's frozen candidate, not the first. The mapping is recorded accurately in evolve/config.json's baselineSha256 field (matches the second experiment's candidate hash) and in this record's metrics.priorRuns; the label itself is cosmetic and should be parameterised in a future pipeline.sh edit made only while no stage is executing.
  • The fair-d4s7 arm is diagnostic only and its contrast against fair-d3s7 crossed zero on this cohort's fresh seeds (-1,409, bounds -62,290 to +58,933); this is expected cohort-to-cohort variation on a small paired sample and is not comparable across the three screens, which drew different seed blocks.

Recorded against Second continuation of the depth-5-distilled NNUE leaf evolution: resume from the plateau population with a slower mutation-size decay (up to 1,000 further generations).

Agent contextHow to extend this record

To add a reader-facing explanation, write web/content/research/RS-20260904T090620Z-e5731bf0.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.

Record file: research/results/RS-20260904T090620Z-e5731bf0.json, validated against research/schemas/result-v1.schema.json.

Run ids
  • RUN-20260903T190433Z-a87fd7fc
Contribution ids
  • CT-20260903T190244Z-ca639e15
Per-game artifact
artifacts/results/EX-20260903-nnue-evolution-continuation2-d3-80eebad3/RUN-20260903T190433Z-a87fd7fc/screen/heldout.json (sha256 e133b2ed1e3519f6d2b0d8f59108831bf5221a142486d35e02c2affc35a98d45, 320 records)
Artifact manifest
artifacts/results/EX-20260903-nnue-evolution-continuation2-d3-80eebad3/RUN-20260903T190433Z-a87fd7fc/manifest.json
Machine profiles
  • research/system-profiles/MACH-20260903T190714Z-20004b2f.json