On this page
Dates
Recorded
Record idRS-20260903T025751Z-6577b33e

No explanation has been written for this record yet.

Technical recordMetrics, gate checks and limitationsRS-20260903T025751Z-6577b33e
valid runoutcome: failnot-supported-as-testedtier: public-developmentRS-20260903T025751Z-6577b33e

Valid run of the frozen protocol EX-20260902-nnue-evolution-d3-v2-49c18bc2 on the whole workstation (32 threads), every stage completed, every artifact with zero illegal and zero incomplete decisions. Stage A: the depth-5 seven-stratum teacher played 177 complete games (21,618 sibling-complete labelled roots, 3 stopped at the 500-move cap) before the 46,800 s new-game cutoff; the teacher ran 77.5 s per root, four to seven times slower than the pilot projected, so the corpus is about a third of the 512 games the protocol allowed for. Stage B: the supervised warm start reached validation Huber 0.6873 rise units at epoch 8 of 16 on a whole-origin split (17,641/3,977 roots), and the deployment-faithful ordering probe put the depth-3 search with that leaf at top-1 agreement 0.4414 with the teacher on 256 held-out roots (mean teacher-value regret 2,395 points). Stage C: 60 generations of the mutation-only GA (population 32, 32 paired games per candidate per fresh block, fair-d3s7 and warm-start controls on every block). The population separated steadily from its warm start - ten-generation paired mean margins over the init control of +4,682, +15,853, +22,704, +32,278, +34,037, +40,431 points, best candidate +76,679 in the last block - but never approached the fair control: the population mean was above the fair leaf in 0 of the final 10 generations and the best candidate in 0 (mean margin -121,896), so the theory's training-signal falsifier fails. Elite re-selection on 128 fresh games froze candidate-29 at 202,237 (finalists spanned 187,352 to 202,237). Stage D: on the 64 never-read held-out games (0xa52e1300, opened once at 2026-09-03T02:41:29Z after the candidate's SHA-256 was recorded), the evolved candidate averaged 190,961 against the frozen fair leaf's 297,926 at the identical depth-3 configuration: paired delta -106,964 (bootstrap 95% bounds -146,580 to -69,983, Student-t lower bound -145,890, detection floor 38,357), W-T-L 14-0-50, both halves negative (-109,139 / -104,790), lower quartile 122,253 against 158,387. Every preregistered pass criterion except artifact integrity and candidate identity fails: scientific outcome fail for this exact configuration. The ablation arm shows what evolution did contribute: the unevolved warm start averaged 155,586 (-142,340 against the fair leaf), and the evolved candidate beat it on the same seeds by +35,375 (bootstrap lower bound +16,899, floor 18,949, W-T-L 36-0-28, both halves positive). Whole-game evolution with common random numbers therefore moves a 572k-weight leaf on the deployed objective, which the first (CMA-ES) leaf evolution could not show; it moved it about a quarter of the way from a warm start that plays at half the fair leaf's level. The reference arm reproduced the program's standing result: fair depth 4 over fair depth 3 +97,059 (lower bound +29,330). Read: the claim is not supported as tested; the mechanism's evolutionary leg is supported, its distillation leg is the weak link (a warm start that holds the teacher's values but not its ordering), and the budget (177 teacher games, 60 generations) was too small for evolution to cover the distance.

What it had to pass
  • All CHECK gates passed before the first training seed was read — observed: 11 of 11 gates passed on the already-open probe block (gates.log), rebuilt binaries, 9/9 unit tests
  • Every generation artifact and the screen artifact: illegalDecisions 0, incompleteDecisions 0, censored games reported — observed: 60 generation artifacts: 0 illegal, 0 incomplete, 0 censored; screen: 0 illegal, 0 incomplete, 0 censored in all four arms
  • Held-out screen, candidate vs fair-d3s7: bootstrap 95% lower bound of the paired score delta > 0 AND Student-t 95% lower bound > 0 — observed: delta -106,964; bootstrap LB -146,580, UB -69,983; t LB -145,890; paired sd 186,538; floor 38,357; W-T-L 14-0-50
  • Held-out screen: paired mean score delta > 0 in both halves (seeds 0-31 and 32-63) — observed: first half -109,139, second half -104,790
  • Held-out screen: candidate lower-quartile score >= fair-d3s7 lower-quartile score — observed: candidate Q25 122,253 vs fair-d3s7 Q25 158,387 (delta -36,134)
  • The screened candidate is the exact frozen winner of the elite re-selection; SHA-256 recorded before the screen lease opened — observed: candidate-29 of generation 59, mean 202,237 over 128 fresh games; sha256 edd0d2efd181de43f35d62c4df784cb2c789db3ae296be93a1c6c59082034a9f; recorded 2026-09-03T02:41:29Z, lease opened 2026-09-03T02:41:29Z by the screen stage
  • Theory falsifier (training signal): population mean fitness > paired fair-d3s7 control in the majority of the final 10 generations — observed: mean above fair in 0/10, top-4 in 0/10, best in 0/10; mean margin -121,896, top-4 margin -96,788
  • Theory falsifier (supervised initialisation): distilled NNUE's within-root top-1 agreement with the depth-5 teacher on whole-origin held-out roots exceeds the 0.30 level materially — observed: deployment-faithful probe (d3s7 search with the NNUE leaf vs the teacher's column) top-1 0.4414 on 256 held-out roots, regret 2,395 points; for scale, the leaf-swing diagnostic put the frozen fair leaf at 0.720 and a zero leaf at 0.450 on a different 400-root sample
  • Theory falsifier (ablative): if the candidate passes but the unevolved init passes equally, evolution contributed nothing — observed: precondition not met (candidate failed); recorded anyway: candidate minus init +35,375 (bootstrap LB +16,899, t LB +16,145, floor 18,949, W-T-L 36-0-28, halves +47,004 / +23,746); init minus fair -142,340 (LB -178,492)
Technical recordRecorded metricsRS-20260903T025751Z-6577b33e
cohort
seedsStartHex
0xa52e1300
games
64
moveCap
2,000
role
public-development, never read before this screen (lease SL-20260825T063000Z-a52e1300, opened once)
arms
candidate
games
64
mean
190961.2188
median
160107.5000
q25
122253.2500
max
477,758
movesMean
58.5000
clearsPerMove
1.7131
revealsPerMove
0.9124
censored
0
illegal
0
incomplete
0
wallSeconds
443.8840
init-d3s7
games
64
mean
155585.9688
median
140,282
q25
107423.2500
max
401,038
movesMean
48.7500
clearsPerMove
1.5510
revealsPerMove
0.7554
censored
0
illegal
0
incomplete
0
wallSeconds
352.7846
fair-d3s7
games
64
mean
297925.5938
median
254999.5000
q25
158387.2500
max
844,318
movesMean
88.0313
clearsPerMove
1.9549
revealsPerMove
1.0722
censored
0
illegal
0
incomplete
0
wallSeconds
265.4953
fair-d4s7
games
64
mean
394984.3125
median
302,556
q25
198524.7500
max
1,663,637
movesMean
114.2969
clearsPerMove
2.0547
revealsPerMove
1.1494
censored
0
illegal
0
incomplete
0
wallSeconds
11302.7039
paired
candidate-vs-fair-d3s7
meanDelta
-106964.3750
bootstrapLower95
-146579.5391
bootstrapUpper95
-69983.3727
studentTLower95
-145890.1465
pairedSd
186537.5328
detectionFloor
38356.7802
wtl
  1. 14
  2. 0
  3. 50
halves
  1. -109138.6563
  2. -104790.0938
q25Delta
-36,134
movesDelta
-29.5313
init-vs-fair-d3s7
meanDelta
-142339.6250
bootstrapLower95
-178491.5461
bootstrapUpper95
-107770.6313
studentTLower95
-178510.6720
pairedSd
173336.5226
detectionFloor
35642.3225
wtl
  1. 9
  2. 0
  3. 55
halves
  1. -156142.6563
  2. -128536.5938
q25Delta
-50,964
movesDelta
-39.2813
candidate-vs-init
meanDelta
35375.2500
bootstrapLower95
16898.9273
bootstrapUpper95
54629.2750
studentTLower95
16144.9914
pairedSd
92153.9860
detectionFloor
18949.1634
wtl
  1. 36
  2. 0
  3. 28
halves
  1. 47,004
  2. 23746.5000
q25Delta
14,830
movesDelta
9.7500
fair-d4s7-vs-fair-d3s7
meanDelta
97058.7188
bootstrapLower95
29329.5438
bootstrapUpper95
165453.2305
studentTLower95
27011.3702
pairedSd
335676.3163
detectionFloor
69023.4425
wtl
  1. 41
  2. 0
  3. 23
halves
  1. 75457.7813
  2. 118659.6563
q25Delta
40137.5000
movesDelta
26.2656
training
generationsCompleted
60
tenGenerationBlockAverages
  1. generations
    0-9
    populationMean
    159340.4412
    top4Mean
    176605.0742
    best
    183544.7219
    meanMinusInit
    4682.4068
    bestMinusInit
    28886.6875
    meanMinusFair
    -172846.5150
    bestMinusFair
    -148642.2344
  2. generations
    10-19
    populationMean
    168372.4929
    top4Mean
    189184.1141
    best
    194671.8250
    meanMinusInit
    15852.6179
    bestMinusInit
    42151.9500
    meanMinusFair
    -154441.2915
    bestMinusFair
    -128141.9594
  3. generations
    20-29
    populationMean
    174464.3515
    top4Mean
    197282.4398
    best
    203909.7531
    meanMinusInit
    22703.5952
    bestMinusInit
    52148.9969
    meanMinusFair
    -155743.1267
    bestMinusFair
    -126297.7250
  4. generations
    30-39
    populationMean
    181219.7211
    top4Mean
    203060.9922
    best
    209035.9844
    meanMinusInit
    32277.6086
    bestMinusInit
    60093.8719
    meanMinusFair
    -125175.9352
    bestMinusFair
    -97359.6719
  5. generations
    40-49
    populationMean
    184679.3310
    top4Mean
    206650.3125
    best
    213143.7469
    meanMinusInit
    34036.9685
    bestMinusInit
    62501.3844
    meanMinusFair
    -160026.6253
    bestMinusFair
    -131562.2094
  6. generations
    50-59
    populationMean
    190907.8234
    top4Mean
    216015.6836
    best
    227155.7656
    meanMinusInit
    40431.1578
    bestMinusInit
    76679.1000
    meanMinusFair
    -121896.0891
    bestMinusFair
    -85648.1469
trainingSignalCheck
rule
population mean fitness > paired fair-d3s7 control in the majority of the final 10 completed generations (theory falsifier 2)
window
  1. 50
  2. 51
  3. 52
  4. 53
  5. 54
  6. 55
  7. 56
  8. 57
  9. 58
  10. 59
generationsMeanAboveFair
0
generationsTop4AboveFair
0
generationsBestAboveFair
0
generationsInitAboveFair
0
meanMarginLast10
-121896.0891
top4MarginLast10
-96788.2289
passed
false
artifactIntegrity
generationArtifacts
60
illegalDecisions
0
incompleteDecisions
0
censoredGames
0
config
population
32
games
32
generations
60
elites
4
tournament
3
sigmaRel
0.0500
seed
0x0e701e58
leaseStart
0xa52e0500
moveCap
2,000
wallSeconds
21,600
selection
finalGeneration
59
finalists
  1. 12
  2. 1
  3. 6
  4. 29
  5. 30
  6. 23
  7. 13
  8. 4
blockStart
0xa52e0c80
selectGames
128
winner
candidate-29
winnerMean
202236.7734
corpus
games
177
roots
21,618
censoredGames
3
score
n
177
mean
433131.1525
sd
323856.2717
min
103,368
q25
229,522
median
348,847
q75
549,029
max
1,889,608
moves
n
177
mean
122.1356
sd
85.7326
min
34
q25
70
median
100
q75
150
max
500
wallSecondsPerGame
n
177
mean
9461.8629
sd
7988.6131
min
1209.2638
q25
4178.7932
median
7476.2450
q75
11333.4374
max
49497.1916
secondsPerRoot
77.4702
rootsPerGame
122.1356
legalColumnsPerRoot
n
21,618
mean
6.8572
sd
0.6843
min
1
q25
7
median
7
q75
7
max
7
teacherSiblingSpreadPoints
n
21,618
mean
40559.6546
sd
160027.6823
min
0
q25
2448.7576
median
4187.9510
q75
7480.5930
max
1017575.6843
pretrain
report
roots
21,618
games
177
trainRoots
17,641
valRoots
3,977
epochs
16
batch
64
lr
0.0003
seed
0x0e701e57
bestEpoch
8
bestValHuber
0.6873
probe
probeRoots
256
top1
0.4414
meanTeacherValueRegret
2394.6708
epochs
  1. epoch
    0
    trainHuber
    2.3016
    valHuber
    1.8859
    valPearson
    0.8790
  2. epoch
    1
    trainHuber
    1.2819
    valHuber
    0.9546
    valPearson
    0.9314
  3. epoch
    2
    trainHuber
    0.7173
    valHuber
    0.8128
    valPearson
    0.9449
  4. epoch
    3
    trainHuber
    0.5488
    valHuber
    0.7449
    valPearson
    0.9491
  5. epoch
    4
    trainHuber
    0.4514
    valHuber
    0.7100
    valPearson
    0.9512
  6. epoch
    5
    trainHuber
    0.3886
    valHuber
    0.6918
    valPearson
    0.9520
  7. epoch
    6
    trainHuber
    0.3371
    valHuber
    0.6906
    valPearson
    0.9519
  8. epoch
    7
    trainHuber
    0.3036
    valHuber
    0.6919
    valPearson
    0.9500
  9. epoch
    8
    trainHuber
    0.2706
    valHuber
    0.6873
    valPearson
    0.9504
  10. epoch
    9
    trainHuber
    0.2443
    valHuber
    0.6939
    valPearson
    0.9502
  11. epoch
    10
    trainHuber
    0.2246
    valHuber
    0.6993
    valPearson
    0.9479
  12. epoch
    11
    trainHuber
    0.2092
    valHuber
    0.6955
    valPearson
    0.9496
  13. epoch
    12
    trainHuber
    0.1915
    valHuber
    0.7119
    valPearson
    0.9480
  14. epoch
    13
    trainHuber
    0.1807
    valHuber
    0.6936
    valPearson
    0.9479
  15. epoch
    14
    trainHuber
    0.1729
    valHuber
    0.6995
    valPearson
    0.9485
  16. epoch
    15
    trainHuber
    0.1620
    valHuber
    0.6969
    valPearson
    0.9470
resources
  1. label
    corpus
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/teacher_corpus
    2. --seeds-start
    3. 0xa52e0300
    4. --games
    5. 512
    6. --threads
    7. 32
    8. --depth
    9. 5
    10. --move-cap
    11. 500
    12. --wall-seconds
    13. 46800
    14. --out
    15. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/corpus
    startedAt
    2026-09-02T04:24:10Z
    endedAt
    2026-09-02T22:39:58Z
    wallSeconds
    65747.5370
    userSeconds
    1667946.0280
    systemSeconds
    641.5270
    peakRssBytes
    1,615,892,480
    exitCode
    0
  2. label
    pretrain
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/pretrain
    2. --corpus
    3. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/corpus/corpus.jsonl
    4. --out
    5. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain
    6. --epochs
    7. 16
    8. --batch
    9. 64
    10. --lr
    11. 3e-4
    12. --seed
    13. 0x0e701e57
    14. --probe-roots
    15. 256
    16. --threads
    17. 32
    startedAt
    2026-09-02T22:42:11Z
    endedAt
    2026-09-02T22:42:25Z
    wallSeconds
    13.9780
    userSeconds
    45.2200
    systemSeconds
    0.1080
    peakRssBytes
    265,060,352
    exitCode
    0
  3. label
    evolve
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
    2. --init
    3. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain/init.bin
    4. --lease-start
    5. 0xa52e0500
    6. --out
    7. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve
    8. --population
    9. 32
    10. --games
    11. 32
    12. --generations
    13. 60
    14. --elites
    15. 4
    16. --tournament
    17. 3
    18. --sigma-rel
    19. 0.05
    20. --seed
    21. 0x0e701e58
    22. --threads
    23. 32
    24. --wall-seconds
    25. 21600
    26. --move-cap
    27. 2000
    startedAt
    2026-09-02T22:42:25Z
    endedAt
    2026-09-03T02:37:19Z
    wallSeconds
    14093.6190
    userSeconds
    441673.1460
    systemSeconds
    184.6580
    peakRssBytes
    477,446,144
    exitCode
    0
  4. label
    select
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/evolve
    2. --select
    3. --lease-start
    4. 0xa52e0500
    5. --out
    6. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve
    7. --generations
    8. 60
    9. --games
    10. 32
    11. --select-games
    12. 128
    13. --threads
    14. 32
    15. --move-cap
    16. 2000
    startedAt
    2026-09-03T02:37:19Z
    endedAt
    2026-09-03T02:41:29Z
    wallSeconds
    250.4040
    userSeconds
    7852.4190
    systemSeconds
    3.9940
    peakRssBytes
    252,526,592
    exitCode
    0
  5. label
    screen
    command
    1. /home/keshav/Developer/drop7-bench/approaches/lifetime-objective/nnue-evolution/target/release/screen
    2. --candidate
    3. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/evolve/candidate-weights.bin
    4. --init
    5. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/pretrain/init.bin
    6. --seeds-start
    7. 0xa52e1300
    8. --games
    9. 64
    10. --threads
    11. 32
    12. --move-cap
    13. 2000
    14. --out
    15. /home/keshav/Developer/drop7-bench/runs/RUN-20260902T035644Z-c1fd8987/nnue-evolution/screen/heldout.json
    startedAt
    2026-09-03T02:41:29Z
    endedAt
    2026-09-03T02:52:22Z
    wallSeconds
    652.9870
    userSeconds
    12346.6780
    systemSeconds
    3.8980
    peakRssBytes
    1,792,266,240
    exitCode
    0
Limitations
  • Single 64-game held-out screen: the paired detection floor is about 38,000 points for the primary contrast and 19,000 for the ablation; the primary result is far outside its floor, the ablation clears its own.
  • The teacher corpus reached 177 of the 512 games the protocol allowed for because the depth-5 teacher ran four to seven times slower than the pilot projected; the protocol makes the completed games the corpus, so the result is valid, but it rejects this configuration at this corpus size, not the design at 512 games.
  • The supervised warm start plays at about half the fair leaf's level; the deployment-faithful probe (0.441 top-1) and the leaf-swing diagnostic (zero leaf 0.450 on a different sample) suggest imitation of state values barely improved the search's ordering over no leaf at all. Evolution then had roughly 150,000 paired points to make up in 60 generations and made up about 35,000 to 40,000.
  • The per-game artifact holds four arms x 64 games (256 rows); the primary contrast pairs the candidate and fair-d3s7 rows by seed.
  • The fair-d4s7 arm is diagnostic only and reproduces the standing depth-4-over-depth-3 result on these seeds; it is not part of the gate.
  • Two operational faults during the unattended chain (a false liveness reading, then a script replacement that crashed the corpus stage's shell after its binary had exited 0) are recorded in the run record; the 'corpus: done' marker in pipeline.log was appended by hand with a note. No artifact was affected.
  • The elite re-selection block (0xa52e0c80, 128 games) and every fitness block are training-lease seeds; the last generation's 32-game leaders read 20,000 to 60,000 above their 128-game re-selection means, which is the best-of-32 bias the re-selection exists to remove.

Recorded against Depth-5-distilled NNUE leaf refined by paired-fitness whole-game evolution inside the depth-3 fair search (v2: operational parameters fixed in the record body).

Agent contextHow to extend this record

To add a reader-facing explanation, write web/content/research/RS-20260903T025751Z-6577b33e.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.

Record file: research/results/RS-20260903T025751Z-6577b33e.json, validated against research/schemas/result-v1.schema.json.

Run ids
  • RUN-20260902T035644Z-c1fd8987
Contribution ids
  • CT-20260902T042620Z-3b4bf7a8
Per-game artifact
artifacts/results/EX-20260902-nnue-evolution-d3-v2-49c18bc2/RUN-20260902T035644Z-c1fd8987/screen/heldout.json (sha256 70009c09a3fdacd40c9c25fc60bfdb7ffe698041c8e837f714732b86a2874383, 256 records)
Artifact manifest
artifacts/results/EX-20260902-nnue-evolution-d3-v2-49c18bc2/RUN-20260902T035644Z-c1fd8987/manifest.json
Machine profiles
  • research/system-profiles/MACH-20260902T035644Z-f5e59b6e.json