Drop7 Research
← Theories

A leaf-affordable NNUE-class student trained on successor-closed exact D4 sibling values reproduces D4's within-root ordering at the d4q gate thresholds

not-supported-as-testedassessedevidence: pilotpublic-policyTH-20260823-nnue-class-holds-d4-ordering-c7b397a5
No explanation has been written for this theory yet. Add web/content/research/TH-20260823-nnue-class-holds-d4-ordering-c7b397a5.mdx and it will appear here. The registered record is shown below.

The registered record

Claim

A LeafNet-shaped student (approaches/lifetime-objective/learned-leaf/leaf_features.py features, EmbeddingBag(8902,64)->32, ~572k parameters, ~1.3 us per state) trained with a within-root listwise plus pairwise loss on the existing successor-closed exact D4 sibling-value labels (runs/RUN-20260821T085042Z-c5cf0e71/d4q-labels, 291,890 afterstates over 8,639 roots) reproduces fair D4's within-root ordering on whole-origin held-out roots at the frozen d4q thresholds: top-1 >= 0.60, pairwise >= 0.78, normalised regret <= 0.13, in each half-fold.

Mechanism

RS-20260821T104500Z-77d21e90 showed a 3.4M-parameter afterstate CNN reaches only top-1 0.375 on these labels, below exact D1 (0.486), and status.md relocated the obstacle to the capacity of a compact board evaluator. The leaf-affordable NNUE class has never been trained on these labels: it was trained on played-action lifetime (saturated at Pearson 0.855-0.857, RS-20260822T024228Z-94090db1) and on planner values (finding-11). The draft self-play theory TH-20260821-search-guided-self-play-at-scale-299ed02f needs a student that both holds D4's ordering and runs at leaf cost; this probe decides whether that student can exist in the only leaf-affordable model class, on data already on disk, for about one CPU-hour. A fail is a valid negative that routes the self-play loop to a root-prior redesign. Design: runs/RUN-20260823T191900Z-b9f8f80d/kimi-k3-open-items.md (K4-D probe, Kimi K3). Designer's prior: P(pass) about 0.08.

What would prove it wrong

  • Held-out (whole-origin, the d4q experiment's own split) top-1 < 0.60 OR pairwise < 0.78 OR normalised regret > 0.13 in either half-fold, for the best of 5 initialisation seeds judged on the validation origins only, refutes the claim for this model class at this size; the held-out panel is already-read diagnostic data, so even a pass is not evidence and licenses only a fresh-manifest confirmation.
  • A pass that fails a fresh whole-origin label manifest confirmation refutes it at one tier higher.
registered 2026-08-23T19:29:58Z by unknown / unknown

Experiments that test it

Results recorded against it

valid run · outcome: failnot-supported-as-testedpilotRS-20260823T194142Z-946e3cd1

The leaf-affordable NNUE-class student does not reproduce fair D4's within-root ordering from the existing successor-closed labels. Five initialisation seeds of the LeafNet h64/m32 student (571,905 parameters; features from learned-leaf/leaf_features.py) were trained for 30 epochs on the 6,551 training roots of runs/RUN-20260821T085042Z-c5cf0e71/d4q-labels with the preregistered listwise + gap-weighted pairwise + 0.1 MSE loss; the seed with the best calibration-fold top-1 (0xA52E02, 0.3281) was then read once on the 3,030-root held-out panel of that experiment. Held-out top-1 was 0.296 / 0.3011 in the two origin-seed half-folds (gate >= 0.60), pairwise 0.6266 / 0.6259 (>= 0.78), normalised regret 0.3354 / 0.3379 (<= 0.13): every criterion fails in both half-folds. The other four seeds lie between 0.2987 and 0.3135 pooled top-1, all below the 3.4M-parameter afterstate CNN of RS-20260821T104500Z-77d21e90 (0.375) and below exact D1 (0.486). The student is leaf-affordable as claimed (1.27 us per state through the existing d7leaf export and leaf-check), but at this size it holds less of D4's ordering than the CNN did. Per the preregistered failure action: no leaf-affordable NNUE-class student of this size reproduces D4's ordering from these labels; the self-play loop TH-20260821-...-299ed02f stays blocked in its leaf form and its redesign is routed to the larger-student question. Diagnostic tier: the held-out panel had already been read once by the d4q experiment, so this is a diagnostic read on existing data, not new evidence about fresh roots.

What it had to pass
  • Label file is the preregistered artifact: 291,890 rows / 8,639 roots — observed: 291890 rows, 8639 roots (train 6,551, calibration 2,088)
  • No non-finite loss in any seed; all five seeds completed 30 epochs within 5,400 wall-seconds — observed: 5 of 5 seeds valid; 166.08 wall-seconds total
  • Selected seed (best validation top-1, 0xA52E02) held-out top1 >= 0.6 in each half-fold — observed: half1 0.2960, half2 0.3011
  • Selected seed (best validation top-1, 0xA52E02) held-out pairwise >= 0.78 in each half-fold — observed: half1 0.6266, half2 0.6259
  • Selected seed (best validation top-1, 0xA52E02) held-out regret <= 0.13 in each half-fold — observed: half1 0.3354, half2 0.3379
Recorded metrics
labelRows
291,890
labelRoots
8,639
trainRoots
6,551
validationRoots
2,088
heldoutRoots
3,030
modelParameters
571,905
seedsRun
5
selectedSeed
0xA52E02
selectedValidationTop1
0.3281
top1Pooled
0.2987
top1Half1
0.2960
top1Half2
0.3011
top2Pooled
0.5099
pairwisePooled
0.6262
pairwiseHalf1
0.6266
pairwiseHalf2
0.6259
regretPooled
0.3367
regretHalf1
0.3354
regretHalf2
0.3379
heldoutTop1AllSeeds
0xA52E01
0.3026
0xA52E02
0.2987
0xA52E03
0.3069
0xA52E04
0.3132
0xA52E05
0.3135
validationTop1FinalAllSeeds
0xA52E01
0.3051
0xA52E02
0.3281
0xA52E03
0.3089
0xA52E04
0.3127
0xA52E05
0.3027
validationTop1BestEpochAllSeeds
0xA52E01
0.3554
0xA52E02
0.3549
0xA52E03
0.3520
0xA52E04
0.3525
0xA52E05
0.3592
heldoutTop1SpreadAcrossSeeds
  1. 0.2987
  2. 0.3135
inferenceMicrosecondsPerState
1.2721
trainWallSeconds
166.0800
comparatorAfterstateCnnTop1Pooled
0.3752
referenceD1Top1
0.4860
referenceD2Top1
0.5680
Limitations
  • Diagnostic-tier read: the 3,030-root held-out panel (d4q-labels-gate) was already opened once by EX-20260821-afterstate-d4q-stage1; no fresh roots were labelled. The result schema has no diagnostic tier, so it is recorded at the lowest tier (pilot); it is not development-tier evidence.
  • The preregistration names only the d4q-labels file, which carries folds train and calibration but no held-out fold; the held-out panel and its origin-seed half-folds were taken from the same run's d4q-labels-gate/d4q-labels.tsv and corpus-gate/roots.tsv, exactly as d4q.py consumed them, which is the reading the protocol's 'train/validation/held-out origins are reused unchanged' clause intends.
  • Epoch count was fixed at 30 by the protocol and the final-epoch model was used. Calibration-fold top-1 peaked at 0.352-0.359 between epochs 5 and 10 for every seed and then declined as the training listwise loss kept falling (overfitting); even the best epoch is far below the 0.60 gate, so early stopping would not change the verdict, but the reported held-out numbers are for the overfitted final epoch.
  • Loss coefficients not fixed by the hypothesis text were set to 1.0 for the pairwise term (d4-q-clone used 0.35) and the prediction-side softmax temperature was 1 in standardised value units; only the listed configuration is rejected.
  • CPU training with 8 torch threads; the EmbeddingBag backward is not bit-deterministic across thread counts, so a re-run reproduces the numbers only approximately.
  • Inference cost is from leaf-check on the RUN-A51D-corpus mix-d3 states on a shared host; informational only.