A leaf-affordable NNUE-class student trained on successor-closed exact D4 sibling values reproduces D4's within-root ordering at the d4q gate thresholds
web/content/research/TH-20260823-nnue-class-holds-d4-ordering-c7b397a5.mdx and it will appear here. The registered record is shown below.The registered record
Claim
A LeafNet-shaped student (approaches/lifetime-objective/learned-leaf/leaf_features.py features, EmbeddingBag(8902,64)->32, ~572k parameters, ~1.3 us per state) trained with a within-root listwise plus pairwise loss on the existing successor-closed exact D4 sibling-value labels (runs/RUN-20260821T085042Z-c5cf0e71/d4q-labels, 291,890 afterstates over 8,639 roots) reproduces fair D4's within-root ordering on whole-origin held-out roots at the frozen d4q thresholds: top-1 >= 0.60, pairwise >= 0.78, normalised regret <= 0.13, in each half-fold.
Mechanism
RS-20260821T104500Z-77d21e90 showed a 3.4M-parameter afterstate CNN reaches only top-1 0.375 on these labels, below exact D1 (0.486), and status.md relocated the obstacle to the capacity of a compact board evaluator. The leaf-affordable NNUE class has never been trained on these labels: it was trained on played-action lifetime (saturated at Pearson 0.855-0.857, RS-20260822T024228Z-94090db1) and on planner values (finding-11). The draft self-play theory TH-20260821-search-guided-self-play-at-scale-299ed02f needs a student that both holds D4's ordering and runs at leaf cost; this probe decides whether that student can exist in the only leaf-affordable model class, on data already on disk, for about one CPU-hour. A fail is a valid negative that routes the self-play loop to a root-prior redesign. Design: runs/RUN-20260823T191900Z-b9f8f80d/kimi-k3-open-items.md (K4-D probe, Kimi K3). Designer's prior: P(pass) about 0.08.
What would prove it wrong
- Held-out (whole-origin, the d4q experiment's own split) top-1 < 0.60 OR pairwise < 0.78 OR normalised regret > 0.13 in either half-fold, for the best of 5 initialisation seeds judged on the validation origins only, refutes the claim for this model class at this size; the held-out panel is already-read diagnostic data, so even a pass is not evidence and licenses only a fresh-manifest confirmation.
- A pass that fails a fresh whole-origin label manifest confirmation refutes it at one tier higher.
Experiments that test it
- NNUE-class student on existing successor-closed D4 labels: within-root ordering probeLeafNet h64/m32 listwise student on d4q labels vs exact fair D4 ordering (d4q labels); exact D1 0.486 and D2 0.568 top-1 as calibration · CHECK · completed
Results recorded against it
The leaf-affordable NNUE-class student does not reproduce fair D4's within-root ordering from the existing successor-closed labels. Five initialisation seeds of the LeafNet h64/m32 student (571,905 parameters; features from learned-leaf/leaf_features.py) were trained for 30 epochs on the 6,551 training roots of runs/RUN-20260821T085042Z-c5cf0e71/d4q-labels with the preregistered listwise + gap-weighted pairwise + 0.1 MSE loss; the seed with the best calibration-fold top-1 (0xA52E02, 0.3281) was then read once on the 3,030-root held-out panel of that experiment. Held-out top-1 was 0.296 / 0.3011 in the two origin-seed half-folds (gate >= 0.60), pairwise 0.6266 / 0.6259 (>= 0.78), normalised regret 0.3354 / 0.3379 (<= 0.13): every criterion fails in both half-folds. The other four seeds lie between 0.2987 and 0.3135 pooled top-1, all below the 3.4M-parameter afterstate CNN of RS-20260821T104500Z-77d21e90 (0.375) and below exact D1 (0.486). The student is leaf-affordable as claimed (1.27 us per state through the existing d7leaf export and leaf-check), but at this size it holds less of D4's ordering than the CNN did. Per the preregistered failure action: no leaf-affordable NNUE-class student of this size reproduces D4's ordering from these labels; the self-play loop TH-20260821-...-299ed02f stays blocked in its leaf form and its redesign is routed to the larger-student question. Diagnostic tier: the held-out panel had already been read once by the d4q experiment, so this is a diagnostic read on existing data, not new evidence about fresh roots.
- ✓Label file is the preregistered artifact: 291,890 rows / 8,639 roots — observed: 291890 rows, 8639 roots (train 6,551, calibration 2,088)
- ✓No non-finite loss in any seed; all five seeds completed 30 epochs within 5,400 wall-seconds — observed: 5 of 5 seeds valid; 166.08 wall-seconds total
- ✕Selected seed (best validation top-1, 0xA52E02) held-out top1 >= 0.6 in each half-fold — observed: half1 0.2960, half2 0.3011
- ✕Selected seed (best validation top-1, 0xA52E02) held-out pairwise >= 0.78 in each half-fold — observed: half1 0.6266, half2 0.6259
- ✕Selected seed (best validation top-1, 0xA52E02) held-out regret <= 0.13 in each half-fold — observed: half1 0.3354, half2 0.3379
Recorded metrics
- 0xA52E01
- 0.3026
- 0xA52E02
- 0.2987
- 0xA52E03
- 0.3069
- 0xA52E04
- 0.3132
- 0xA52E05
- 0.3135
- 0xA52E01
- 0.3051
- 0xA52E02
- 0.3281
- 0xA52E03
- 0.3089
- 0xA52E04
- 0.3127
- 0xA52E05
- 0.3027
- 0xA52E01
- 0.3554
- 0xA52E02
- 0.3549
- 0xA52E03
- 0.3520
- 0xA52E04
- 0.3525
- 0xA52E05
- 0.3592
- 0.2987
- 0.3135
- Diagnostic-tier read: the 3,030-root held-out panel (d4q-labels-gate) was already opened once by EX-20260821-afterstate-d4q-stage1; no fresh roots were labelled. The result schema has no diagnostic tier, so it is recorded at the lowest tier (pilot); it is not development-tier evidence.
- The preregistration names only the d4q-labels file, which carries folds train and calibration but no held-out fold; the held-out panel and its origin-seed half-folds were taken from the same run's d4q-labels-gate/d4q-labels.tsv and corpus-gate/roots.tsv, exactly as d4q.py consumed them, which is the reading the protocol's 'train/validation/held-out origins are reused unchanged' clause intends.
- Epoch count was fixed at 30 by the protocol and the final-epoch model was used. Calibration-fold top-1 peaked at 0.352-0.359 between epochs 5 and 10 for every seed and then declined as the training listwise loss kept falling (overfitting); even the best epoch is far below the 0.60 gate, so early stopping would not change the verdict, but the reported held-out numbers are for the overfitted final epoch.
- Loss coefficients not fixed by the hypothesis text were set to 1.0 for the pairwise term (d4-q-clone used 0.35) and the prediction-side softmax temperature was 1 in standardised value units; only the listed configuration is rejected.
- CPU training with 8 torch threads; the EmbeddingBag backward is not bit-deterministic across thread counts, so a re-run reproduces the numbers only approximately.
- Inference cost is from leaf-check on the RUN-A51D-corpus mix-d3 states on a shared host; informational only.