ResultPilot iteration 3: K=256-scenario corpus, fresh held-out origins, full training
With K=256 aligned scenarios the label-stability floor passed decisively (decisive-root split-half Spearman 0.818, unconditioned 0.638, both >= 0.5), so for the first time in this line the ranking target is stable enough to certify a verdict.
On this page
No explanation has been written for this record yet.
Technical recordMetrics, gate checks and limitations
With K=256 aligned scenarios the label-stability floor passed decisively (decisive-root split-half Spearman 0.818, unconditioned 0.638, both >= 0.5), so for the first time in this line the ranking target is stable enough to certify a verdict. The verdict is a valid negative for the direct-play configuration: the distributional afterstate ranker trailed exact fair D4 on every preregistered comparison in both fresh held-out half-folds (pooled top-1 0.424 vs 0.499; pairwise 0.685 vs 0.737; regret 0.241 vs 0.178). On the 411 decisive roots (spread > 20k) the gap narrowed but persisted (top-1 0.526 vs 0.584; regret 0.082 vs 0.064). The model clearly beat its D1 continuation teacher (top-1 0.424 vs 0.319) and its quantile head is calibrated (0.864 coverage). Conclusion: a compact action-free afterstate network trained on successor-closed H40 D1-continuation labels is a viable, calibrated long-horizon evaluator, but it is NOT a viable direct replacement for 4-ply expectimax; any future use must be as a companion signal inside search, and the same ranking gate applies. No gameplay tier was opened.
- ✓Successor-closed corpus (all legal siblings x 256 scenarios) — observed: 24,270,592 rows over 14,009 roots across both corpora; completeness 1.0
- ✓Label stability floor: decisive-root scenario-half (128 vs 128) Spearman >= 0.5 — observed: 0.8181 on 411 decisive roots; unconditioned 0.6379 on 2,523 roots
- ✕Model top-1 >= D4 top-1 - 0.02 on each half-fold — observed: half1 0.3981 vs 0.4776; half2 0.4439 vs 0.5141
- ✕Model pairwise >= D4 pairwise - 0.02 on each half-fold — observed: half1 0.6828 vs 0.7280; half2 0.6868 vs 0.7429
- ✕Model regret <= D4 regret + 0.02 on each half-fold — observed: half1 0.2496 vs 0.1870; half2 0.2344 vs 0.1720
- ✓Nominal 76.5% quantile interval coverage within [0.66, 0.87] — observed: 0.8635
- ✓All CHECK-tier tests pass before any label is inspected — observed: SELFTEST OK (9 checks)
Technical recordRecorded metrics
- Labels are H40 returns under a phase-greedy D1 continuation (a weak fixed teacher); a stronger-teacher corpus was not tested and might shift the verdict.
- Model undertrained: 11 of 20 epochs at the 2h GPU budget stop; ranking loss was still decreasing.
- Compact 3.4M-parameter ResNet; capacity and input resolution (single afterstate, no root context) were not scaled.
- Roots are harvested from fair-D1 games; deployment-distribution roots were not represented.
- Single machine profile; FP32 on the shared-memory iGPU.
Recorded against Pilot iteration 3: K=256-scenario corpus, fresh held-out origins, full training.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/RS-20260820T142500Z-8f4a2d17.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.
Record file: research/results/RS-20260820T142500Z-8f4a2d17.json, validated against research/schemas/result-v1.schema.json.
- Run ids
RUN-20260820T111656Z-f771f173
- Contribution ids
CT-20260820T140540Z-f8458dc0
- Per-game artifact
–- Machine profiles
research/system-profiles/MACH-20260820T080056Z-376ada90.json