ResultPilot: action-complete H40 afterstate corpus and distributional ranker offline gate
The successor-closed corpus pipeline works exactly as designed (11,379 roots, 616,048 sibling labels, 100% action completeness, all CHECK tests pass), but the frozen label-stability floor failed: with K=8 aligned scenarios and an H40 phase-greedy-D1 continuation, the split-half Spearman correlation of within-root action rankings is 0.246, far below the 0.50 floor.
On this page
- Recorded
No explanation has been written for this record yet.
Technical recordMetrics, gate checks and limitations
The successor-closed corpus pipeline works exactly as designed (11,379 roots, 616,048 sibling labels, 100% action completeness, all CHECK tests pass), but the frozen label-stability floor failed: with K=8 aligned scenarios and an H40 phase-greedy-D1 continuation, the split-half Spearman correlation of within-root action rankings is 0.246, far below the 0.50 floor. Per the preregistered gate the outcome is inconclusive: the scenario-mean H40 outcome ranking is too noisy at K=8 to certify any model's ranking. Diagnosis: within-action scenario std is ~21k points while the median between-action gap is ~7.1k, so the target ranking is noise-dominated; noise math predicts K=64 (split-half 32v32) lifts stability to roughly 0.7. For the record, against this noisy target the model trailed fair D4 (top-1 0.247 vs 0.330 pooled; pairwise 0.571 vs 0.637), and D4 beat D1 (0.330 vs 0.254), consistent with the historical ranking of those searches. No gameplay tier was opened.
- ✓Successor-closed corpus: every non-trivial root labeled for every legal sibling under all 8 scenarios — observed: 616,048 rows over 11,379 roots; completeness 1.0 by construction and verified by self-test legality checks
- ✕Label stability floor: mean scenario-half Spearman >= 0.5 on held-out roots — observed: 0.2457 mean, 0.286 median over 2,470 held-out roots; per the frozen rule this makes the outcome inconclusive, not a pass and not a theory rejection
- ✕Model top-1 >= fair D4 top-1 - 0.02 on each held-out half-fold — observed: half1 model 0.2478 vs D4 0.3296; half2 model 0.2462 vs D4 0.3239 (moot given the stability failure)
- ✕Model pairwise accuracy >= fair D4 pairwise - 0.02 on each half-fold — observed: half1 model 0.5678 vs D4 0.6413; half2 model 0.5745 vs D4 0.6322 (moot given the stability failure)
- ✕Model normalized regret <= fair D4 regret + 0.02 on each half-fold — observed: half1 model 0.4193 vs D4 0.3318; half2 model 0.4076 vs D4 0.3418 (moot given the stability failure)
- ✕80% quantile interval coverage within [0.70, 0.90] — observed: 0.631 coverage of the nominal 76.5% outer-quantile interval on held-out afterstates
- ✓All CHECK-tier mechanics, legality, determinism, reflection, information-boundary, and resource-bound tests pass before any label is inspected — observed: build/afterstate/self-test prints SELFTEST OK (9 checks); make test (TypeScript, native, parity) also passes
Technical recordRecorded metrics
- Labels are H40 returns under a phase-greedy D1 continuation: a fixed, weak public teacher. Even a perfectly stable version of this target may not transfer to strong-play rankings.
- Roots are harvested from fair-D1 games, so the state distribution is D1's, not the deployment policy's.
- The model trailed D4 against the noisy target; with stability 0.246 it is impossible to say how much of that gap is real.
- The protocol text was authored before data access but the researchctl freeze hash was computed after the run; the frozen content did not change during the run.
- Single machine profile; FP32 training on the shared-memory iGPU.
Recorded against Pilot: action-complete H40 afterstate corpus and distributional ranker offline gate.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/RS-20260820T094500Z-5c1e9a04.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.
Record file: research/results/RS-20260820T094500Z-5c1e9a04.json, validated against research/schemas/result-v1.schema.json.
- Run ids
RUN-20260820T082542Z-7866d15c
- Contribution ids
CT-20260820T140540Z-f8458dc0
- Per-game artifact
–- Machine profiles
research/system-profiles/MACH-20260820T080056Z-376ada90.json