ResultPilot iteration 2: K=64-scenario action-complete H40 corpus with fresh held-out origins
K=64 improved label stability from 0.246 to 0.446 mean split-half Spearman on fresh held-out origins, but still below the frozen 0.5 floor, so the preregistered verdict is again inconclusive.
On this page
No explanation has been written for this record yet.
Technical recordMetrics, gate checks and limitations
K=64 improved label stability from 0.246 to 0.446 mean split-half Spearman on fresh held-out origins, but still below the frozen 0.5 floor, so the preregistered verdict is again inconclusive. Quantile calibration passed (0.807 coverage of the nominal 76.5% interval). Against the still-noisy target the model trailed fair D4 consistently in both half-folds (pooled top-1 0.342 vs 0.433; pairwise 0.637 vs 0.697; regret 0.303 vs 0.232) while beating its D1 continuation teacher (top-1 0.342 vs 0.299). Training stopped at epoch 15 of 20 on the GPU time budget with the ranking loss still decreasing. No gameplay tier was opened.
- ✓Successor-closed corpus (all legal siblings x 64 scenarios) — observed: 6,167,936 rows over 14,228 roots across both corpora; completeness 1.0
- ✕Label stability floor: mean scenario-half (32 vs 32) Spearman >= 0.5 — observed: 0.4462 on 2,750 fresh held-out roots; inconclusive per the frozen rule
- ✕Model top-1 >= D4 top-1 - 0.02 on each half-fold — observed: half1 0.3401 vs 0.4365; half2 0.3442 vs 0.4299 (moot given the stability failure)
- ✕Model pairwise >= D4 pairwise - 0.02 on each half-fold — observed: half1 0.6341 vs 0.6970; half2 0.6391 vs 0.6974 (moot)
- ✕Model regret <= D4 regret + 0.02 on each half-fold — observed: half1 0.3031 vs 0.2388; half2 0.3028 vs 0.2247 (moot)
- ✓Nominal 76.5% quantile interval coverage within [0.66, 0.87] — observed: 0.8069
- ✓All CHECK-tier tests pass before any label is inspected — observed: SELFTEST OK (9 checks) after the K-parameterization change
Technical recordRecorded metrics
- Labels remain H40 returns under a phase-greedy D1 continuation (a weak fixed teacher).
- Model undertrained: 15 of 20 epochs at the GPU budget stop.
- Iteration-1 held-out roots were folded into training data here; the gate read only fresh origins 0x5da70100-0x5da7013f.
- Single machine profile; FP32 on the shared-memory iGPU.
Recorded against Pilot iteration 2: K=64-scenario action-complete H40 corpus with fresh held-out origins.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/RS-20260820T114500Z-2b7c9e31.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.
Record file: research/results/RS-20260820T114500Z-2b7c9e31.json, validated against research/schemas/result-v1.schema.json.
- Run ids
RUN-20260820T090411Z-73e93859
- Contribution ids
CT-20260820T140540Z-f8458dc0
- Per-game artifact
–- Machine profiles
research/system-profiles/MACH-20260820T080056Z-376ada90.json