ResultStronger-teacher (D2 continuation) afterstate corpus: ranking gate and frozen override rule
The stronger-teacher hypothesis fails as tested.
On this page
- Recorded
No explanation has been written for this record yet.
Technical recordMetrics, gate checks and limitations
The stronger-teacher hypothesis fails as tested. A successor-closed corpus relabeled with a fair-D2 (five-sample) continuation teacher at K=64 (2.88M rows over 6,535 roots; partial at the generator's 4h default wall stop) trained a model that, on the fresh 0x5da70500 ranking gate against the fixed D1-continuation H40 target, reached top-1 0.3365 - far below fair D4's 0.5020, below the D1-teacher model's 0.4245 from iteration 3, and only at exact-D1's own 0.3339. The frozen criterion (top-1 >= 0.4616 on each half, i.e. closing half the iteration-3 gap to D4) failed by a wide margin in both halves (0.342, 0.331). The frozen override gate on fresh 0x5da70600 roots also failed (eligible-root regret: half1 -0.0089, half2 +0.0110, pooled +0.0012). IMPORTANT CONFOUND, disclosed: the D2-teacher model was evaluated against D1-continuation outcomes (frozen for comparability with iteration 3), so part of its regression may reflect the teacher/target mismatch rather than teacher quality alone. Read narrowly, the result says a stronger-teacher corpus did not produce a better ranker of the fixed public-continuation outcome, and the afterstate line's ranking deficit is robust to the teacher choice within the tested configurations.
- ✓Corpus successor-closed at K=64 with the D2 teacher, completeness 1.0 — observed: per-root completeness 1.0 over 6,535 fully labeled roots; the corpus is partial (6,535 of 8,192 planned) at the generator's 4h default wall stop
- ✕Ranking gate: model top-1 >= 0.4616 on each half-fold — observed: half1 0.3418, half2 0.3314
- ✕Override gate: eligible-root override regret <= D4 regret - 0.01 in EACH half-fold, rate >= 5% — observed: half1 -0.0089, half2 +0.0110, pooled +0.0012; override rate 35.9%
- ✓Decisive-root label stability >= 0.5 for both evaluation corpora — observed: 0.8058 (override corpus), 0.8010 (ranking corpus)
- ✓Quantile coverage within [0.66, 0.87] — observed: 0.7244
Technical recordRecorded metrics
- Teacher/target mismatch confound: the model was trained on D2-continuation outcomes but evaluated against D1-continuation outcomes (frozen for comparability with iteration 3). A matched D2-continuation target was not generated (cost); the regression may partly reflect the mismatch.
- The D2 corpus is partial (6,535 of 8,192 roots) at the generator's 4h default wall stop.
- The compact 3.4M-param architecture and the D1-harvested root distribution are unchanged from iteration 3.
- Single machine profile; FP32 on the shared-memory iGPU.
Recorded against Stronger-teacher (D2 continuation) afterstate corpus: ranking gate and frozen override rule.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/RS-20260821T134500Z-4b9d2f68.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.
Record file: research/results/RS-20260821T134500Z-4b9d2f68.json, validated against research/schemas/result-v1.schema.json.
- Run ids
RUN-20260821T084812Z-0d366205
- Contribution ids
CT-20260820T140540Z-f8458dc0
- Per-game artifact
runs/RUN-20260821T084812Z-0d366205/gate/override-gate-report.json(sha256424970572fc11c652e3d28e5ef1c2bd5340ea49123ad81c521842afe73b7354b, 1 records)- Machine profiles
research/system-profiles/MACH-20260820T080056Z-376ada90.json