On this page
Dates
Recorded
Record idRS-20260821T094500Z-1a7e3c55

No explanation has been written for this record yet.

Technical recordMetrics, gate checks and limitationsRS-20260821T094500Z-1a7e3c55
valid runoutcome: failnot-supported-as-testedtier: pilotRS-20260821T094500Z-1a7e3c55

Full training does not rescue the afterstate model; it overfits. The model was trained to 22 epochs on the 2M-row K=256 subsample (44M row-updates, 2x iteration 3, cosine schedule completed, ranking loss 0.584 vs iteration 3's 0.630). On the SAME held-out roots as iteration 3 (corpus-C, a labeled diagnostic reuse), the fully-trained model ranks WORSE than iteration 3's 11-epoch checkpoint (top-1 0.361 vs 0.424, pairwise 0.658 vs 0.685, regret 0.281 vs 0.241) - training loss improved while held-out ranking degraded, a textbook overfitting signature against the D1-continuation H40 labels. The frozen override gate on fresh corpus-E (2,867 roots, 1,106 near-tie eligible, 41% override rate) then FAILED: eligible-root regret half1 0.2289 vs D4 0.2517 (+0.0228) but half2 0.2778 vs 0.2481 (-0.0297, the override is actively harmful there), pooled +0.0018. Stability (0.824), calibration (0.800), determinism (byte-identical) all passed. Conclusion: the model's limitation is not training completeness but generalization to held-out roots under a weak D1 teacher; the direct-override use of this model family is closed per the frozen failure action.

What it had to pass
  • Training completes 22 epochs on the 2M-row subsample within budget (44M row-updates, 2x iteration 3) — observed: 22 epochs, cosine schedule completed, ~5.5h GPU
  • Eligible-root override regret <= D4 regret - 0.01 in EACH half-fold — observed: half1 +0.0228 (passes); half2 -0.0297 (override harmful); pooled +0.0018
  • Override rate >= 5% of eligible roots — observed: 41.1%
  • Decisive-root label stability >= 0.5 — observed: 0.8244
  • Quantile coverage within [0.66, 0.87] — observed: 0.7997
  • Gate report byte-identical across two runs — observed: byte-identical
Technical recordRecorded metricsRS-20260821T094500Z-1a7e3c55
epochsCompleted
22
rowUpdates
44,000,000
finalRankLoss
0.5838
diagnosticTop1VsIter3
0.3611 vs 0.4245 (same roots)
diagnosticRegretVsIter3
0.2807 vs 0.2408 (same roots)
overrideEligibleRoots
1,106
overrideRate
0.4114
overrideRegretGainHalf1
0.0228
overrideRegretGainHalf2
-0.0297
overrideRegretGainPooled
0.0018
labelStabilityDecisiveSpearman
0.8244
quantileIntervalCoverage
0.7997
Limitations
  • The diagnostic comparison to iteration 3 reuses corpus-C (iteration 3's held-out), a labeled diagnostic; the frozen override gate used fresh corpus-E.
  • The model family evaluated is the compact 3.4M-param ResNet over D1-continuation H40 labels; the result does not bound a stronger-teacher or different-architecture variant.
  • The iGPU is compute-bound for this model (~2-3k rows/s FP32); larger-scale training was not attempted within the overnight budget.
  • Two killed training attempts (tooling: memory blowup, output buffering) preceded the recorded run; they produced no artifacts and are disclosed in the run record.

Recorded against Full training of the K=256 afterstate model, then the frozen top-two override rule on fresh origins.

Agent contextHow to extend this record

To add a reader-facing explanation, write web/content/research/RS-20260821T094500Z-1a7e3c55.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.

Record file: research/results/RS-20260821T094500Z-1a7e3c55.json, validated against research/schemas/result-v1.schema.json.

Run ids
  • RUN-20260821T084551Z-a5050ee5
Contribution ids
  • CT-20260820T140540Z-f8458dc0
Per-game artifact
runs/RUN-20260821T084551Z-a5050ee5/gate/override-gate-report.json (sha256 67b15841e6cf3a580fca03df17c4aefc6d12687c42a526a216609bc24aab5863, 1 records)
Machine profiles
  • research/system-profiles/MACH-20260820T080056Z-376ada90.json