On this page
Dates
Recorded
Record idRS-20260820T184500Z-63c0a8e2

No explanation has been written for this record yet.

Technical recordMetrics, gate checks and limitationsRS-20260820T184500Z-63c0a8e2
valid runoutcome: failnot-supported-as-testedtier: pilotRS-20260820T184500Z-63c0a8e2

The calibrated top-two near-tie override of fair D4 by the frozen iteration-3 afterstate model FAILED its frozen gate by a narrow margin, with genuinely positive signal. On 2,689 fresh held-out roots, 1,030 (38.3%) were near-tied at the frozen 500-point threshold and the override fired on 37.0% of those. On eligible roots the override reduced mean normalized regret versus unchanged D4 in both half-folds (half1 0.2484 vs 0.2835, +0.0351; half2 0.2343 vs 0.2418, +0.0075; pooled +0.0214), but the frozen criterion required at least +0.01 in EACH half-fold, and half2 fell short. On the 84 decisive eligible roots the improvement was larger (+0.0426: 0.0880 vs 0.1306). Label stability passed (decisive Spearman 0.814), model quantile calibration passed (0.861), the gate is deterministic (byte-identical repeated reports), and the corpus is successor-closed. Interpretation: the conservative-override MECHANISM works as designed (fires often, calibrated, positive in both halves) and this is the first preregistered held-out test in this repository where a learned model's intervention improved on fair D4 at all; the CHECKPOINT's ranking quality is the bottleneck, consistent with the model being undertrained (11 of 20 epochs) and trained on a weak D1-continuation teacher. Per the frozen failure action, no further override variants of this same checkpoint are opened; a retry requires new evidence such as a fully trained or stronger-teacher model under the same frozen rule.

What it had to pass
  • Successor-closed corpus (all legal siblings x 256 scenarios) — observed: 4,858,880 rows over 2,793 roots; completeness 1.0
  • Label stability floor: decisive-root scenario-half Spearman >= 0.5 — observed: 0.8144
  • Eligible-root override regret <= D4 regret - 0.01 in EACH half-fold — observed: half1 +0.0351 (passes); half2 +0.0075 (below the frozen 0.01 margin); pooled +0.0214
  • Override rate >= 5% of eligible roots — observed: 37.0%
  • Quantile interval coverage within [0.66, 0.87] — observed: 0.8606
  • Gate script deterministic (byte-identical repeated reports) — observed: two runs byte-identical after moving wall time out of the report
  • All existing CHECK-tier self-tests pass before any label is inspected — observed: SELFTEST OK (9 checks)
Technical recordRecorded metricsRS-20260820T184500Z-63c0a8e2
roots
2,689
eligibleRoots
1,030
nearTieRate
0.3830
overrideRateEligible
0.3699
d4RegretEligiblePooled
0.2629
overrideRegretEligiblePooled
0.2414
regretGainHalf1
0.0351
regretGainHalf2
0.0075
regretGainDecisive
0.0426
d4RegretWholeSet
0.1932
overrideRegretWholeSet
0.1850
labelStabilityDecisiveSpearman
0.8144
quantileIntervalCoverage
0.8606
Limitations
  • The evaluation target is the H40 phase-greedy-D1 continuation outcome - the same label family the model was trained to predict; D4 is compared against the same target, but the offline gate cannot measure real gameplay interaction, which only a SCREEN tier can.
  • The model was undertrained (11 of 20 epochs) and used a weak D1-continuation teacher; both are documented as the likely bottleneck and motivate any retry.
  • The near-tie threshold (500) and regret margin (0.01) were frozen choices; the half2 miss (0.0075) is close to the margin and the result should be read as a narrow failure, not as evidence of no effect.
  • Single machine profile; FP32 on the shared-memory iGPU.

Recorded against Offline gate: calibrated top-two near-tie override of fair D4 by the frozen afterstate model.

Agent contextHow to extend this record

To add a reader-facing explanation, write web/content/research/RS-20260820T184500Z-63c0a8e2.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.

Record file: research/results/RS-20260820T184500Z-63c0a8e2.json, validated against research/schemas/result-v1.schema.json.

Run ids
  • RUN-20260820T175726Z-f4f86d62
Contribution ids
  • CT-20260820T140540Z-f8458dc0
Per-game artifact
runs/RUN-20260820T175726Z-f4f86d62/gate/override-gate-report.json (sha256 de6e7b70dacc8ccdfa95c8486d941b9249ce28675e96daa45d4460618f7582df, 1 records)
Machine profiles
  • research/system-profiles/MACH-20260820T080056Z-376ada90.json