ResultStage D0: fair relabelling of oracle-visited vs matched fair-D4 public states
Stage D0 refuses the H-pool theory on two of its three preregistered criteria.
On this page
- Recorded
No explanation has been written for this record yet.
Technical recordMetrics, gate checks and limitations
Stage D0 refuses the H-pool theory on two of its three preregistered criteria. Pool O was 1,984 public states sampled from 64 oracle games (depth 4, beam 128; 63 of 64 games reached the 500-move cap); 1,271 were matched 1:1 to fair-D4 states on exact (rise phase, occupancy/4, max-height/2) buckets over all 256 comparator seeds (match rate 0.641; 713 O states dropped and counted). Under 32 common public futures with fair D1 continuation at horizon 25, oracle-visited boards are only marginally better than matched fair boards: R_fair(O) = 24.783 vs R_fair(F) = 24.520 (difference +0.263 moves, cluster bootstrap 95% [+0.156, +0.367], same positive sign in both origin-game halves). The preregistered transferable fraction is negative, tau = -0.959 (95% cluster interval [-3.069, -0.390]), because its denominator degenerates at this horizon: R_tape(O) = 24.338 vs R_real(F) = 24.611 (94-95% of both realised remainders were capped at 25), so the realised-lifetime gap the fraction was defined against is -0.274 moves rather than positive. Independently of that degeneracy, the action-quality criterion fails outright: the oracle's own column is fair-top-1 at its own roots less often than fair D4's column at the same roots (0.766 vs 0.814 over 1,271 roots, difference -0.049, 95% [-0.069, -0.029]; on the 194 unique-maximum roots 0.268 vs 0.387). Blocked-flow-band fraction over matched O states is 0.0047 (F: 0.026). Per the preregistration the theory is assessed not-supported-as-tested, the pool-comparison follow-on is not registered, and audit-05's H-pool program closes.
- ✓CHECK gates before any seed is read: privilege boundary (relabel binary links no oracle tape accessor), determinism of R_fair across two runs and thread counts, domain separation of restart streams, mirror invariance of R_fair — observed: ALL GATES PASS in runs/RUN-20260823T191900Z-b9f8f80d/d0/gates.log on probe seeds 0xa5278000-0xa527810f: 0 oracle symbols in d0-relabel, generate byte-identical at 16 vs 4 threads, relabel byte-identical at 1 vs 8 threads across two runs, 0 mirror/metadata/sibling/stream failures on 47 probe states (mirror invariance is exact: the restart plays in the canonical frame), label edits changed no relabel byte
- ✕tau >= 0.25 pooled — observed: tau = -0.9590, 95% cluster interval [-3.0691, -0.3897]; denominator R_tape(O)-R_real(F) = -0.274 moves is degenerate at horizon 25 (94-95% of both realised remainders capped) while the numerator R_fair(O)-R_fair(F) = +0.263 [+0.156, +0.367]
- ✓sign of R_fair(O) - R_fair(F) agrees in both origin-game halves — observed: half 0 (games 0-31): +0.2319; half 1 (games 32-63): +0.3146; both positive
- ✕oracle-column fair-top-1 rate at O roots >= fair-D4-column fair-top-1 rate at the same roots — observed: oracle 0.7655 vs fair D4 0.8143 over all 1,271 matched O roots under the same 32 common futures (no subsampling; ties count for every tied column); difference -0.0488, 95% cluster interval [-0.0691, -0.0288]; strict-maximum roots only: 0.268 vs 0.387
Technical recordRecorded metrics
- -3.0691
- -0.3897
- half0
- -0.8357
- half1
- -1.1761
- mean
- 24.7826
- seStates
- 0.0365
- seOriginGames
- 0.0633
- mean
- 24.5201
- seStates
- 0.0491
- seOriginGames
- 0.1772
- mean
- 24.3375
- seStates
- 0.0904
- seOriginGames
- 0.0827
- mean
- 24.6113
- seStates
- 0.0525
- seOriginGames
- 0.2163
- 0.1561
- 0.3670
- half0
- 0.2319
- half1
- 0.3146
- oStates
- 1,984
- oMatched
- 1,271
- oUnmatched
- 713
- matchRate
- 0.6406
- fStates
- 1,271
- oOriginGames
- 64
- fOriginGames
- 101
- bucketsMatched
- 63
- R_tape_O_cappedAt25Fraction
- 0.9473
- R_real_F_cappedAt25Fraction
- 0.9426
- R_fair_O_survivedHorizonFraction
- 0.9624
- R_fair_F_survivedHorizonFraction
- 0.9167
- oracleGamesCensoredAt500
- 63
- oracleOriginCensoredStates
- 62
- O_matched_blocked
- 0.0047
- O_all_blocked
- 0.0030
- F_blocked
- 0.0260
- roots
- 1,271
- subsampled
- false
- oracleColumnTop1Rate
- 0.7655
- fairD4ColumnTop1Rate
- 0.8143
- difference
- -0.0488
- difference95ClusterBootstrap
- -0.0691
- -0.0288
- oracleEqualsD4ColumnRate
- 0.3895
- uniqueMaximumRoots
- 194
- oracleColumnTop1RateStrictRoots
- 0.2680
- fairD4ColumnTop1RateStrictRoots
- 0.3866
- method
- cluster bootstrap over O origin games carrying matched F partners
- resamples
- 10,000
- seed
- 0xb0071eaf
- clusters
- 64
- sensitivityIndependentClustersTau95
- -9.3496
- 5.1228
- The preregistered tau is ill-conditioned as measured: with 63 of 64 oracle games censored at the 500-move cap and comparator states drawn from the same shape buckets, both realised remainders sit at the 25-move horizon cap (R_tape(O) 24.34, R_real(F) 24.61, 94-95% capped), so the denominator is -0.274 moves instead of the large positive gap the definition assumed. The tau < 0.25 refusal is therefore driven by a degenerate denominator, not by a negative numerator; the numerator (the fair-value advantage of oracle boards) is positive but small, +0.263 moves on a 25-move horizon.
- The fair-top-1 criterion is unaffected by that degeneracy and fails on its own: the oracle's action is fair-best at its own roots significantly less often than fair D4's action at the same roots.
- Match rate 0.641: the 713 dropped O states skew toward buckets fair D4 rarely visits (matched O states are shape-matched by construction), so the numerator is a lower bound on the raw O-vs-F fair-value difference over all oracle states, and the comparison is conditional on shape overlap as preregistered.
- R_fair uses a fixed public D1 continuation and horizon 25 with 32 futures; 92-96% of futures survive the horizon, so both pools are near the measurement ceiling and the +0.263 difference is compressed by censoring at both ends.
- F origin games contribute clustered states (mean 12.6 matched states over 101 of 256 games); uncertainty uses cluster bootstrap over O origin games carrying matched partners, with the independent-clusters sensitivity interval also recorded (it is much wider: [-9.35, +5.12]).
- Diagnostic tier (CHECK): no gameplay evidence and no policy claim; training-role seeds only, opened once.
Recorded against Stage D0: fair relabelling of oracle-visited vs matched fair-D4 public states.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/RS-20260823T205143Z-ead14c9d.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.
Record file: research/results/RS-20260823T205143Z-ead14c9d.json, validated against research/schemas/result-v1.schema.json.
- Run ids
RUN-20260823T195750Z-1cb2f9de
- Contribution ids
CT-20260823T195817Z-5f14e7f7
- Per-game artifact
runs/RUN-20260823T191900Z-b9f8f80d/d0/pools.json(sha256257dadcc26611e593350f1498a7dfce3fe476f9f208d801d1b5a69bdc6312d39, 3255 records)- Artifact manifest
runs/RUN-20260823T191900Z-b9f8f80d/d0/d0-result.json- Machine profiles
research/system-profiles/MACH-20260820T080056Z-376ada90.json