Drop7 Research
← Experiments

Successor SCREEN: paired 256-game test of the reveal-construction leaf terms (two doses and a chain-to-crack bundle) against the unchanged fair D4 at seven strata with the memo engine, after the corpus value gate failed

completedtier SCREENalgorithmicpublic-developmentpublic-policyEX-20260823-reveal-construction-screen-v2-63d73b6a
No explanation has been written for this experiment yet. Add web/content/research/EX-20260823-reveal-construction-screen-v2-63d73b6a.mdx and it will appear here. The registered protocol is shown below.

The registered protocol

Successor to EX-20260823-reveal-construction-screen-371fd638, whose seed-free corpus VALUE gate failed (RS-20260823T110000Z-baecb816: aligned_double_hit held-out partial r -0.0443, incremental R^2 +0.00047) while its ACTION statistic passed (60.47% of live same-wave double-hit setups uncollected by the corpus depth-4 policy within two moves). The independent review written before that result (kimi-k3-prereg-review.md section 3) argued that a partial-correlation kill is a value criterion that can reject a term whose purpose is to re-rank sibling actions at the horizon boundary. This experiment therefore tests the owner's hypothesis in play, with no corpus gate, and discloses that the corpus result was read first and points the wrong way. With the fast fair-D4 search at seven strata, the bit-exact one-entry leaf memo, cache 60,000, work bound worstCaseWork(4,7)+1, terminal utility -1,000,000, policy seed 0xd7075eed, 2,000-move cap, four arms play the same 256 ordered seeds 0xa52d0200-0xa52d02ff: 'frozen' (zero extra weights, gated bit-identical to FastSearch); 'A300' (aligned_double_hit +300, the preregistered primary); 'A900' (aligned_double_hit +900, a dose arm added because CHECK-gate coverage at +300 was 0.3-1.9% of decisions on probe seeds, near the rarity bar); 'B' (aligned_double_hit +300 ungated + chain_to_crack_cracked_gated +150 + chain_to_crack_solid_gated +300, where the gate is 1/0.5/0 at max height <=4/5/>=6). Term definitions are frozen in approaches/lifetime-objective/chain-reveal-leaf/extra-terms.hpp (sha256 recorded in expectedArtifacts at freeze). A shadow unchanged search runs at every decision of the non-frozen arms to measure coverage. Primary contrast: A300 minus frozen. Secondary, declared underpowered: A900 minus frozen (dose), B minus A300 (chain-to-crack bundle).

Candidate
fast-d4s7-memo-reveal-construction
approaches/lifetime-objective/chain-reveal-leaf/run.cpp
Comparator
fair-d4
approaches/fair-expectimax/reference/fair-only-depth4.cpp

Primary metric

paired mean whole-game score delta, A300 minus frozen, 256 games, depth 4 seven strata memo engine

Statistical unit: whole-game

Pass criteria, fixed in advance

  • CHECK gates passed at the frozen term source after the review fixes: zero-weight leaf bit parity, search parity with FastSearch (column, work, nodes, cache hits, completed depth) at d4s5 and d4s7, mirror invariance, thread determinism at live weights, metadata blindness, memo-vs-fresh identity; the extra-term signature takes only (board, moves_remaining, scratch).
  • Every artifact carries seedLease SL-20260823T110000Z-a52d0200 and dataRole public-development; incompleteDecisionsTotal 0 and illegalDecisionsTotal 0 in every arm.
  • A300 minus frozen: bootstrap 95% lower bound > 0 AND Student-t 95% lower bound > 0.
  • A300 minus frozen: candidate Q25 >= comparator Q25 AND both halves > 0.
  • A300 coverage >= 2% of its decisions; below 2% the A300 result is inconclusive by rarity regardless of the deltas, and the A900 arm is then read as the primary dose with the same gate but recorded as 'pass at a dose chosen after a rarity null', never as a clean pass.
  • Mechanism check (reported, not a pass condition): reveals per move of the passing arm exceed frozen's; a score gain with reveals per move flat is recorded as 'pass, mechanism unsupported'.

On pass: Record valid + pass at public-development tier; register a replication on the reserved 0xa52d0300-0xa52d03ff block as the next experiment before any dose sweep; register the arm as a playground leaderboard policy. No tier above SCREEN is claimed.

On fail: Record valid + fail for these exact terms and doses; the theory is not-supported-as-tested at public-development tier; report A900, B, coverage and flow as diagnostics; open no further cohort. If all three non-frozen arms lose with reveals per move not higher than frozen, the owner's H2 reveal-construction mechanism is recorded as closed at this tier alongside fair-reveal-reward and t1-exposures, and the predeclared H1 fallback (entombed_high soft penalty at -150, 128 games, declared a non-measurement for score) is registered as a successor, not run here.

Data and reuse

CHECK gates read previously opened development probe seeds 0xa5278000-0xa52784ff only. The corpus value statistics on runs/RUN-A51D-corpus/all.states (training role) were read once under the predecessor experiment and are disclosed in the hypothesis; they are not a gate here. The screen opens the never-read lease 0xa52d0200-0xa52d02ff exactly once; 0xa52d0300-0xa52d03ff of the same lease stays reserved for a replication under a successor protocol. The predecessor's cancelled lease 0xa52d0000-0xa52d01ff was never read and is not regranted.

seed leases: SL-20260823T110000Z-a52d0200

What happened

valid run · outcome: failnot-supported-as-testedpublic-developmentRS-20260823T131226Z-16564ed9

Paired 256-game SCREEN on fresh seeds 0xa52d0200-0xa52d02ff, fair D4 at seven strata with the memo engine. The preregistered primary arm A300 (aligned_double_hit +300) and the bundle arm B changed the unchanged search's column in 0.69% and 0.70% of decisions at the 32-game rarity check, below the 1% rule, and were stopped as no-measurements (partials: A300 -4,298 on 39 games, B +18,213 on 36 games, both far inside their floors). The dose arm A900 (aligned_double_hit +900), read as primary per the gate text, completed 256 games: mean 389,749 vs 386,545, paired delta +3,204 points (one-sided 95% bootstrap lower bound -26,860; Student-t -27,863; detection floor 30,957; paired sd 301,098), moves 112.55 vs 111.59 (+0.97), W-T-L 99-53-104, Q25 +18,534 (non-regression met), halves +48,762 and -42,354 (opposite signs: fail), coverage 598/28,814 = 2.08% (measured, not rarity). Predeclared mechanism directions were absent: cover reveals per move 1.1520 vs 1.1536, numbered clears per move 2.0541 vs 2.0551, occupancy 23.32 vs 23.22. Gate: FAIL. The term re-ranks about one decision in fifty at +900 and those re-rankings add no reveal flow; the score delta is a non-measurement for effects under about 31,000 points, but the flat flow statistics, whose paired noise is far smaller, reject the mechanism itself. The frozen arm is also the largest fresh-seed fair-D4 seven-stratum cohort on record: 386,545 points, 111.59 moves, 2.0551 clears and 1.1536 reveals per move over 256 never-read games, consistent with the 64-game 398,498 reference.

What it had to pass
  • CHECK gates at the frozen term source — observed: 0 mismatches (gates.log, gates-phase3.log)
  • Artifacts carry the lease and role; 0 incomplete and 0 illegal decisions — observed: SL-20260823T110000Z-a52d0200 / public-development; 0 / 0
  • A300 coverage >= 2% (else inconclusive by rarity, A900 read as primary) — observed: 0.69% at the 32-game rarity check; arm stopped; A900 read as primary at 2.08%
  • Primary minus frozen: bootstrap and Student-t one-sided 95% lower bounds > 0 — observed: A900: +3,204; LB -26,860 / -27,863
  • Primary: Q25 non-regression AND both halves > 0 — observed: Q25 +18,534 (met); halves +48,762 / -42,354 (opposite signs)
  • Mechanism (reported): reveals per move above frozen — observed: 1.1520 vs 1.1536
Recorded metrics
A900_minus_frozen
games
256
meanScoreCandidate
389,749
meanScoreReference
386,545
pairedDelta
3,204
bootstrapLB95
-26,860
studentTLB95
-27,863
detectionFloor
30,957
pairedSd
301,098
winTieLoss
99-53-104
medianCandidate
326,048
medianReference
313,466
q25Candidate
215,187
q25Reference
196,653
q25Delta
18,534
minCandidate
102,610
minReference
102,894
maxCandidate
1,976,127
maxReference
1,656,350
halves
  1. 48,762
  2. -42,354
movesCandidate
112.5500
movesReference
111.5900
movesDelta
0.9700
movesBootstrapLB95
-7.1700
flow
A900
numberedClearsPerMove
2.0541
coverRevealsPerMove
1.1520
meanOccupiedCells
23.3200
maxChainDepth
14
frozen
numberedClearsPerMove
2.0551
coverRevealsPerMove
1.1536
meanOccupiedCells
23.2200
maxChainDepth
12
coverage
A900
divergentDecisions
598
shadowDecisions
28,814
rate
0.0208
A300_at_stop
divergentDecisions
27
shadowDecisions
4,052
rate
0.0067
completedGames
39
B_at_stop
divergentDecisions
26
shadowDecisions
3,745
rate
0.0069
completedGames
36
stoppedArms
A300
games
39
pairedDelta
-4,298
bootstrapLB95
-41,726
B
games
36
pairedDelta
18,213
bootstrapLB95
-35,413
integrity
incompleteDecisions
0
illegalDecisions
0
censoredGames
0
memoHitRate
frozen
0.6483
A900
0.6484
maxWork
11,892,399
wallSecondsStage2
5,299
frozenArmFreshBaseline
games
256
meanScore
386,545
meanMoves
111.5900
numberedClearsPerMove
2.0551
coverRevealsPerMove
1.1536
gamesAtOrAboveMillion
Limitations
  • A score null inside a 30,957-point floor is a non-measurement for effects below that size; the rejection rests on the flat flow statistics and the opposite-sign halves, not on the score delta.
  • The preregistered primary dose (+300) could not be measured: it changed under 1% of decisions and was stopped by the rarity rule; the +900 dose was read as primary under the gate's own clause and is therefore a pass-or-fail at a dose chosen after a rarity null.
  • The runner has no per-arm stop; the operator killed the four-arm run after the 32-game check and relaunched frozen + A900 on the same seeds (deterministic; the frozen arm reproduces stage-1 games exactly). Stage-1 partials are retained under stopped-at-32/.
  • The corpus value gate for this term failed before this screen (RS-20260823T110000Z-baecb816) and was disclosed in the protocol; this result is the in-play test that the blind review asked for.
  • Single cohort, public-development tier; not independently replicated.
  • The frozen-arm fresh baseline is a by-product and has not been registered as a benchmark manifest.