Reveal construction: pricing the aligned same-wave double hit on a solid gray, and chains that terminate on a cover, raises reveal flow and lifetime
web/content/research/TH-20260823-reveal-construction-leaf-e8f7c78d.mdx and it will appear here. The registered record is shown below.The registered record
Claim
Adding to the frozen fair-D4 leaf a term that prices the JOINT same-wave readiness of two support-disjoint numbered neighbours of a solid gray (the cheapest reveal in the game: two hits in one cascade, no cracked intermediate), optionally with a second gray-terminating release step (a chain whose release delivers a hit to a cover) and a danger gate that switches these construction terms off when max height exceeds 4, increases cover reveals per move and mean paired lifetime and score against the unchanged search on fresh development games at depth 4 / seven strata, by more than a 256-game paired cohort's detection floor.
Mechanism
Owner's hypotheses (2026-08-23): (H2) the search does not try to maximise gray reveals, and lining up a chain reaction that fully exposes a gray disc is especially valuable; (H1) the policy should track discs needing resolution and build long chains when no emergency is pending. Kimi K3's design (runs/RUN-20260823T091530Z-cbe65468/kimi-k3-theory-design.md) locates what is structurally missing rather than mis-weighted: the chance model already enumerates reveals inside the 4-ply horizon and the leaf already pays ~+600 to +5,000 per reveal through cell counts and height relief; every prior reveal arm priced the EVENT (fair-reveal-reward) or the MARGINAL neighbour readiness (t1-exposures, -67,719) and achieved fewer reveals. No term expresses the joint same-wave condition of fast-engine.hpp:349-357 (solid_exposure is best*0.35+second*0.65 of marginals), no term extends latent_chain_potential a second release step, and no sum of linear terms expresses the 'no emergency -> build' product. The term pays for a configuration one move before a reveal, at the horizon boundary, adds no chain-depth incentive (finding-06: chain seeking trades lifetime), and is hard to farm without being one move from a reveal. H1's ledger clause is demoted: the entombed-3..7 count has held-out partial r -0.023 beyond the leaf and the hard action filter changed 1.97% of decisions (RS-20260822T233343Z-12becce9); it survives only as a declared-diagnostic fallback soft penalty.
What would prove it wrong
- Seed-free corpus gate (runs/RUN-A51D-corpus/all.states, entombed-discs pipeline, thresholds fixed before the dump is read): the candidate term is killed if its held-out partial correlation with log1p remaining lifetime beyond the 18 leaf features, occupancy and rise clock is below +0.05, or incremental R^2 below 0.005, or prevalence below 5% of positions, or (aligned_double_hit) the fraction of live setups (joint readiness >= 0.5) the depth-4 policy fails to convert into a reveal within 2 moves is below 30%. Gameplay gate: arm A (aligned_double_hit +300) minus frozen on 256 fresh paired games at d4s7 has a one-sided 95% bootstrap or Student-t lower bound at or below zero, or Q25 regression, or opposite-sign halves, or root-argmax coverage below 2% (inconclusive by rarity). Predeclared mechanism directions: reveals/move up, clears/move not down, mean occupancy not up; a score gain with reveals/move flat is not support for this mechanism. Designer's prior stated before the run: P(arm A passes) about 0.15; modal outcome a fourth reveal-pricing null.
Experiments that test it
- Seed-free corpus gate, then a paired 256-game SCREEN of fair D4 (seven strata, memo engine) with reveal-construction leaf terms against the unchanged leaffast-d4s7-memo-reveal-construction vs fair-d4 · SCREEN · completed
- Successor SCREEN: paired 256-game test of the reveal-construction leaf terms (two doses and a chain-to-crack bundle) against the unchanged fair D4 at seven strata with the memo engine, after the corpus value gate failedfast-d4s7-memo-reveal-construction vs fair-d4 · SCREEN · completed
Results recorded against it
Seed-free corpus gate for the reveal-construction leaf terms, read once on the depth-4 non-explored subset of runs/RUN-A51D-corpus/all.states (62,831 rows, 768 games, whole-origin split, base held-out R^2 0.6483): the gating term aligned_double_hit FAILS the preregistered four-part gate. Held-out partial correlation with log1p remaining lifetime beyond the 18 leaf features, occupancy and rise clock is -0.0443 (threshold >= +0.05) and incremental R^2 +0.00047 (>= 0.005); prevalence 19.75% and the uncollected-setup rate 60.47% (56.51% excluding ambiguous-empty outcomes; 6,969 setups) pass. chain_to_crack_cracked (partial r -0.0325, R^2 +0.00026, prevalence 28.92%) and chain_to_crack_solid (partial r -0.0385, prevalence 1.77%, uncollected 62.85%) also fail; entombed_high on depth 4 alone reads partial r -0.0595, incremental R^2 +0.00125. Per the protocol the 256-game screen was NOT run and no leased seed was read. All CHECK gates passed at the final term source (0 bit, parity, mirror, determinism, metadata mismatches at d4s5 and d4s7; runs/RUN-20260823T091530Z-cbe65468/gates.log). The value-criterion part of the gate rejects these terms as lifetime predictors; the action-level statistic says the depth-4 (s5, epsilon 0.03) behaviour policy leaves 60% of live same-wave double-hit setups uncollected within two moves.
- ✕aligned_double_hit held-out partial r >= +0.05 — observed: -0.0443
- ✕aligned_double_hit incremental held-out R^2 >= 0.005 — observed: 0.00047
- ✓aligned_double_hit prevalence >= 5% of depth-4 positions — observed: 19.75%
- ✓aligned_double_hit uncollected-setup rate >= 30% — observed: 60.47% (excluding ambiguous 56.51%)
- ✓CHECK gates (bit parity, search parity, mirror, determinism, metadata blindness, memo identity) at d4s5 and d4s7 — observed: 0 mismatches in every gate; runs/RUN-20260823T091530Z-cbe65468/gates.log
- –Screen gate (A minus frozen lower bounds, Q25, halves, coverage) — observed: not evaluated: the screen was not run because the corpus gate failed
Recorded metrics
- records
- 5,257,181
- depth4Rows
- 62,831
- depth4Games
- 768
- depth4Train
- 50,003
- depth4Test
- 5,528
- baseHeldOutR2
- 0.6483
- partialR
- -0.0443
- incrementalR2
- 0.0005
- prevalenceDepth4
- 0.1975
- uncollectedRate
- 0.6047
- uncollectedRateExcludingAmbiguous
- 0.5651
- setups
- 6,969
- revealedWithin2
- 2,750
- partialR
- -0.0325
- incrementalR2
- 0.0003
- prevalenceDepth4
- 0.2892
- partialR
- -0.0385
- incrementalR2
- 0.0001
- prevalenceDepth4
- 0.0177
- uncollectedRate
- 0.6285
- uncollectedRateExcludingAmbiguous
- 0.5775
- setups
- 324
- revealedWithin2
- 120
- partialR
- -0.0595
- incrementalR2
- 0.0013
- prevalenceDepth4
- 0.1678
- partialR
- -0.0504
- incrementalR2
- 0.0007
- prevalenceDepth4
- 0.1698
- partialR
- -0.0277
- incrementalR2
- 0.0003
- prevalenceDepth4
- 0.2383
- partialR
- -0.0064
- incrementalR2
- 0.0000
- prevalenceDepth4
- 0.0151
- leafBitMismatches
- 0
- parityMismatches
- 0
- mirrorMismatches
- 0
- determinismMismatches
- 0
- metadataBlindnessMismatches
- 0
- liveDivergence
- A_d4s5
- 2/600
- B_d4s5
- 5/645
- A_d4s7
- 7/370
- B_d4s7
- 6/310
- Mechanics-only tier: no gameplay was run; this rejects the terms as held-out predictors of remaining lifetime beyond the frozen leaf on depth-4 corpus positions, not as action-changing leaf terms in play.
- The corpus behaviour policy is the reference depth-4 search at FIVE strata with epsilon 0.03 exploration, not the seven-stratum memo configuration the screen would have run; the uncollected-setup rate is therefore inflated by epsilon moves and by s5's coarser chance model (Kimi K3 review, written before the result was read: runs/RUN-20260823T091530Z-cbe65468/kimi-k3-prereg-review.md section 3).
- The same review argued, blind to the result, that a partial-correlation kill criterion is a value criterion and can reject a term whose purpose is to re-rank sibling actions at the horizon boundary; the preregistered gate is nonetheless applied as written. Any successor that drops the value criterion must disclose that this gate was read first.
- The uncollected-setup rate counts a reveal of a different gray as not collecting this setup; a high rate can mean the search had better lines.
- The formal freeze command ran after the corpus gate was read; the thresholds were fixed by file content 9.5 minutes before the read (experiment amendment).
- The implementer's support-disjoint rule R2 (skip adjacent-side pairs whose completion paths share a cell) excludes configurations where one dropped disc completes both runs in the same wave; the term as frozen is narrower than the design's prose.
Paired 256-game SCREEN on fresh seeds 0xa52d0200-0xa52d02ff, fair D4 at seven strata with the memo engine. The preregistered primary arm A300 (aligned_double_hit +300) and the bundle arm B changed the unchanged search's column in 0.69% and 0.70% of decisions at the 32-game rarity check, below the 1% rule, and were stopped as no-measurements (partials: A300 -4,298 on 39 games, B +18,213 on 36 games, both far inside their floors). The dose arm A900 (aligned_double_hit +900), read as primary per the gate text, completed 256 games: mean 389,749 vs 386,545, paired delta +3,204 points (one-sided 95% bootstrap lower bound -26,860; Student-t -27,863; detection floor 30,957; paired sd 301,098), moves 112.55 vs 111.59 (+0.97), W-T-L 99-53-104, Q25 +18,534 (non-regression met), halves +48,762 and -42,354 (opposite signs: fail), coverage 598/28,814 = 2.08% (measured, not rarity). Predeclared mechanism directions were absent: cover reveals per move 1.1520 vs 1.1536, numbered clears per move 2.0541 vs 2.0551, occupancy 23.32 vs 23.22. Gate: FAIL. The term re-ranks about one decision in fifty at +900 and those re-rankings add no reveal flow; the score delta is a non-measurement for effects under about 31,000 points, but the flat flow statistics, whose paired noise is far smaller, reject the mechanism itself. The frozen arm is also the largest fresh-seed fair-D4 seven-stratum cohort on record: 386,545 points, 111.59 moves, 2.0551 clears and 1.1536 reveals per move over 256 never-read games, consistent with the 64-game 398,498 reference.
- ✓CHECK gates at the frozen term source — observed: 0 mismatches (gates.log, gates-phase3.log)
- ✓Artifacts carry the lease and role; 0 incomplete and 0 illegal decisions — observed: SL-20260823T110000Z-a52d0200 / public-development; 0 / 0
- ✕A300 coverage >= 2% (else inconclusive by rarity, A900 read as primary) — observed: 0.69% at the 32-game rarity check; arm stopped; A900 read as primary at 2.08%
- ✕Primary minus frozen: bootstrap and Student-t one-sided 95% lower bounds > 0 — observed: A900: +3,204; LB -26,860 / -27,863
- ✕Primary: Q25 non-regression AND both halves > 0 — observed: Q25 +18,534 (met); halves +48,762 / -42,354 (opposite signs)
- ✕Mechanism (reported): reveals per move above frozen — observed: 1.1520 vs 1.1536
Recorded metrics
- games
- 256
- meanScoreCandidate
- 389,749
- meanScoreReference
- 386,545
- pairedDelta
- 3,204
- bootstrapLB95
- -26,860
- studentTLB95
- -27,863
- detectionFloor
- 30,957
- pairedSd
- 301,098
- winTieLoss
- 99-53-104
- medianCandidate
- 326,048
- medianReference
- 313,466
- q25Candidate
- 215,187
- q25Reference
- 196,653
- q25Delta
- 18,534
- minCandidate
- 102,610
- minReference
- 102,894
- maxCandidate
- 1,976,127
- maxReference
- 1,656,350
- halves
- 48,762
- -42,354
- movesCandidate
- 112.5500
- movesReference
- 111.5900
- movesDelta
- 0.9700
- movesBootstrapLB95
- -7.1700
- A900
- numberedClearsPerMove
- 2.0541
- coverRevealsPerMove
- 1.1520
- meanOccupiedCells
- 23.3200
- maxChainDepth
- 14
- frozen
- numberedClearsPerMove
- 2.0551
- coverRevealsPerMove
- 1.1536
- meanOccupiedCells
- 23.2200
- maxChainDepth
- 12
- A900
- divergentDecisions
- 598
- shadowDecisions
- 28,814
- rate
- 0.0208
- A300_at_stop
- divergentDecisions
- 27
- shadowDecisions
- 4,052
- rate
- 0.0067
- completedGames
- 39
- B_at_stop
- divergentDecisions
- 26
- shadowDecisions
- 3,745
- rate
- 0.0069
- completedGames
- 36
- A300
- games
- 39
- pairedDelta
- -4,298
- bootstrapLB95
- -41,726
- B
- games
- 36
- pairedDelta
- 18,213
- bootstrapLB95
- -35,413
- incompleteDecisions
- 0
- illegalDecisions
- 0
- censoredGames
- 0
- memoHitRate
- frozen
- 0.6483
- A900
- 0.6484
- maxWork
- 11,892,399
- wallSecondsStage2
- 5,299
- games
- 256
- meanScore
- 386,545
- meanMoves
- 111.5900
- numberedClearsPerMove
- 2.0551
- coverRevealsPerMove
- 1.1536
- gamesAtOrAboveMillion
- A score null inside a 30,957-point floor is a non-measurement for effects below that size; the rejection rests on the flat flow statistics and the opposite-sign halves, not on the score delta.
- The preregistered primary dose (+300) could not be measured: it changed under 1% of decisions and was stopped by the rarity rule; the +900 dose was read as primary under the gate's own clause and is therefore a pass-or-fail at a dose chosen after a rarity null.
- The runner has no per-arm stop; the operator killed the four-arm run after the 32-game check and relaunched frozen + A900 on the same seeds (deterministic; the frozen arm reproduces stage-1 games exactly). Stage-1 partials are retained under stopped-at-32/.
- The corpus value gate for this term failed before this screen (RS-20260823T110000Z-baecb816) and was disclosed in the protocol; this result is the in-play test that the blind review asked for.
- Single cohort, public-development tier; not independently replicated.
- The frozen-arm fresh baseline is a by-product and has not been registered as a benchmark manifest.