ResultThe two arms finding-09 left unfinished: reveal sampling on top of the fourth ply, and the depth-3 ladder at full joint coverage
Chance-node decorrelation and search depth do not compound; they substitute.
On this page
- Recorded
No explanation has been written for this record yet.
Technical recordMetrics, gate checks and limitations
Chance-node decorrelation and search depth do not compound; they substitute. The primary arm is complete at 64 of 64 games: depth 4 with seven disc samples and two reveal samples scores 356,548 against the 398,498 of the same search with one reveal sample, a paired -41,950 with a one-sided 95% whole-game bootstrap lower bound of -100,137 and W-T-L 28-0-36, for 4.07x the logical work per move. The gate asked for a lower bound above zero and got a negative point estimate, so the compounding theory is rejected as tested. STATE THE DOSE WHEN QUOTING THIS: two reveal samples raises joint (disc, reveal) coverage from 14.3% to 28.6%, which is a smaller increment than the six samples (85.7%) that first cleared noise at depth 3; at depth 3 the three-sample dose (42.9%) was also not significant (+24,980, lower bound -23,451). This result therefore rejects a doubling of reveal samples on top of the fourth ply, and does not establish that a wide reveal estimator at depth 4 would fail - that arm was never affordable. What the arm does establish is that the depth-4 search is not starved for the thing the extra samples supply. The striking positive finding is an equivalence at near-equal work: depth 3 with six reveal samples costs 4,244,020 work per move and scores 376,442, while depth 4 with one reveal sample costs 4,956,614 and scores 398,498, and the paired delta between them is -22,056 with a lower bound of -89,867 and W-T-L 30-0-34 - two different ways of spending the same budget landing in the same place, which is the same exchangeability the depth factorial shows from the other side. The new arm is not worthless: against the frozen five-stratum depth-4 reference it is +59,221 with a lower bound of +9,134 and W-T-L 37-0-27, so the gain comes from the seven disc samples, not from the reveal samples. Second arm, partial at 32 of 64 games and still running: depth 3 with twelve reveal samples takes joint coverage to 100% and scores -4,495 against the six-sample arm on the 32 shared seeds (lower bound -91,425, W-T-L 15-0-17) and +31,413 against the one-sample arm (lower bound -70,729), so the ladder that ordered with coverage from M=1 to M=6 stops ordering at M=12. Every arm audited 0 decisions below target depth, 0 work-limit events, 0 censored games and 0 score-decomposition identity failures, and the chunk-pooling used to survive interruption was verified to reproduce a single 64-game run field-for-field with identical summed logical work.
- ✓Pooling validity: four pooled 16-game chunks reproduce a single 64-game run on all 13 per-game fields with identical summed logical work — observed: 64 games, 0 field mismatches, logical work 312,966,881 vs 312,966,881 (equal)
- ✓Bound diagnostics: 0 decisions below target depth and 0 work-limit events in every arm — observed: d4 N=7 M=2: 6,633 decisions, 0 below target, 0 work-limit events, busiest decision at 44% of its bound; d3 N=7 M=12 (partial): 3,335 decisions, 0 below target, 0 work-limit events, busiest at 31% of its bound
- ✕Depth 4 with two reveal samples beats depth 4 with one reveal sample, 95% bootstrap lower bound above zero, 64 complete games — observed: -41,950 with a 95% lower bound of -100,137 (upper bound +17,541), W-T-L 28-0-36, at 4.07x the logical work per move; negative point estimate, and not distinguishable from zero in either direction
- –Depth 3 with twelve reveal samples is at least as strong as depth 3 with six — observed: undecidable as run: the arm holds 32 of 64 games and is still executing. On the 32 shared seeds the delta is -4,495 with a 95% lower bound of -91,425 and W-T-L 15-0-17 - the coverage ladder has stopped ordering, but 32 games cannot decide this criterion
- ✓0 censored games and 0 score-decomposition identity failures in every arm — observed: 0 and 0 in both new arms and in every comparator arm re-read here
Technical recordRecorded metrics
- M1
- 312,327
- M3
- 337,306
- M6
- 376,442
- M12partial32
- 356,890
- M1
- 0.1430
- M2
- 0.2860
- M3
- 0.4290
- M6
- 0.8570
- M12
- 1
- n
- 64
- meanScoreDelta
- -41,950
- lowerBound95
- -100,137
- upperBound95
- 17,541
- meanMoveDelta
- -11.0200
- winTieLoss
- 28-0-36
- workRatio
- 4.0700
- n
- 64
- meanScoreDelta
- 59,221
- lowerBound95
- 9,134
- meanMoveDelta
- 16.4800
- winTieLoss
- 37-0-27
- workRatio
- 15.5700
- n
- 64
- meanScoreDelta
- -19,894
- lowerBound95
- -76,456
- meanMoveDelta
- -5.8100
- winTieLoss
- 37-0-27
- workRatio
- 4.7500
- n
- 64
- meanScoreDelta
- -22,056
- lowerBound95
- -89,867
- meanMoveDelta
- -5.2000
- winTieLoss
- 30-0-34
- workRatio
- 0.8600
- n
- 64
- meanScoreDelta
- 24,980
- lowerBound95
- -23,451
- winTieLoss
- 32-0-32
- n
- 64
- meanScoreDelta
- 64,116
- lowerBound95
- 7,475
- winTieLoss
- 36-0-28
- n
- 32
- meanScoreDelta
- -4,495
- lowerBound95
- -91,425
- winTieLoss
- 15-0-17
- note
- partial arm, paired on the 32 shared seeds
- n
- 32
- meanScoreDelta
- 31,413
- lowerBound95
- -70,729
- winTieLoss
- 15-0-17
- note
- partial arm, paired on the 32 shared seeds
- n
- 32
- meanScoreDelta
- -31,616
- lowerBound95
- -143,344
- winTieLoss
- 12-0-20
- note
- partial arm, paired on the 32 shared seeds
- decisions
- 6,633
- decisionsBelowTargetDepth
- 0
- workLimitEvents
- 0
- minCompletedDepth
- 4
- maxDecisionWork
- 81,686,570
- declaredBound
- 187,336,114
- decisions
- 3,335
- decisionsBelowTargetDepth
- 0
- workLimitEvents
- 0
- minCompletedDepth
- 3
- maxDecisionWork
- 128,386,272
- declaredBound
- 407,634,528
- games
- 64
- fieldMismatches
- 0
- summedLogicalWorkCandidate
- 312,966,881
- summedLogicalWorkComparator
- 312,966,881
- note
- four pooled 16-game chunks against the single 64-game d3 N=5 M=1 run
- PARTIAL ARM: the depth-3 twelve-reveal-sample arm holds 32 of 64 games (chunks 0 and 1 of 4) and was still executing when this result was written. Every M=12 number here is over those 32 paired seeds, none of them decides its gate, and the arm's 32-game mean is not an estimate of its 64-game mean. The frozen snapshot assessed is runs/RUN-20260821T143541Z-4c4370b6/d3-n7-m12-partial-32games.json; re-pooling the live artifact after chunk 2 will change its bytes.
- DOSE, NOT AXIS: the depth-4 arm tested two reveal samples (28.6% joint coverage), which is below the depth-3 dose that first cleared noise (six samples, 85.7%). The negative result is about that dose at that depth. A wide depth-4 reveal estimator was not run and remains unmeasured; at the observed 20.2M work per move for M=2, an M=6 depth-4 arm would be roughly another 3x on top and was not affordable on a shared machine.
- The depth-4 M=2 delta is negative but not significantly negative: the same estimator's one-sided 95% upper bound is +17,541. The honest reading is 'buys nothing measurable for 4.07x the work', not 'harms'.
- The comparator arms for the M ladder (d3 M=1, M=3, M=6) and the two depth-4 single-sample arms come from finding-09 and from the earlier chance-strata study; they were re-read here from their retained artifacts but have no machine-readable run record of their own, so this result's run records cover only the two new arms.
- 64 paired games with a score standard deviation of 55-62% of the mean. Effects smaller than roughly 60,000 points are invisible at this sample size.
- The cohort 0xa51d1000-0xa51d103f was opened by finding-05 and read again by finding-09; it is permanently development data. STANDARD tier, diagnostic only, never a freeze gate and never confirmation evidence.
- Each arm ran as four sequential 16-game chunks so it could survive interruption. Pooling was verified to reproduce a single run exactly, but chunking does change thread scheduling and wall time, and chunk 1 of the depth-4 arm ran at 8 threads rather than 12 after a load back-off.
- Work per move is not comparable across arms with different declared cache capacities. The new arms auto-size their cache from their branching factor (960,695 and 346,921 entries) while the recorded comparators declare 60,000; every work figure must be read with its capacity.
- The experiment record was written after the runs, by a different agent than the one that executed them. The two arms and their launch protocol were fixed in prose in finding-09 and in run-arms.sh before either produced a game, but this is retroactive registration and the amendment on the experiment record says so.
- All artifacts for this family live under runs/, which is gitignored, so the evidence is not committed with the record. The frozen snapshots and the content manifest under the two run directories are the durable reference.
- This rejects compounding at the tested dose only. It says nothing about the reveal-by-reveal correlation, which finding-09 identified as untouched by this knob and which no arm here attacks.
Recorded against The two arms finding-09 left unfinished: reveal sampling on top of the fourth ply, and the depth-3 ladder at full joint coverage.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/RS-20260821T181918Z-ea7076a3.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.
Record file: research/results/RS-20260821T181918Z-ea7076a3.json, validated against research/schemas/result-v1.schema.json.
- Run ids
RUN-20260821T035407Z-00483c6cRUN-20260821T143541Z-4c4370b6
- Contribution ids
CT-20260821T181919Z-05882e86
- Per-game artifact
runs/RUN-A525-reveal/d4-n7-m2-pooled.json(sha25606fc7bdacb9849962e7369515a8084afd3d1b6e4b9a92b7b0cb9c5526384d861, 64 records)- Artifact manifest
runs/RUN-20260821T035407Z-00483c6c/manifest.json- Machine profiles
research/system-profiles/MACH-20260820T080056Z-376ada90.json