Drop7 Research
← Theories

Complete-sibling long-horizon outcome labels fix the learned leaf's within-root discrimination (P-SOL)

untestedproposedevidence: proposalpublic-policyTH-20260823-sibling-outcome-labels-fix-leaf-0c3d480d
No explanation has been written for this theory yet. Add web/content/research/TH-20260823-sibling-outcome-labels-fix-leaf-0c3d480d.mdx and it will appear here. The registered record is shown below.

The registered record

Claim

A leaf-cost student (the existing 572k-parameter finding-08 NNUE, architecture unchanged) trained on ALL legal siblings per root, with labels that are long-horizon outcome distributions (per-rise hazard vector plus censored restricted-mean lifetime, Kaplan-Meier over K CRN-paired continuations under a fixed cheap public continuation policy) and a within-root pairwise ranking loss on the deployed scalar (KM expected lifetime), achieves within-root discrimination that played-action outcome training cannot provide (coverage) and exact-search-value distillation cannot provide (objective): offline, top-1 agreement with exact D4's argmax >= 0.55 and >= the played-action incumbent + 0.03; deployed as the existing blended leaf, it improves the paired 256-game d4s7 mean with a positive one-sided 95% lower bound without giving back the incumbent's d4s5 gain.

Mechanism

Design: Kimi K3, runs/RUN-20260823T191900Z-b9f8f80d/kimi-k3-main-design.md (P-SOL-v1), grounded in tonight's three results. M1 coverage: one row per played move gives no within-root gradient; complete panels supply 21 ordered pairs per root. M2 objective: exact D4 values encode search depth a 1.33us leaf cannot represent (RS-20260823T194142Z-946e3cd1: top-1 0.296/0.301 at 572k, 0.375 at 3.4M, vs exact D1 0.486); outcome distributions under a fixed cheap continuation are properties of the afterstate and leaf-representable. M3 label economics: under within-root CRN pairing, sibling outcome differences are driven by afterstate structure, so a cheap continuation engine (D1/D2) can preserve the ranking of D3 N7M6 (the C0-validated strong cheap policy, RS-20260823T194200Z-42b113db) - measured by the G0 fidelity ladder before any corpus spend, never assumed. Single-factor discipline: model, features, head weights and deploy path held identical to finding-08; only sibling-complete data, the K-sample CRN estimator and the ranking term change.

What would prove it wrong

  • F1: no cheap continuation engine achieves within-root label Kendall-tau vs D3 N7M6 with 95% LB >= 0.85 on the G0 ladder (label-fidelity failure; stop at ~14 CPU-h).
  • F2: offline gate G1 fails (student top-1 vs exact-D4 argmax < 0.52, or < incumbent LeafNet + 0.01) while the G2 decomposition shows target quality adequate (label-argmax top-1 vs D4 >= 0.60): sibling-complete data did not fix discrimination.
  • F3: G2 shows label-argmax top-1 vs D4 < 0.60: the outcome-label family itself does not track decision quality; falsifies the label mechanism regardless of fit.
  • F4: gameplay co-primary fails: paired 256-game d4s7 lower bound <= 0 (successor screen experiment).
  • A negative rejects only the tested configuration {continuation in D1/D2, H=48, K=16, 572k student, ranking+regression mix}; D3-continuation labels at corpus scale remain open via the scale-out path.
  • Status 2026-08-23: F1 is triggered in its v2 form - no affordable continuation engine could be certified (the fast M=1 proxy failed the D3 N7M6 guardrail: mean tau 0.370, top-1 4/6) - so the core claims M1/M2 remain untested rather than refuted; the theory is blocked pending the fast-engine M=6 port (E-FAST-M6) or a scale-out lease for native-M6 labels.
registered 2026-08-23T21:18:28Z by unknown / unknown

Experiments that test it

Results recorded against it

valid run · outcome: failnot-supported-as-testedpilotRS-20260823T225753Z-0fbd48c3

P-SOL-2 stage G0 fails per the frozen failureAction, with no leased seed opened. Re-gates 7/7 pass, including fast-vs-native d3s7 parity (240 decisions, 0 action/work mismatches). Rung 1 S1-halted on measured cost: fast-d3s7 continuations run at 54.6 ms/move, ~13x the design's 4 ms estimate, projecting 47.9 CPU-h against the 6.76 ceiling; no tau was computed and the 1,200-root seed-free pool is retained. Rung 2 executed within envelope and the guardrail KILLED the proxy: on 6 CRN-matched C0 roots (exact replay verified), fast-d3s7 vs native D3 N7M6 within-root KM-lifetime orderings agree at mean tau 0.370 (LB95 0.283, min 0.053), top-1 4/6 - both kill conditions trigger, worst at late-game roots. Scientific consequence: the M=6 reveal quadrature genuinely changes within-root orderings, so an M=1 continuation corpus cannot carry D3 N7M6 label semantics; the fast-engine M=6 port (E-FAST-M6) is the required reopening route for any cheap-continuation label programme. Total ~7.5 CPU-h of the 35 cap; training lease still reserved and unopened.

What it had to pass
  • Re-gates: v1 byte-identity, fast-vs-native d3s7 decision parity, CRN tape parity, mirror invariance, thread/run determinism, legality — observed: 7/7 pass; parity 240 decisions, 0 action/work mismatches (work 38,179,907 both sides); CRN 2,016 comparisons 0 mismatches
  • Rung 1: at least one of D1/D2 with tau LB95 >= 0.75 and top-1 LB95 >= 0.80 vs fast-d3s7 on 1,200 roots — observed: not run: S1 halt - measured fast-d3s7 54.6 ms/move (design assumed 4), projection 47.9 CPU-h vs 6.76 ceiling; root pool retained
  • Rung 2 guardrail: top-1 > 4/6 AND mean tau >= 0.5 AND no root tau < 0 (fast-d3s7 vs native D3 N7M6, 6 CRN-matched C0 roots) — observed: mean tau 0.370 (LB95 0.283, median 0.304, min 0.053), top-1 4/6 - kill conditions 'mean tau < 0.5' and 'top-1 <= 4/6' both trigger; native 3.40 s/move realised
Recorded metrics
regates
7/7 pass; fast-vs-native parity 240 decisions 0 mismatches, work 38,179,907 both sides
rung1
halted
S1 on projection
ratesMsPerMove
d1
0.0740
d2
1.9900
fastD3s7
54.6000
projectionCpuH
47.9000
ceilingCpuH
6.7600
rung2
roots
6
K
4
H
40
meanTau
0.3700
tauLB95
0.2830
medianTau
0.3040
minTau
0.0530
top1Agreement
4/6
kill
  1. mean tau < 0.5
  2. top-1 <= 4/6
nativeSecondsPerMove
3.4000
cpuHours
4.6000
totalCpuHours
7.5000
Limitations
  • Rung 2 is a 6-root guardrail: it detects gross misalignment (kill power ~0.89 at true agreement 0.5) and its kill here is decisive for the proxy, but the tau point estimates carry wide intervals.
  • Rung 1's powered comparison never ran, so D1/D2 fidelity to fast-d3s7 is unmeasured.
  • The 20-minute wall overage over the subagent's 2.5 h cap was spent in the single-threaded determinism gate and is disclosed.
valid run · outcome: passsupported-as-testedmechanics-onlyRS-20260824T010000Z-8f3e9b4f

E-FAST-M6 passes: factored reveal sampling (N disc strata x M reveal samples, native scenario indexing s = r*N + d over T = N*M) is ported into the fast memo engine as drop7::fastr::FastFactoredSearch and is trace-equivalent to the native FactoredSearch. Grid gate: 6,300 live probe decisions (525 per point over d3/d4 x N5/N7 x M1/M2/M6) with 0 column, 0 work-count and 0 completed-depth mismatches, including 549 work-limited decisions on the two budget-capped d4-M6 points where both engines degrade to completed depth 3 identically. All 3 retained C0 games replay to byte-for-value final identity with the fast engine driving (335 decisions, 0 mismatches). M=1 regression: 2,100 decisions x {memo-on, memo-off} bit-identical to the untouched fast::FastSearch on all six metric fields. Determinism byte-identical across repeated runs and {26,7,1} threads; mirror invariance exact with symmetric boards excluded (finding-13 4C; one disclosed gate-harness iteration); the one-entry leaf memo stays enabled under M>1 (board-memcmp keying cannot alias across reveal samples) with memo-on/off trace identity across the grid. Continuation duty: 8 CRN continuations (2 C0 roots, K=4, H=40) byte-identical to native. Realised speedup at d3 N7M6 is 5.5-5.8x (play duty 2.140 -> 0.385 s/move; continuation duty 0.990 -> 0.177 s/move; grid 4.38-5.87x across points), measured under the gate's own 26-thread load - well below the 10-40x hoped for in the hypothesis, because that figure divided native M6 seconds by fast M1 seconds and ignored the ~27x work ratio. P-SOL continuation labels at native D3 N7M6 semantics now cost ~0.18 s/move instead of ~1-3.4 s/move.

What it had to pass
  • Equivalence: >= 500 live decisions per grid point over d3/d4 x N5/N7 x M1/M2/M6 on probe seeds and >= 3 full replayed C0 games, 0 column/work/completed-depth mismatches vs the native factored search — observed: 525 decisions per point (6,300 total), 0/0/0 mismatches; per-point native and fast trace hashes identical; 3 C0 replays with 0 mismatches and all finals byte-for-value identical to runs/RUN-A525-reveal/d3-n7-m6.json
  • M=1 regression: the port at M=1 remains bit-identical to the existing fast search — observed: 2,100 decisions x memo-on and memo-off vs fast::FastSearch: 0 mismatches on action, completed_depth, nodes, work, cache_hits, cache_entries
  • Determinism: byte-identical outputs across two runs and across thread counts; mirror invariance of decisions — observed: subset artifact byte-identical across {repeat at 26, 7, 1} threads; mirror invariance 0 mismatches on 114 asymmetric-board decisions (18 mirror-symmetric boards excluded per finding-13 4C; iteration 1's harness compared symmetric boards and is disclosed in gates.log)
  • npm test and make test pass; default random-reveal behaviour byte-identical (latent-mode contract untouched) — observed: npm test pass; make test pass including TypeScript/native parity (256 seeds, 6,852 moves, exact); the change is additive (new approach directory only) and touches no engine source
Recorded metrics
gridDecisions
6,300
gridMismatches
column
0
work
0
completedDepth
0
gridPointsAtLeast500Decisions
12/12 (525 each)
c0Replay
games
3
decisions
335
mismatches
0
finalsIdentical
3/3 (320871/95, 381355/110, 463094/130)
m1Regression
decisions
2,100
comparators
fast::FastSearch vs port memo-on and memo-off
metricFieldMismatches
0
determinism
subset output byte-identical across {repeat, 26, 7, 1} threads (sha256 645a2bb8977f4e86ad95c048964186061970183179ae615159dde8cccb0aed97)
memoUnderM6
enabled; memo-on/off trace identity on all 12 grid points and 2,100 M=1 decisions; no aliasing possible (full-board memcmp key below the work increment)
continuationDuty
continuations
8
outcomeMismatches
0
nativeCpuSecondsPerMove
0.9904
fastCpuSecondsPerMove
0.1767
speedup
5.6000
playDuty
nativeCpuSecondsPerMove
2.1397
fastCpuSecondsPerMove
0.3847
speedup
5.5600
gridSpeedupRange
4.38-5.87x (per-point table in timing.json)
testSuites
npm test pass; make test pass (parity 256 seeds / 6,852 moves exact)
Limitations
  • The two d4 M=6 grid points run under disclosed budget-capped work bounds (51,084,852 and 100,000,000) rather than their infeasible worst-case bounds; this deliberately exercises the work-limit and LRU-eviction paths, but full-depth d4-M6 completion is not itself gated (the C0 configuration d3 N7M6 uses its exact retained bounds).
  • Timing was measured with CLOCK_THREAD_CPUTIME_ID while the gate itself loaded 26 threads; idle-host per-move times will be lower for both engines, and the speedup ratio (4.4-5.9x) is the robust figure, not the absolute times.
  • The realised speedup is ~5.5-5.8x at d3 N7M6, not the 10-40x the hypothesis projected from an M1-vs-M6 comparison; the P-SOL-3 successor's budget must be re-planned from the measured 0.177 s/move continuation rate.
  • Total CPU seconds are measured for stage E and rate-derived for the resumed stages (run record measurementNotes).
  • Self-reported; the run includes one disclosed external task kill (resumed between stages) and one disclosed gate-harness iteration on stage D.