ResultP-SOL-2: corrected two-rung G0 fidelity ladder (fast-d3s7 proxy + native D3 N7M6 guardrail), fast-engine corpus continuations, and the G1/G2 offline gates
P-SOL-2 stage G0 fails per the frozen failureAction, with no leased seed opened.
On this page
- Recorded
No explanation has been written for this record yet.
Technical recordMetrics, gate checks and limitations
P-SOL-2 stage G0 fails per the frozen failureAction, with no leased seed opened. Re-gates 7/7 pass, including fast-vs-native d3s7 parity (240 decisions, 0 action/work mismatches). Rung 1 S1-halted on measured cost: fast-d3s7 continuations run at 54.6 ms/move, ~13x the design's 4 ms estimate, projecting 47.9 CPU-h against the 6.76 ceiling; no tau was computed and the 1,200-root seed-free pool is retained. Rung 2 executed within envelope and the guardrail KILLED the proxy: on 6 CRN-matched C0 roots (exact replay verified), fast-d3s7 vs native D3 N7M6 within-root KM-lifetime orderings agree at mean tau 0.370 (LB95 0.283, min 0.053), top-1 4/6 - both kill conditions trigger, worst at late-game roots. Scientific consequence: the M=6 reveal quadrature genuinely changes within-root orderings, so an M=1 continuation corpus cannot carry D3 N7M6 label semantics; the fast-engine M=6 port (E-FAST-M6) is the required reopening route for any cheap-continuation label programme. Total ~7.5 CPU-h of the 35 cap; training lease still reserved and unopened.
- ✓Re-gates: v1 byte-identity, fast-vs-native d3s7 decision parity, CRN tape parity, mirror invariance, thread/run determinism, legality — observed: 7/7 pass; parity 240 decisions, 0 action/work mismatches (work 38,179,907 both sides); CRN 2,016 comparisons 0 mismatches
- ✕Rung 1: at least one of D1/D2 with tau LB95 >= 0.75 and top-1 LB95 >= 0.80 vs fast-d3s7 on 1,200 roots — observed: not run: S1 halt - measured fast-d3s7 54.6 ms/move (design assumed 4), projection 47.9 CPU-h vs 6.76 ceiling; root pool retained
- ✕Rung 2 guardrail: top-1 > 4/6 AND mean tau >= 0.5 AND no root tau < 0 (fast-d3s7 vs native D3 N7M6, 6 CRN-matched C0 roots) — observed: mean tau 0.370 (LB95 0.283, median 0.304, min 0.053), top-1 4/6 - kill conditions 'mean tau < 0.5' and 'top-1 <= 4/6' both trigger; native 3.40 s/move realised
Technical recordRecorded metrics
- halted
- S1 on projection
- ratesMsPerMove
- d1
- 0.0740
- d2
- 1.9900
- fastD3s7
- 54.6000
- projectionCpuH
- 47.9000
- ceilingCpuH
- 6.7600
- roots
- 6
- K
- 4
- H
- 40
- meanTau
- 0.3700
- tauLB95
- 0.2830
- medianTau
- 0.3040
- minTau
- 0.0530
- top1Agreement
- 4/6
- kill
- mean tau < 0.5
- top-1 <= 4/6
- nativeSecondsPerMove
- 3.4000
- cpuHours
- 4.6000
- Rung 2 is a 6-root guardrail: it detects gross misalignment (kill power ~0.89 at true agreement 0.5) and its kill here is decisive for the proxy, but the tau point estimates carry wide intervals.
- Rung 1's powered comparison never ran, so D1/D2 fidelity to fast-d3s7 is unmeasured.
- The 20-minute wall overage over the subagent's 2.5 h cap was spent in the single-threaded determinism gate and is disclosed.
Recorded against P-SOL-2: corrected two-rung G0 fidelity ladder (fast-d3s7 proxy + native D3 N7M6 guardrail), fast-engine corpus continuations, and the G1/G2 offline gates.
Agent contextHow to extend this record
To add a reader-facing explanation, write web/content/research/RS-20260823T225753Z-0fbd48c3.mdx; it renders above this record on the next request. The recorded metrics, gate checks and limitations are in the technical record above.
Record file: research/results/RS-20260823T225753Z-0fbd48c3.json, validated against research/schemas/result-v1.schema.json.
- Run ids
RUN-20260823T224351Z-6ad4e09e
- Contribution ids
CT-20260823T224410Z-869dd2f6CT-20260823T215412Z-e59c1b8cCT-20260823T205825Z-40e17d96
- Per-game artifact
runs/RUN-20260823T215500Z-sol/g0v2/rung2.json(sha25635b1bac76ccdf4f94aa90454f6f64db2b79b97e3cb5b034ba59b1416d2bf68d5, 6 records)- Machine profiles
research/system-profiles/MACH-20260820T080056Z-376ada90.json