The reveal the search leaves on the table
Morning: the owner's two strategic hypotheses became one structural leaf term that failed its corpus gate and its screen. Evening: a six-brief Kimi drafting fleet, two preregistered cheap experiments (both decisive negatives: no leaf-cost NNUE holds D4's ordering, and the oracle's boards carry no transferable fair value), one retained positive bound (D3 N7M6 beats D4 s5, LB +29,033), and the console learned to draw its own evidence.
The owner brought two hypotheses from watching the search play. First, a strategic mode: keep a ledger of discs that need resolving, and when nothing is urgent, build long chain reactions instead of taking the clear in front of you. Second, reveal maximisation: the search does not try to uncover gray discs, and lining up a cascade that fully exposes one — two hits in a single chain — is worth more than anything the leaf currently prices. The owner also asked that long-horizon and strategic thinking be offloaded to Kimi K3, so today's design, red-team review and forward portfolio are Kimi's, and the coordinator's job was to turn them into records, gates and runs.
What the search already knows
Before designing anything, a read-only map of the frozen leaf
(runs/RUN-20260823T091530Z-cbe65468/leaf-map.md) settled what is and is not
priced. The chance model already samples every cascade reveal inside the
four-ply horizon, and the leaf already pays for a reveal through its cell
counts: a solid gray costs −620, a cracked one −220, a numbered cell −18, plus
height relief, so a reveal is worth roughly +600 to +5,000 at the leaf. Every
earlier attempt to value reveals more — a +600 reward per reveal event, the
cover-exposure weights raised fifty- to a hundred-fold — produced fewer
reveals and lost. And the clairvoyant planner that maximises chains earns
1.75x the points per move and dies 58% sooner. Neither hypothesis could be a
bigger weight on an existing term.
Kimi's design: price the joint event, not the marginal
Kimi K3's design (kimi-k3-theory-design.md) located what is structurally
missing rather than mis-weighted. The cheapest reveal in the game is the
same-wave double hit — two numbered neighbours of a solid gray popping in one
cascade — and no term expresses it: solid_exposure is a weighted sum of the
two neighbours' marginal readiness, which cannot tell two discs that are each
80% ready from a pair that is jointly reachable in one wave. The proposal: an
aligned_double_hit term, the product of the two neighbours' readiness over
support-disjoint pairs, at +300; a second release step restricted to chains
that terminate on a cover; the owner's ledger as a soft entombment penalty for
values 3–7; and a danger gate that switches the construction terms off above
height four — the "no emergency" clause as a product no linear leaf can
express. Kimi's own prior: about one chance in seven that the primary arm
passes, and the ledger clause demoted to a declared-diagnostic fallback
because the entombed count already reads −0.023 against remaining life.
Gate first, seeds never
The experiment (EX-20260823-reveal-construction-screen-371fd638) put a
seed-free gate in front of the cohort: on the already-open training corpus,
the term had to predict remaining lifetime beyond the eighteen leaf features
(held-out partial correlation ≥ +0.05, incremental R² ≥ 0.005), be present in
at least 5% of depth-4 positions, and — the action-level statistic — the
depth-4 policy had to be leaving at least 30% of live setups uncollected
within two moves. The engine substrate was built the way the siblings are: the
search generated from the gated fast search by a guarded substitution, a
board-keyed memo, and a structural rule that the extra terms can only see the
board and the rise clock. Every CHECK gate passed at five and seven strata
(gates.log: zero bit, parity, mirror, determinism and metadata mismatches).
The corpus value gate failed
cost: one corpus pass, no seedRead once, the gate failed on its value half and passed on its action half
(RS-20260823T110000Z-baecb816). aligned_double_hit has a held-out partial
correlation of −0.044 with remaining lifetime beyond the leaf, and an
incremental R² of +0.0005 — the wrong sign, and the same lesson finding-14
taught: positions with ready discs parked next to grays are positions carrying
occupancy. Prevalence was 19.75%. But the depth-4 policy leaves 60.5% of
live same-wave setups uncollected within two moves (6,969 setups; 56.5%
excluding ambiguous outcomes). The chain-to-crack terms failed the same way,
and the solid variant is too rare (1.77%) to move a mean. Per the protocol, no
seed was opened.
Kimi's red-team review of the preregistration
(kimi-k3-prereg-review.md, written at 09:54 UTC, two minutes before the
corpus result existed) had argued exactly this risk: a partial-correlation
kill is a value criterion, and a term whose job is to re-rank sibling actions
at the horizon boundary can show zero or negative partial correlation and
still change what the search does. It also found four defects that would have
bitten a passing run — artifacts hard-labelled with the wrong seed lease, a
pinned comparison script that could not read the runner's output, a
comment-enforced information boundary, and a confounded arm B — all fixed
before any seed was read.
In play: one decision in fifty, and no extra reveals
cost: 256 paired games x 2 arms at d4s7, plus 39/36 stopped gamesThe successor (EX-20260823-reveal-construction-screen-v2-63d73b6a, lease
0xa52d0200–0xa52d02ff, opened once) dropped the corpus gate, disclosed its
result, and tested the term in the only place it can be settled: 256 paired
games at depth 4, seven strata, memo engine, with a shadow unchanged search
at every decision to count how often the term changes the column. At the
preregistered 32-game rarity check the +300 dose had changed 0.69% of
decisions and the chain-to-crack bundle 0.70% — below the 1% rule — so
both were stopped as no-measurements and the run continued with the
unchanged search and a +900 dose. That arm changed 2.08% of decisions and
finished at +3,204 points (RS-20260823T131226Z-16564ed9): bootstrap lower bound −26,860,
W–T–L 99–53–104, halves +48,762 and −42,354, and — the number that matters —
1.152 reveals per move against 1.154, with clears and occupancy equally
flat. The gate reads fail. A term that pays for the configuration one move
before a same-wave double hit does re-rank the search occasionally, and the
re-rankings buy no reveals. This is the fourth reveal-pricing null on record,
and the first that measured the mechanism directly rather than the score.
A by-product worth keeping: the unchanged arm of that run is the largest fresh-seed seven-stratum fair-D4 cohort so far — 386,545 points and 111.59 moves over 256 never-read games, 2.055 clears and 1.154 reveals per move, in line with the 64-game 398,498 reference.
What the owner's eye caught, and what it did not
The owner's observation was correct at the board: three fifths of the same-wave setups the policy could cash go uncashed. What the day showed is that cashing them is not where the game is lost. The search already samples those reveals inside its four plies and declines most of them for reasons the leaf prices better than the term did; the value gate said so first, and the flow statistics said it again with far less noise than the score could. The strategic-mode ledger was not tested in play today — its value form is closed three times over and the review argued the only live version is an action counterfactual at the entombing roots.
Kimi's portfolio: measure the instrument, then the ceiling
Asked to think past this screen, Kimi K3 wrote a conditional portfolio
(kimi-k3-portfolio.md). Two points stand out. First, same-seed pairing is
already exact common-random-numbers in this engine — the disc tape and each
move's reveal stream are pure functions of (seed, move) — so the 200k–370k
paired spread is divergence amplification, and no replay scheme can reduce it;
only a model-based lifetime estimator, validated seed-free on the retained
arms of the shared cohort, can lower the 50k floor. Second, the 1.54-million
clairvoyant ceiling was measured against a persistent hidden board, but the
native engine draws a gray's value at reveal time, so that planner knew things
no legal policy could know even in principle. A stream-clairvoyant planner
over the engine's actual future — about 25 CPU-hours — would bound every legal
policy per tape; if it sustains under about 2.25 clears per move, the
million-point mean is very likely out of reach under these semantics, and
saying so with a measured bound becomes the program's most valuable result.
Both are owner-level direction calls and were not started.
Evening: a drafting fleet and three cheap answers
The owner widened the directive: keep advancing the learned-model program, with Kimi K3 doing the long-horizon design work, toward either depth-3 search with an NNUE evaluator or a from-scratch network that reads the board and chooses the move. Six Kimi briefs ran in parallel from a coordinator survey of every learned attempt in the repository: the main corpus-and-target redesign, a D3+NNUE program, a direct-policy program, theory drafts for the remaining open items, a sweep for genuinely untried directions, and a figure plan for the documentation. The engineering digest that grounded them (runs/RUN-20260823T191900Z-b9f8f80d/substrate-map.md) corrected the coordinator's own plan before it ran: the per-rise hazard vector is already a trained head of the learned leaf, and the sibling panel carries only a one-move label per sibling, so the naive "turn the panel on and refit" experiment would have measured nothing new.
Kimi's open-items ranking put two nearly free experiments first, and both ran to completion tonight.
No leaf-affordable NNUE holds D4's ordering
top-1 0.30 vs gate 0.60The NNUE d4q ordering probe (EX-20260823-nnue-d4q-ordering-probe-0ca09bb1, RS-20260823T194142Z-946e3cd1) trained the 572k-parameter, 1.27 µs LeafNet architecture directly on the existing successor-closed exact-D4 sibling values with a within-root listwise loss, five initialisation seeds, whole-origin splits. Held-out top-1 was 0.296/0.301 by half-fold — below the 3.4M-parameter CNN's 0.375 and exact D1's 0.486, against a 0.60 gate. Total cost 166 seconds. The search-guided self-play loop's drafted form (evaluator as the search leaf) is now blocked by evidence at both model scales; any revival is root-prior shaped, or needs a different architecture class at leaf cost.
The oracle's boards are not better boards
tau = -0.96; top-1 76.6% vs 81.4%Stage D0 (EX-20260823-hpool-stage-d0-e0ad1c65, RS-20260823T205143Z-ead14c9d) finally measured the number audit-05 called the crux of the optimistic- curriculum idea: the fair-relabelled value of oracle-visited states. On a fresh training lease (0xa52e0000), 1,984 oracle states matched 1:1 to fair-D4 states on phase, occupancy and height, 32 public futures each: the transferable fraction came out negative (tau = −0.959, interval [−3.069, −0.390]), with the honest caveat that the denominator degenerated — 63 of 64 oracle games hit the 500-move cap. The criterion the cap cannot excuse also failed: at the oracle's own roots, under fair futures, the oracle's chosen column is top-1 less often (76.6%) than fair D4's choice at the same board (81.4%). Optimism lives in the tape, not the board. The H-pool program closes as preregistered.
Depth 3 with the full reveal quadrature beats D4 s5, now with a bound
+79,115, LB +29,033The free reanalysis Kimi's D3 program listed as C0 (EX-20260823-d3n7m6-vs-d4s5-paired-reanalysis-ea66f4ec, RS-20260823T194200Z-42b113db) computed the paired bound the +79,115 point estimate never had: D3 N7M6 minus D4 s5 on the shared 64-game cohort is lower-bound positive (+29,033 bootstrap, +27,548 Student-t, W-T-L 35-0-29, both halves positive), and against D4 s7 it is a wash at 0.86x the work. Already-read cohort, so diagnostic tier — but it is the first retained lower-bound-positive paired contrast for a non-default configuration, and it reprices the D3+NNUE program's fallback (T5a/T5b: absorb the disc quadrature, keep the reveal quadrature) as the strongest live route.
What Kimi's drafting fleet concluded
owner: kimi-k3 (OpenCode)The D3+NNUE program (kimi-k3-d3-nnue-program.md) puts its mass on the exact control T5a — collapse the leaf-ply disc loop closed-form, D4-s7-shaped strength at 0.82x D4 s5 work, prior 0.50 — and stages everything to be killable offline. The direct-policy program (kimi-k3-direct-policy-program.md) settles the recurrence question: the public state is Markov, a cross-move hidden state is illegal as the protocol is written, and what memory would buy is amortised computation, which per-decision attention can supply legally; its recommended first rung is a sibling-set transformer scored offline. The unexplored-space sweep (kimi-k3-unexplored.md) cut ten of thirteen candidate directions as already tried and kept three, led by a half-hour board-clear accounting measurement. The evening's two negatives sharpen all three documents: the set-ranker probe is now the only untried route to a D4-ordering student.
The console learned to draw its own evidence
owner: claude-fable-5 (Claude Code), kimi-k3 (OpenCode)A provenance-checked SVG figure toolchain landed in the web console: every chart point must name the record it came from or the generator refuses to render it, docs pages embed figures with a fenced block, and hover popovers carry values, bounds, n and the source record. Kimi authored twelve chart specs from the figure plan — resolving two number drifts in status.md's favor of the records, including the +26,605/+26,468 discrepancy — and five hand-drawn mechanism diagrams (sibling extrapolation, the two-hit reveal rule, strategy fusion, the occupancy fixed point, the information ladder). The dense prose these figures replace is being condensed next.
P-SOL G0: the panel2 generator shipped its gates, and the budget arithmetic killed the ladder before a seed was spent
owner: claude-fable-5 (Claude Code)Stage G0 of the preregistered sibling-outcome-label experiment
(EX-20260823-sol-corpus-and-offline-gate-4d3d86e4) built the machinery in
full: the corpus generator gained a --panel2 mode writing 992-byte
PanelRecordV2 rows — all seven siblings per root, K CRN continuation outcomes
per sibling, tapes derived from canonical public-root hashes under dedicated
domains and never from the seed — plus a reader, a legality checker, and the
Kendall-τ ladder analysis with the preregistered cluster bootstrap. Every
pre-seed gate passed on the already-open smoke range
(runs/RUN-20260823T215500Z-sol/g0/gates.log): v1 output byte-identical to a
pristine binary built from git HEAD, panel2 output byte-identical across
repeated runs and thread counts, mirror invariance exact by construction
(continuations run in the canonical orientation), and labels present only for
legal columns.
Then the seed-free feasibility check ended the stage. The design's D3 N7M6
budget lines rest on 292 moves/s — the throughput of the single-knob depth-3
search from RUN-A51D — but the frozen fidelity reference is the C0 factored
engine (N=7, M=6, work 51,084,852), which the retained C0 artifact and a
fresh smoke pilot both price at about 1.75–6.83 CPU-s per decision. The
64-root mini-ladder alone projects to 62.9–244.8 CPU-h against its 3.2 CPU-h
line, and every D3 N7M6 stage lands at roughly 20–77× its allocation, so stop
rule S1 halted G0 with the training lease SL-20260823T215000Z-a5216000
still unopened — no leased seed was read, and no ladder τ exists. The partial
is recorded in runs/RUN-20260823T215500Z-sol/g0/ladder.json and
research/runs/RUN-20260823T215500Z-fc74e0b4.json; the pipeline demo on 16
smoke roots (D1 vs D2, diagnostic only) is in smoke-ladder-demo.json. The
decision now owed by the coordinator: a P-SOL-v2 with honest throughput
arithmetic — a cheaper fidelity reference, shorter or fewer reference
continuations, or a fast-engine factored search — or a scale-out request,
since the ladder as frozen is a multi-hundred-CPU-hour object, not a
14-CPU-hour one.
Late evening: the main experiment survives its own stop rules, then fails honestly
The redesigned main experiment (P-SOL: complete-sibling Kaplan-Meier outcome labels for the learned leaf, designed by Kimi K3 after three silent OpenCode session crashes forced a fully inlined brief) went through two preregistrations in one night, and both behaved exactly as written.
P-SOL G0: the cheap-continuation label route
cost: ~8 CPU-h, zero leased seedsP-SOL-1 froze a budget that priced D3 N7M6 continuations at the depth-3 N5M1 throughput - a 20-77x error - and its S1 stop rule halted the run on projection before a single leased seed was read (the panel2 generator and all its byte-identity, CRN and mirror gates passed and carry forward). Kimi's corrected P-SOL-2 replaced the unaffordable powered ladder with a two-rung design, and rung 2's guardrail then killed the whole cheap route: on six CRN-matched roots replayed exactly from the C0 cohort, the fast M=1 engine's within-root sibling orderings agree with native D3 N7M6 at only mean tau 0.370 (top-1 4/6), while engine parity itself was exact over 240 decisions. The reveal quadrature is not a refinement - it changes which sibling looks best. No M=1 continuation corpus can carry these labels (RS-20260823T225753Z-0fbd48c3).
E-FAST-M6: the port is now the critical path
owner: coordinatorC0 made D3 N7M6 the most interesting policy in the program; tonight's guardrail kill made its cost the binding constraint on the whole learned-label programme. Porting M=6 reveal sampling into the fast memo engine is semantics-preserving engineering (days, plus about 5 CPU-h of preregistered equivalence gates, zero leased seeds) and would turn the 6-root guardrail into a powered test, re-open the P-SOL corpus at ~40x lower cost, and drop 100-game D3 N7M6 studies to about half a CPU-hour. The sibling-outcome theory itself remains untested, not refuted - it is blocked on this port, and the training lease is still reserved and unopened.
Log entries are a narrative account written by the contributors listed above. They are not evidence records: run validity, scientific outcome and evidence tier live with the experiment and result records an entry refers to.