Research log
A dated account of what was actually tried, by whom, and what came of it. Most days end with more rejected ideas than accepted ones, and those days are written up in the same detail as the good ones: a failed experiment rejects the configuration it tested, and that is a finished piece of work, not a wasted night.
The log is narrative. It is written by the people and models doing the work and is not itself evidence. Every number in an entry belongs to a record — a theory, an experiment or its result — and those records, with their run validity, scientific outcome and evidence tier, are the authority. Entries live in web/content/log/.
Source data
Counts are the outcome tallies recorded in each day's log frontmatter (web/content/log/2026-08-2*.mdx, the 'outcomes' block), copied verbatim: 2026-08-20 positive 2 / negative 5 / open 3; 2026-08-21 positive 2 / negative 7 / open 6; 2026-08-22 positive 2 / negative 2 / open 1; 2026-08-23 positive 1 / negative 5 / open 2. The figure plan (written mid-day on 2026-08-23) carried 0 / 2 / 1 for 2026-08-23; the finalized log frontmatter reads 1 / 5 / 2 and is charted — plan-vs-record drift resolved in the record's favour. Milestones behind these counts (verbatim anchors): the 400,675-point eight-game D4 confirmation (docs/strategies.md, Current reference table); the 398,498 d4s7 64-game confirmation mean (finding-05); the 386,545 frozen-arm mean over 256 fresh games (RS-20260823T131226Z-16564ed9, metrics.frozenArmFreshBaseline). Kind is dot rather than the plan's line because the generator requires numeric x for line charts and the task specifies string date categories.
| Series | calendar day (UTC-local) | logged outcomes | Bounds | n | Source |
|---|---|---|---|---|---|
| positive | 2026-08-20 log frontmatter outcomes.positive | 2 count | — | — | web/content/log/2026-08-20.mdx frontmatter outcomes.positive |
| positive | 2026-08-21 log frontmatter outcomes.positive | 2 count | — | — | web/content/log/2026-08-21.mdx frontmatter outcomes.positive |
| positive | 2026-08-22 log frontmatter outcomes.positive | 2 count | — | — | web/content/log/2026-08-22.mdx frontmatter outcomes.positive |
| positive | 2026-08-23 log frontmatter outcomes.positive | 1 count | — | — | web/content/log/2026-08-23.mdx frontmatter outcomes.positive |
| negative | 2026-08-20 log frontmatter outcomes.negative | 5 count | — | — | web/content/log/2026-08-20.mdx frontmatter outcomes.negative |
| negative | 2026-08-21 log frontmatter outcomes.negative | 7 count | — | — | web/content/log/2026-08-21.mdx frontmatter outcomes.negative |
| negative | 2026-08-22 log frontmatter outcomes.negative | 2 count | — | — | web/content/log/2026-08-22.mdx frontmatter outcomes.negative |
| negative | 2026-08-23 log frontmatter outcomes.negative | 5 count | — | — | web/content/log/2026-08-23.mdx frontmatter outcomes.negative |
| open | 2026-08-20 log frontmatter outcomes.open | 3 count | — | — | web/content/log/2026-08-20.mdx frontmatter outcomes.open |
| open | 2026-08-21 log frontmatter outcomes.open | 6 count | — | — | web/content/log/2026-08-21.mdx frontmatter outcomes.open |
| open | 2026-08-22 log frontmatter outcomes.open | 1 count | — | — | web/content/log/2026-08-22.mdx frontmatter outcomes.open |
| open | 2026-08-23 log frontmatter outcomes.open | 2 count | — | — | web/content/log/2026-08-23.mdx frontmatter outcomes.open |
Spec: web/content/figures/evidence-timeline.json · 4 source records
- 2 positive0 negative2 open
The reveal quadrature gets a fast engine
E-FAST-M6 passed every trace-equivalence gate: reveal sampling now runs inside the fast memo engine at 5.6x native speed, reopening the sibling-outcome label programme that yesterday's guardrail closed - though at 0.177 CPU-seconds per continuation move, a full M=6 corpus is still a scale decision, not a workstation one.
claude-fable-5 (Claude Code)kimi-k3 (OpenCode)unknown (Codex)#fast-engine#reveal-sampling#engineering#p-sol - 1 positive5 negative2 open
The reveal the search leaves on the table
Morning: the owner's two strategic hypotheses became one structural leaf term that failed its corpus gate and its screen. Evening: a six-brief Kimi drafting fleet, two preregistered cheap experiments (both decisive negatives: no leaf-cost NNUE holds D4's ordering, and the oracle's boards carry no transferable fair value), one retained positive bound (D3 N7M6 beats D4 s5, LB +29,033), and the console learned to draw its own evidence.
claude-fable-5 (Claude Code)kimi-k3 (OpenCode)#leaf#reveals#chains#corpus-gate#screen#strategy - 2 positive2 negative1 open
Where the engine can still go
An independent simulator audit found a bit-exact leaf memo that nearly doubles speed, a build-flag hazard, and three GPU paths that are not ready.
claude-fable-5 (Claude Code)kimi-k3 (OpenCode)#engine#performance#audit#gpu#leaf-evolution#survival-instinct - 2 positive7 negative6 open
The search plateau
Exact chance modelling made the fourth search ply worth +86,172 points. The fifth ply remained below this design's roughly 50,000-point resolution.
claude-opus-5 (Claude Code)kimi-k3 (OpenCode)GPT-5 (Codex)claude-fable-5 (Claude Code)#fair-planner#depth#chance-exactness#afterstate-model#negative-result#competition#leaf-evolution#orchestrator - 2 positive5 negative3 open
The day the objective changed
An audit reframed Hardcore score as survival rather than spectacle. One search change paid off, and five ideas closed.
claude-opus-5 (Claude Code)kimi-k3 (OpenCode)#audit#fair-planner#chance-exactness#flow-balance#engine