Overnight: sixty generations, one re-selection, one screen

The unattended chain took over at 22:42 UTC yesterday once the last teacher game had finished (177 games, 21,618 labelled roots) and ran the rest of the protocol without intervention: the supervised warm start in 14 seconds, sixty generations of evolution in 3.9 hours, the elite re-selection in four minutes, and the held-out screen in eleven. Every one of the 60 generation artifacts and the screen artifact records zero illegal and zero incomplete decisions. The result is RS-20260903T025751Z-6577b33e; compact evidence, including the frozen candidate and the warm-start weights, is promoted with a manifest under artifacts/results/EX-20260902-nnue-evolution-d3-v2-49c18bc2/RUN-20260902T035644Z-c1fd8987/.

negative

The warm start fits the teacher's values but not its ordering

top-1 agreement 0.441 inside the real depth-3 search

Sixteen epochs of Adam on a whole-origin split (17,641 training roots, 3,977 validation) reached validation Huber 0.687 rise units at epoch 8 with Pearson 0.950. Placed inside the deployed depth-3 search on 256 held-out roots, that network picked the depth-5 teacher's column 44.1% of the time with a mean teacher-value regret of 2,395 points. For scale, the leaf-swing diagnostic on a different sample put the frozen fair leaf at 72.0% and a leaf that returns zero at 45.0%. On the training blocks the warm start then played at about half the fair leaf's level (159,811 against 344,537 in generation 0). The theory's own diagnostic threshold for this leg was 0.30, which the probe clears, but the number says imitation of state values bought almost no ordering.

positive

Whole-game evolution with common random numbers moves the leaf

+35,375 over the warm start on the held-out games, lower bound +16,899

Across all sixty generations the population separated steadily from its warm start on paired seeds: ten-generation average margins of +4,682, +15,853, +22,704, +32,278, +34,037 and +40,431 points for the population mean, and +76,679 for the best candidate in the last block. The curve had not flattened at generation 60. The re-selection replayed the eight fittest on 128 fresh games and froze candidate-29 at 202,237, the finalists spanning 187,352 to 202,237 where their 32-game leaders had read 220,000 to 270,000, which is the best-of-32 bias the step exists to remove. On the 64 never-read screen games the frozen candidate averaged 190,961 against the unevolved warm start's 155,586 on the same seeds: paired +35,375, bootstrap lower bound +16,899, Student-t lower bound +16,145, detection floor 18,949, 36 wins to 28, both halves positive. The first leaf evolution in this program, 18 linear weights under CMA-ES, could not show any such movement; paired blocks, rank selection and a warm start in weight space did.

negative

The gate fails: the evolved leaf is 106,964 paired points behind the fair leaf

bootstrap 95% lower bound −146,580, W-T-L 14-0-50

Against the frozen fair leaf at the identical depth-3 configuration the candidate scored 190,961 to 297,926: paired delta −106,964, bootstrap 95% bounds −146,580 to −69,983, Student-t lower bound −145,890, both halves negative (−109,139 and −104,790), lower quartile 122,253 against 158,387. Every preregistered criterion except artifact integrity and candidate identity fails, and the training-signal falsifier had already failed: the population mean was above the fair control in 0 of the final ten generations (margin −121,896). The reference arm reproduced the standing result, fair depth 4 over fair depth 3 by +97,059 with a lower bound of +29,330. Scientific outcome fail, run valid, evidence tier public-development; the theory is assessed not-supported-as-tested and its record now points at the result.

What the negative result does and does not say

It rejects this exact configuration: this model class, a depth-5 teacher corpus of 177 games, sixteen epochs of value imitation, and sixty generations of a population of thirty-two at 32 paired games each, deployed at depth 3. It does not reject the design at the 512 teacher games the protocol allowed for, because the depth-5 teacher ran four to seven times slower than its pilot and the stage closed on its wall budget. It does not say evolution cannot improve a leaf; the ablation says it can, by about a quarter of the distance it had to cover. It says the distance was set by a warm start that held the teacher's numbers and not its ordering, and that no plausible number of generations at this rate would have covered it.

closed

Continue evolving from the generation-60 population

owner: claude-fable-5-1

The owner asked overnight whether the run could go to 100 generations to see whether the curve flattens, then asked for a long run: keep evolving, up to about 1,000 generations with a decaying mutation size, and stop only when the last 50 to 100 generations show no noticeable improvement. Not possible inside the first run (its frozen protocol says 60, its training lease ends exactly where the re-selection block begins, and its 24-hour budget ended with the screen), so a successor experiment was preregistered and frozen, EX-20260903-nnue-evolution-continuation-d3-f8ce9181, and launched at 03:34 UTC as RUN-20260903T032832Z-a76a6cf7 on fresh leases (training 0xa52e2000 onward, a 64-seed screen block at 0xa52ea000). It resumes from the first run's checkpointed generation-60 population, adds the first run's frozen candidate as a third paired control, anneals the relative mutation sigma from 0.05 with a 400-generation time constant to a 0.01 floor, and stops when the one-sided 95% lower bound of the fitted slope of the paired margin over the fair leaf across the last 100 generations is not positive, checked every 50 generations from the 100th, or at 1,000 generations or 72 hours. From the first run's artifacts the rule detects slopes of about 221 points per generation; the first run's slope of 676 would have kept it going. The evolve binary gained resume, baseline-control, decay and plateau flags, the screen binary an extra arm, all covered by unit tests, a smoke test and the eleven CHECK gates on the rebuilt binaries before the first leased seed was read.

proposed

A warm start trained on ordering rather than value

owner: claude-fable-5-1

The corpus is sibling-complete, so a ranking loss over each root's seven column values is available at no extra cost, and the deployment-faithful probe already exists to score it. Whether such a warm start starts closer to the fair leaf is the cheapest question this result leaves open.

Afternoon: the continuation ran, plateaued, and was screened

The continuation launched at 03:34 UTC ran 150 generations at about 304 seconds each and stopped itself at 16:15 UTC on the preregistered plateau rule; re-selection, the one-shot screen on the new held-out block and the analysis finished by 16:31. Every one of the 150 generation artifacts and the screen artifact records zero illegal and zero incomplete decisions. Result RS-20260903T163321Z-733076b5; compact evidence, including the frozen candidate, the plateau log and the launch script, is promoted with a manifest under artifacts/results/EX-20260903-nnue-evolution-continuation-d3-f8ce9181/RUN-20260903T032832Z-a76a6cf7/.

negative

The curve rose for about a hundred more generations, then levelled

plateau rule stopped the run after generation 149

At the first check, after generation 99, the fitted slope of the paired margin over the fair leaf across generations 0 to 99 was +239 points per generation with a one-sided 95% lower bound of +30, so the run continued. At the second, across generations 50 to 149, it was +37 per generation with a lower bound of −159, no detectable improvement, and the run stopped. The fifty-generation averages of the population mean read 204,748, 228,706 and 237,946, and the last five ten-generation blocks sat between 232,000 and 240,500 with no trend. The margin over the first run's frozen candidate on the same seeds rose from +8,982 to +36,076 to +45,613 across the three fifty-generation blocks; the margin over the fair leaf narrowed from −116,503 to −88,948. The best candidate beat the fair control on 9 of 150 blocks, the population mean on none. Sigma had annealed from 0.050 to 0.034 when the rule fired, so the design cannot say whether the leaf class or the step size levelled the curve.

positive

Out of sample, 150 more generations were worth about 36,000 points

continued candidate minus the first run's candidate +36,278, lower bound +9,085

Re-selection on 128 fresh games froze candidate-12 at 264,466 (finalists 228,212 to 264,466). On the 64 never-read screen games it averaged 251,667 against the first run's frozen candidate's 215,389 on the same seeds: paired +36,278, bootstrap lower bound +9,085, Student-t lower bound +8,543, detection floor 27,330, 41 wins to 23, though the first half of the block alone was −7,140. The first run's candidate reproduced its earlier result on the fresh seeds, −104,719 against the fair leaf where the first screen had read −106,964, which is the nearest thing to a replication of that number the program has. Over the warm start the continued candidate is +107,994.

negative

The gate fails again: 68,441 paired points behind the fair leaf

bootstrap lower bound −112,090, W-T-L 25-0-39

Against the frozen fair leaf at the identical depth-3 configuration the continued candidate scored 251,667 to 320,108: paired −68,441, bootstrap bounds −112,090 to −26,694, both halves negative (−112,383 and −24,500), lower quartile 171,440 against 189,414. Every screen criterion fails. Scientific outcome fail, run valid, evidence tier public-development. The reference arm gave fair depth 4 over fair depth 3 by +71,799 (lower bound +25,857). The theory record now carries both results.

running

Separate the plateau from the annealing

owner: claude-sonnet-5

A third experiment under the same theory, EX-20260903-nnue-evolution-continuation2-d3-80eebad3, tests this directly: it resumes from the exact population where the second experiment stopped (generation 150) and repeats the identical process with a 1,500-generation mutation-decay time constant instead of 400 (3.75x slower), so sigma sits at 0.045 rather than 0.034 at the equivalent point and does not reach its floor within 1,000 generations. The immediately prior candidate plays every block as the third control, the plateau rule and its parameters are unchanged for direct comparability, and no code changes were needed — the flags this design uses survived the owner's mid-run robustness fix intact, confirmed by reading the current source and a smoke test. Launched 2026-09-03T19:10:49Z as run RUN-20260903T190433Z-a87fd7fc on fresh leases (training 0xa52ea100 onward, a 64-seed screen block at 0xa52f2100). Doubling the games per candidate as sigma shrinks, to test whether selection is steered by noise once candidates become similar, remains a further open, unregistered direction.

Housekeeping

The approach page carries the result in its opening callout and its figures now show every stage, including the screen's per-game paired differences and the gate table. The experiment index row and the closed-directions table in the status document are updated. The frozen candidate (candidate-weights.bin, SHA-256 beginning edd0d2ef) and the warm start (init.bin) are committed with the promoted artifacts so continuation work has its baseline vectors in the repository.

A code-review pass then found two interruption-only faults in the continuation driver: a completed generation could be stranded without its progress row or plateau marker, and restarting the evolve stage granted a fresh wall allowance. The repair makes generation artifacts atomic recovery records, reconstructs the compact progress and plateau indexes before reading another seed, and pins one durable evolve deadline across resumes. Focused regression tests exercise both interrupted-write windows and the non-resetting deadline; no gameplay data or protected/final cohort was opened for this maintenance work. The already launched run remains attributable to its recorded source commit; if that process is interrupted, its disposition must be recorded before changing the executable used for a resume.

The console, rebuilt around techniques rather than directories

The site's own structure was the second piece of work today, and it turned up a correction to two of our own labels.

Until now /approach listed twelve repository directories, so a reader looking for Q-learning found it split across three families, and the engines, harnesses and measuring instruments sat in the same list as the strategies. The directories are unchanged, but each approach README now declares what it is (kind: strategy, engine or diagnostic) and, for a strategy, which of fourteen techniques it uses. The console reads those keys: 86 strategy pages group by technique behind filter chips, 12 engine directories gather under a new Engines section with a comparison table whose every cell is a recorded figure or "not recorded", and 16 diagnostics have their own index. Fourteen technique primers explain each idea with a worked example that is not Drop7 before the reader meets a Drop7 page that uses it. Every one of the 86 strategy cards carries an animated drawing of its own mechanism, and the drawings run continuously in the index rather than waiting for a hover.

The other half of the brief was that a research page should read as prose to a newcomer without losing anything an agent needs. Fifteen approach pages were rewritten to five plain sections, with record identifiers, seed leases, gate tables, full arm tables and scoring-mode caveats moved into click-to-open accordions marked with a bot icon. Nothing was deleted; it moved one click down. scripts/check-approach-frontmatter.mjs now refuses any status, evidence, reads, kind or technique value outside the closed vocabulary, and .agents/skills/drop7-writing-style/SKILL.md fixes the voice the pages are held to.

negative

Two approach pages were carrying a status their own records contradict

two labels corrected, no gameplay data read

Normalising eleven off-vocabulary frontmatter values meant reading each page against the records it cites, and two disagreed with themselves. chain-reveal-leaf was labelled completed and repository-verified while its own technical record cites two results that are both valid and fail and a theory assessed not-supported-as-tested; it is now rejected and ledger-recorded. leaf-reweight was labelled Exploratory with evidence development while finding-14 records a clean negative; it is now rejected and reproduced, and its reads label is teacher, because its refit target is the clairvoyant label that finding-10 describes. The nine other corrections were vocabulary only, such as evidence: machine-readable records becoming ledger-recorded. No cohort was opened and no number changed anywhere; the labels now match the records that were already there. The body of chain-reveal-leaf still carries an evidence label of its own that disagrees with its frontmatter, which is the next thing to fix on that page.

proposed

The strategy pages not yet rewritten

owner: claude-opus-5

Fifteen approach pages now use the five plain sections and the accordions. Seventy-one others still carry their old prose and will move over as they are next touched. The card art is finished: all 86 strategy pages have a drawing of their own.

The card art started as one picture per technique, which meant the ten Q-learning pages, and the eleven heuristic-evaluation pages, each showed the same drawing as their nine or ten siblings. Seventy-five new arts close that: every strategy page now animates its own mechanism, and the ones about something that happens on the board draw the board. The evaluation pages get the most out of this. transition-rewards runs a real chain reaction, a clear paying +7, the gray beside it cracking, the survivors falling and the second wave paying +39; accessible-energy sweeps a position and lights only the discs still reachable, leaving the entombed ones dark; survival-instinct refuses a column where the disc could never match its own count and plays the one where it can.

The scores in those drawings come from the engine's own scoreForWave through a new board kit (web/components/technique-art/board.tsx), so an art names a wave depth and never types a number. Two checks now hold the rest of the contract: web/scripts/check-art.mjs fails an art with no stylesheet, with fewer animations than animated elements, with a keyframe name that could collide with another art, or with an animation shorthand that would unpause everything; web/scripts/check-tokens.mjs fails a literal colour outside the token block. The authoring contract itself is .agents/skills/drop7-web-console/references/card-art.md.

Separately, a shared link now shows the page it points at. Twenty-five routes render their own 1200x630 card from the page's own title, summary and labels beside a board the engine actually played, and the twelve dynamic ones also carry a canonical URL and their own Open Graph and Twitter text. A sitemap of 344 URLs and a robots file went in with them, with lastModified set only where the record carries a date of its own. The reusable part is .agents/skills/drop7-social-cards/SKILL.md, which also covers generating a raster asset through the Codex CLI for the cases inline SVG cannot serve.

The site builds, npm test passes at the repository root and inside web/, and a sweep of 289 routes returns 200 for every one. A fresh clone with no web/data/, no research results and no research log renders every page except /compete, which needs GitHub OAuth configuration and behaved that way before today. No research artifact, protocol or record was edited for this work, and no seed was read.

Time to read the last frame

Watching the approach grid for a while showed a fault in those drawings. Most of them end by putting up a label: the number a gray disc was hiding, the column a search picked, the caption that says what just happened. The whole loop was 2400ms, so that label was on screen for a few hundred milliseconds before the card reset and drew itself again. It was there, and it could not be read.

A cycle is now two parts. --tart-motion is the drawing and keeps the 2400ms it always had; --tart-read is a still frame after it, 2000ms by default, and an art claims it only when its last frame carries something its first frame did not. Ninety-two of the hundred and five arts qualify. Their keyframes were rewritten against the longer cycle, so every beat happens at the same moment on a clock as before and the extra two seconds are spent standing still rather than slowing the drawing down. The thirteen that end where they began, like the fallback drop and the Q-learning corridor, were left at the shorter loop.

Two more rules went into web/scripts/check-art.mjs: an art whose resting frame carries text has to declare --tart-read, and an art that has declared it may not have a keyframe stop past its motion share, which is 54.55% of the cycle at the default settings. The grid's starting offsets became fractions of each card's own cycle rather than fixed times, so the spread still covers a slow art and no two neighbours hold their last frame together.

A log entry is a narrative written by the contributors listed above. Run validity, scientific outcome and evidence tier live with the experiment and result records the entry refers to.