Drop7 Research
← Docs
docs/exploratory/finding-04-terminal-utility-saturated.md

Finding 04 — The fair-D4 risk constant is saturated; it is not a lever

Status: exploratory, evidence tier development. Valid negative result. Namespace: approaches/lifetime-objective/risk-calibration, run runs/RUN-A51D-risk/, lease SEEDLEASE-A51D, seeds 0xa51d00000xa51d003f. No existing file was modified. The frozen reference is untouched; this is a new parameterized candidate that reproduces it exactly at default settings.

Question

research/benchmarks/baselines-v1.json pins terminalUtility = -1000000. An independent audit of the reference (audit-02-fair-d4.md) established that leaf points are score pointskImmediateScoreWeight = 1.0 — so −1,000,000 is 58.8 row rises, or 294 moves, or 3.2× the entire recorded mean score. Against a measured non-death sibling spread of 217–10,924 leaf points, the audit concluded the decision rule is effectively lexicographic: minimize modeled four-ply death probability first, maximize the heuristic leaf only as a tie-break.

Since ~94% of Hardcore score is survival (finding-01), how the search prices death is the single constant most directly aimed at the objective — and it had never been varied. This experiment varies it.

Method

A parameterized re-implementation of the depth-4 driver exposes terminalUtility, depth, and chanceSamples as runtime options. Everything else — the leaf, chance stratification, canonicalization, cache keying, column order, work accounting, iterative deepening, and legal fallback — is the unmodified frozen code, included as a library.

CHECK gate before any gameplay: at default parameters the parameterized driver must select the identical column as chooseDepth4Action on every move. Result: 50 moves compared, 0 mismatches, PARITY OK.

Six arms, 64 paired games each, same seeds as the reference cohort, 2,000-move cap, common seeds across arms.

Result

Terminal utilityMean scoreMean movesMedianClears/movePaired Δ vs referenceW–T–L95% lower bound on Δ
−50,000,000321,99294.06266,2821.97300–64–00
−10,000,000321,99294.06266,2821.97300–64–00
−3,000,000321,99294.06266,2821.97300–64–00
−1,000,000 (frozen)321,99294.06266,2821.9730
−300,000321,65293.98266,2821.973−3401–62–1−1,039
−100,000320,16192.92264,9421.987−1,8314–29–31−4,563

The flat segment is the finding, and it is easiest to see drawn:

Terminal utility is saturated — byte-identical play across a fiftyfold increase in death cost Terminal utility is saturated — byte-identical play across a fiftyfold increase in death cost. Six arms, 64 paired whole games each on seeds 0xa51d0000-0xa51d003f, depth 4, corrected 17,000-point Hardcore scoring (finding-04 Result table). The four arms at magnitude 1,000,000 and beyond produced BYTE-IDENTICAL games to the frozen reference — paired delta exactly 0, W-T-L 0-64-0 — so those points are exact and carry no statistical bound; none is drawn and none exists. The two smaller magnitudes are single-cohort estimates: -300,000 scores 321,652 (paired delta -340, one-sided 95% lower bound -1,039, W-T-L 1-62-1) and -100,000 scores 320,161 (delta -1,831, lower bound -4,563, 4-29-31). Valid negative result, evidence tier development: the risk constant is at its stop and cannot be made a lever; the figure plan allowed a dot form or a line on log|utility| — the dot form is used because the generator draws linear axes only and string categories keep the fiftyfold range readable. Sources: docs/exploratory/finding-04-terminal-utility-saturated.md. Terminal utility is saturated — byte-identical play across a fiftyfoldincrease in death cost320,000320,500321,000321,500322,000-100,000-300,000-1,000,000(frozen)-3,000,000-10,000,000-50,000,000mean score (points)terminal utility (categories ordered by ascending magnitude; frozen value marked)mean score, 64 paired games per arm — -100,000 | paired delta -1,831 vs frozen, 95% LB -4,563, W-T-L 4-29-31; 35 | of 64 games differ | mean score: 320,161 points | n = 64 games | source: | docs/exploratory/finding-04-terminal-utility-saturated.md · | Result tablepaired delta -1,831 vs frozen, 95% LB -4,563, W-T-L 4-29-31; 35 of 64 games differmean score, 64 paired games per arm — -100,000paired delta -1,831 vs frozen, 95% LB -4,563, W-T-L 4-29-31; 35of 64 games differmean score: 320,161 pointsn = 64 gamessource:docs/exploratory/finding-04-terminal-utility-saturated.md ·Result tablemean score, 64 paired games per arm — -300,000 | paired delta -340, 95% LB -1,039, W-T-L 1-62-1 | mean score: 321,652 points | n = 64 games | source: | docs/exploratory/finding-04-terminal-utility-saturated.md · | Result tablepaired delta -340, 95% LB -1,039, W-T-L 1-62-1mean score, 64 paired games per arm — -300,000paired delta -340, 95% LB -1,039, W-T-L 1-62-1mean score: 321,652 pointsn = 64 gamessource:docs/exploratory/finding-04-terminal-utility-saturated.md ·Result tablemean score, 64 paired games per arm — -1,000,000 (frozen) | the frozen reference value; 94.06 mean moves, 1.973 clears/move | mean score: 321,992 points | n = 64 games | source: | docs/exploratory/finding-04-terminal-utility-saturated.md · | Result tablethe frozen reference value; 94.06 mean moves, 1.973 clears/movemean score, 64 paired games per arm — -1,000,000 (frozen)the frozen reference value; 94.06 mean moves, 1.973 clears/movemean score: 321,992 pointsn = 64 gamessource:docs/exploratory/finding-04-terminal-utility-saturated.md ·Result tablemean score, 64 paired games per arm — -3,000,000 | byte-identical games to the frozen arm; delta exactly 0, W-T-L | 0-64-0 | mean score: 321,992 points | n = 64 games | source: | docs/exploratory/finding-04-terminal-utility-saturated.md · | Result tablebyte-identical games to the frozen arm; delta exactly 0, W-T-L 0-64-0mean score, 64 paired games per arm — -3,000,000byte-identical games to the frozen arm; delta exactly 0, W-T-L0-64-0mean score: 321,992 pointsn = 64 gamessource:docs/exploratory/finding-04-terminal-utility-saturated.md ·Result tablemean score, 64 paired games per arm — -10,000,000 | byte-identical games to the frozen arm; delta exactly 0, W-T-L | 0-64-0 | mean score: 321,992 points | n = 64 games | source: | docs/exploratory/finding-04-terminal-utility-saturated.md · | Result tablebyte-identical games to the frozen arm; delta exactly 0, W-T-L 0-64-0mean score, 64 paired games per arm — -10,000,000byte-identical games to the frozen arm; delta exactly 0, W-T-L0-64-0mean score: 321,992 pointsn = 64 gamessource:docs/exploratory/finding-04-terminal-utility-saturated.md ·Result tablemean score, 64 paired games per arm — -50,000,000 | byte-identical games to the frozen arm; delta exactly 0, W-T-L | 0-64-0 — fifty times the frozen magnitude changes no decision | mean score: 321,992 points | n = 64 games | source: | docs/exploratory/finding-04-terminal-utility-saturated.md · | Result tablebyte-identical games to the frozen arm; delta exactly 0, W-T-L 0-64-0 — fifty times the frozen magnitude changes no decisionmean score, 64 paired games per arm — -50,000,000byte-identical games to the frozen arm; delta exactly 0, W-T-L0-64-0 — fifty times the frozen magnitude changes no decisionmean score: 321,992 pointsn = 64 gamessource:docs/exploratory/finding-04-terminal-utility-saturated.md ·Result tablemean score, 64 paired games per arm
Mean score against the terminal utility on a log-magnitude axis. The four identical points across a fiftyfold parameter range are byte-identical games — the search is already at its risk-averse limit — and only softening the penalty changes anything, for the worse.
Source data

Six arms, 64 paired whole games each on seeds 0xa51d0000-0xa51d003f, depth 4, corrected 17,000-point Hardcore scoring (finding-04 Result table). The four arms at magnitude 1,000,000 and beyond produced BYTE-IDENTICAL games to the frozen reference — paired delta exactly 0, W-T-L 0-64-0 — so those points are exact and carry no statistical bound; none is drawn and none exists. The two smaller magnitudes are single-cohort estimates: -300,000 scores 321,652 (paired delta -340, one-sided 95% lower bound -1,039, W-T-L 1-62-1) and -100,000 scores 320,161 (delta -1,831, lower bound -4,563, 4-29-31). Valid negative result, evidence tier development: the risk constant is at its stop and cannot be made a lever; the figure plan allowed a dot form or a line on log|utility| — the dot form is used because the generator draws linear axes only and string categories keep the fiftyfold range readable.

Seriesterminal utility (categories ordered by ascending magnitude; frozen value marked)mean scoreBoundsnSource
mean score, 64 paired games per arm-100,000 paired delta -1,831 vs frozen, 95% LB -4,563, W-T-L 4-29-31; 35 of 64 games differ320,161 points64docs/exploratory/finding-04-terminal-utility-saturated.md Result table
mean score, 64 paired games per arm-300,000 paired delta -340, 95% LB -1,039, W-T-L 1-62-1321,652 points64docs/exploratory/finding-04-terminal-utility-saturated.md Result table
mean score, 64 paired games per arm-1,000,000 (frozen) the frozen reference value; 94.06 mean moves, 1.973 clears/move321,992 points64docs/exploratory/finding-04-terminal-utility-saturated.md Result table
mean score, 64 paired games per arm-3,000,000 byte-identical games to the frozen arm; delta exactly 0, W-T-L 0-64-0321,992 points64docs/exploratory/finding-04-terminal-utility-saturated.md Result table
mean score, 64 paired games per arm-10,000,000 byte-identical games to the frozen arm; delta exactly 0, W-T-L 0-64-0321,992 points64docs/exploratory/finding-04-terminal-utility-saturated.md Result table
mean score, 64 paired games per arm-50,000,000 byte-identical games to the frozen arm; delta exactly 0, W-T-L 0-64-0 — fifty times the frozen magnitude changes no decision321,992 points64docs/exploratory/finding-04-terminal-utility-saturated.md Result table

Spec: web/content/figures/terminal-utility-saturation.json · 1 source record

Interpretation

The constant is saturated. Every value at or beyond −1,000,000 produces byte-identical games — 0 wins, 64 ties, 0 losses, across a fiftyfold increase in magnitude. The search is already at its risk-averse limit: making death 50× more expensive changes no decision anywhere in 64 games.

Moving in the other direction only costs. At −300,000 the policy still agrees with the reference on 62 of 64 games. At −100,000 it finally disagrees often (35 of 64 games differ) and is worse: −1,831 mean, with a 95% lower bound of −4,563, losing 31 games and winning 4.

Note the direction of the flow rates at −100,000: clears per move rise slightly (1.987 vs 1.973) while lifetime falls (92.92 vs 94.06 moves). A less death-averse policy clears marginally more per move and dies sooner. Throughput alone is not the objective; throughput sustained without dying is.

What this rules out, and what it points at

This closes a cheap hypothesis: fair D4 cannot be made to survive longer by re-pricing death. The knob is at its stop. Any remaining gap between D4's 94-move mean and the ~294 moves a million-point mean requires must come from somewhere else.

That "somewhere else" is now better constrained. The search already minimizes death within its horizon as its first priority. Its horizon is four plies, and a row rise occurs every five moves, so it observes at most one rise boundary. The quantity that actually kills it — a clearance deficit of 1.973 against the required 2.400 clears per move, accumulating over the forty-plus rise cycles that separate a 94-move game from a 294-move one — is invisible at that horizon and is carried entirely by a hand-tuned leaf that was calibrated to nothing.

The lever is therefore the leaf's long-horizon content, not the search's risk appetite. Two follow-ups are motivated directly by this result and by audit-02:

  1. Chance strata. Five stratified samples cannot represent a seven-atom disc distribution; the audit measured that an average 2.41 of 7 disc values receive zero weight at each node, deterministically. Independently, audit-03 found that the historical seven-stratum rejection reverses sign under corrected scoring (−164 → +4,336), because it lost on score while gaining moves. Both point at the same arm.
  2. A learned long-horizon leaf predicting survival rather than score. This is where the remaining headroom must be, since the risk term is exhausted and the horizon is structurally too short.

Limitations

  • 64 paired games on one exploratory lease; the identical-game result for magnitudes ≥1,000,000 is exact and needs no statistics, but the −300,000 and −100,000 deltas are single-cohort estimates.
  • Only the terminal utility was varied. The leaf weights are held at their frozen values; a jointly re-tuned leaf could in principle move the saturation point.
  • These are the repository simulator's semantics, including the two rise-boundary scoring discrepancies documented in audit-01-engine-fidelity.md.

Reproduce

./approaches/lifetime-objective/risk-calibration/build.sh
./build/lifetime/risk-calibration --parity --seed-start 0xa51d0100 --parity-games 2 --parity-moves 25
for tu in -100000 -300000 -1000000 -3000000 -10000000 -50000000; do
  ./build/lifetime/risk-calibration --terminal-utility $tu --depth 4 \
    --seed-start 0xa51d0000 --games 64 --max-moves 2000 --threads 32 \
    --output runs/RUN-A51D-risk/tu${tu}.json
done
For a walkthrough with board animations, start at how the game works and the concepts primer; every term is defined in the glossary.