The verdict did not predict the cost

Every number below comes from the public data contract: register.json and methodik.json in the snapshot of 2026-09-05, with the register ID named beside it. Estimates are quoted in nElo Normalised Elo: the Elo difference divided by the spread of the game results, so that the SPRT bounds mean the same at different time controls., the unit the test bounds are set in; the register reports Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. as well. Totals are summed from the per-stage games and durations the register reports; the register also carries a per-entry duration a few seconds larger than the sum of its stages. Ratios are computed from these figures and from nothing else.

Six candidates cost 149,104 games and close to thirty hours of machine time. Whether one passed did not predict what it cost.

Between 3 and 5 September, six candidates were measured against their own predecessor. Four passed, two were rejected. Together they cost 149,104 games and 106,585 seconds of machine time — a little under thirty hours.

The six did not cost the same.

What each one cost

All six began with the same short stage: 8+0.08 on book A against their own predecessor, with bounds separating “no gain” from “a gain of 5 nElo”, log-likelihood bounds at ∓2.94, and a ceiling of 80,000 pairs (methodik.json).

CandidateGamesTimeEstimateVerdict
T-00662,3721,036 s+32.96 nElopassed
T-00686,1962,687 s−8.98 nElorejected
T-00679,7904,246 s+9.78 nElopassed
T-006514,3266,211 s+7.47 nElopassed
T-006944,79419,733 s+0.90 nElorejected
T-006448,01620,701 s+3.99 nElopassed

The cheapest run took seventeen minutes, the most expensive close to six hours: 20.2 times the games, 20.0 times the time. None of the six came near the ceiling — the largest used 24,008 of the 80,000 pairs available — and every one of them ended by crossing a log-likelihood bound rather than by running out of budget.

The order that is not in the table

Across these six the verdict did not order the cost. The cheapest run passed, the second cheapest was rejected, the second most expensive was rejected, and the most expensive passed. The two outcomes are spread across the whole range rather than sorted by it.

That is worth saying because the opposite is easy to assume — that a sound idea proves itself quickly and an unsound one drags on. Here the two rejections are the second cheapest run and the second most expensive. Those two, T-0068 and T-0069, were the subject of an earlier entry; they appear here only as two of six.

Six candidates from epoch 3, ordered by games: T-0066 with 2,372 games, T-0068 with 6,196, T-0067 with 9,790, T-0065 with 14,326, T-0069 with 44,794 and T-0064 with 48,016. Passed and rejected candidates are spread across the whole range; the two most expensive runs are the ones whose estimate falls inside the indifference region of zero to five nElo.
The six candidates by the games their short stage needed, with the estimate each ended on. Passed and rejected are spread across the range; the two longest runs are the two whose estimate fell inside the indifference region. T-0066: · passed; T-0068: · rejected; T-0067: · passed; T-0065: · passed; T-0069: · rejected; T-0064: · passed.

A reading that fits

The six can be ordered another way. What follows is a reading rather than a finding, and the reasons it could be wrong come with it. The bounds place the indifference region between 0 and 5 nElo. Call a run’s distance from that region the amount by which its estimate falls outside the nearer edge, and zero when the estimate lands inside it. Sorted by that distance, the games fall into the same order, with one tie.

T-0066 ended at +32.96 nElo, 27.96 outside the upper edge, and settled in 2,372 games. T-0068 ended at −8.98, that far outside the lower edge, in 6,196. T-0067 at +9.78 stands 4.78 outside and took 9,790; T-0065 at +7.47 stands 2.47 outside and took 14,326. T-0069 at +0.90 and T-0064 at +3.99 both land inside the region, so both have distance zero. They are tied — that one took 44,794 games and the other 48,016 is not part of the ordering.

A sequential test runs until the accumulated evidence reaches a bound, and near the indifference region each game contributes less of it. That is the shape this ordering has.

There are at least three reasons it could still be wrong. Sequential tests differ widely in when they stop even when the true effect is the same, and across six runs a monotone ordering can arise by chance. The same likelihood drives both the stopping time and the estimate, so the sorting variable is not independent of what it is being used to explain; in part it restates the stopping rule. And each candidate ran against its own predecessor, so the six did not face a common opponent, and their draw rates and variances need not match.

One further thing the table does not show. The verdict does not follow from where the estimate landed: T-0064 passed at +3.99 nElo and T-0069 was rejected at +0.90, both inside the region. What decides is which log-likelihood bound the accumulated evidence reaches, not the position of the point estimate.

The number that is not here

The natural summary of all this would be a cost per nElo point. This data does not carry one.

Two of the six were rejected, so neither carries an accepted gain to divide by. For the four that passed there is no single figure either, because the two stages disagree: T-0064 measured +3.99 nElo short and +15.7 long, T-0067 +9.78 and +21.79, T-0066 +32.96 and +23.8, T-0065 +7.47 and +8.98 (register.json). The stages run at different time controls on different books, so they answer different questions, and neither is the effect of the change.

Both are also point estimates from tests that stopped at a bound. Dividing hours by a figure that carries that bias would produce something shaped like a rate but not usable as one.

What confirmation costs

The four that passed went on to a long stage — 40+0.4 on book B — spending 23,610 games and 51,971 seconds between them. The two rejected candidates never reached it.

That is where the two shares part. The rejected pair accounts for 50,990 of the 149,104 games, 34.2 %, but for 22,420 of the 106,585 seconds, 21.0 %. A short-stage game averaged 0.435 seconds of elapsed time and a long-stage game 2.201, a factor of 5.06 against a time-control ratio of 5. Those averages are elapsed time with sixty games running at once (methodik.json), not the length of a game on its own clock.

The expensive stage is spent only on what has already survived a cheaper one. That is a property of the arrangement, not an accident of these six.

What this does not say

It does not say what progress costs. Six candidates, one epoch, one machine, three days, and only the runs that were recorded as finished.

It does not say the four accepted changes are worth +3.99, +7.47, +32.96 and +9.78 nElo. Each is a point estimate from a test that stopped when a bound was reached, and is biased by that stopping rule.

It does not say the two rejected runs were the waste in this total. They cost about a fifth of the time and returned two answers; whether that is a good price is not something the register can settle.

And it does not say a seventh candidate can be budgeted from these six. What these six show is a twentyfold spread that the verdict did not predict; what produced it is a question the register cannot settle.