Two runs that were meant to fail
Every number below comes from the public data contract: register.json in the snapshot of 2026-09-05, with the register ID given beside it. Percentages and ratios are computed from the register fields named beside them; the arithmetic is shown where it is used.
An engine playing itself should measure nothing. Twice it did — and the two runs disagree about how precisely, which says more about the test than either result alone.
Most entries in the register test whether a change helped. Two of them test something else: whether the apparatus that answers that question works at all. They put the engine up against an unchanged copy of itself, where the correct answer is known in advance — no difference — and check that the machinery finds it.
Both were rejected, which is what success looks like here.
What the two controls measured
T-0001 | T-0055 | |
|---|---|---|
| Date | 2026-08-28 | 2026-09-02 |
| Opponent | an unchanged copy of v0 | the predecessor source against itself |
| Elo | −0.50 ± 1.71 | −0.17 ± 2.99 |
| Games | 17,984 | 25,880 |
| Terminal LLR | −2.97 | −2.94 |
| Verdict | rejected | rejected |
Both ran at 8+0.08 with thresholds set to separate “no gain” from “a gain of 5 nElo Normalised Elo: the Elo difference divided by the spread of the game results, so that the SPRT bounds mean the same at different time controls.” (elo0 0.0, elo1 5.0 nElo, log-likelihood bounds at ∓2.94; T-0001 and T-0055, register.json). The bounds are stated in normalised Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating.; the estimates quoted below are logistic Elo, which is a different scale. Both estimates sit within half an Elo point of zero, and both runs reached the lower bound. A control that came out any other way would have been the interesting one.
Both estimates come from tests that stopped when a bound was reached, so they are biased by that stopping rule, and the intervals beside them are nominal 95 % intervals rather than guarantees that hold under sequential stopping. They are quoted here as the register reports them, not as unbiased effect sizes.
The register records the opponent role explicitly for T-0055 — the field holds the value vorgaenger, predecessor — but leaves it empty for T-0001; there, the self-comparison is documented in the published hypothesis instead — “unchanged source of reference v0”, with the expected outcome named in advance.
The pair counts are almost symmetric
A test of this kind plays openings in pairs, once from each side, and records the pair’s score. That gives five counts per run — the register carries them per stage as pentanomial (T-0001 and T-0055, register.json). The middle count is the neutral category, a pair scoring one point out of two: it holds pairs that ended one win and one loss as well as pairs that ended in two draws, and the register does not separate the two. The totals and percentages below are sums and quotients of exactly those five numbers.
| Pair result | T-0001 | T-0055 |
|---|---|---|
| both games to one side | 112 | 985 |
| both games to the other | 125 | 979 |
| split, one side ahead | 551 | 2,535 |
| split, other side ahead | 551 | 2,560 |
| neutral pair (one point of two) | 7,653 | 5,881 |
| pairs total | 8,992 | 12,940 |
Read the mirrored rows against each other. In T-0001 the two “split” counts are 551 and 551 — identical. In T-0055 they are 2,535 and 2,560, and the decisive pairs come out 985 against 979. This is worth stating carefully, because the obvious reading of it is wrong. A fair apparatus produces symmetric counts in expectation — but symmetric counts do not prove a fair apparatus. Errors that hit both sides equally, errors that cancel, and plain chance all produce the same picture, and nothing here establishes how close to symmetric these counts would land by chance alone. The observation is descriptive: no alarm, not a certificate.
The other direction does hold. An asymmetry here would point at the harness — a colour bias, an unequal clock, a book that favours one seat — before it pointed at the engine, because in self-play there is no engine difference left to explain it.
More games, wider interval
The two runs disagree about precision, and not in the direction one would guess. The reason is not that extra games hurt — it is that neither the games nor the interval were fixed in advance.
T-0055 played 25,880 games against T-0001’s 17,984 — about 44 % more. Its 95 % half-width is 2.99 against 1.71, about 75 % wider (register.json). More evidence produced a less precise answer.
The pair counts explain it. In T-0001, 7,653 of 8,992 pairs were neutral — roughly 85 %. In T-0055, 5,881 of 12,940 — roughly 45 %. The statistical unit here is the pair, not the game, and a neutral pair sits exactly at the mean: it moves the estimate not at all and adds almost nothing to the spread. A run in which most pairs are neutral therefore has a small variance per pair; a run in which half of them come apart has a large one.
These are two different distributions, not one process running longer. Within a single distribution of fixed variance, more pairs narrow the interval in expectation. Across these two, the variance per pair differs by more than the sample sizes do, and the variance wins.
And the sample size is not an independent choice. A sequential test runs until the evidence reaches a bound, so a noisier run tends to take longer to get there. The larger game count of T-0055 is therefore not a competing explanation for its wider interval; both observations are consistent with the same higher variance per pair. Two runs cannot establish that causal direction, only that the two facts fit it.
The pair counts let this be checked rather than asserted. Treat each pair as one observation scoring 0, 0.25, 0.5, 0.75 or 1, and compute the variance of those scores from the five counts above: it comes out about 4.4 times larger for T-0055 than for T-0001. The standard error of a mean scales as the square root of variance divided by sample size, so those two inputs predict a half-width ratio of about 1.75. The register reports 2.99 against 1.71 — a ratio of 1.75. Had sample size alone decided, the half-width would have fallen by a factor of about 0.83.
This is why the number of games alone says little about how sharp a result is. What matters is how much the pair outcomes scatter. What drives that scatter in these two runs is not established here.
What this does not say
It does not say the measurement chain is calibrated. Two controls that came out near zero are consistent with a working apparatus; they do not establish that it is working, and two runs cannot estimate how often a false positive occurs. That would need a series of controls, not a pair of them.
It does not say the engine is unchanged between the two dates. T-0001 compares v0 with itself, T-0055 compares a later build with itself. The two runs test the same property of the apparatus at two points in time, not the same code.
It does not say a wide interval is a defect. Both intervals contain zero, which is the outcome a control is supposed to produce. The comparison between them is about the relationship between sample size, scatter and precision — not about one run being better than the other.
And it does not say why the two runs differ in how often their pairs come out neutral. The counts show that they do, and that this accounts for the direction of the precision difference. Why the later build produces fewer neutral pairs against itself than v0 did is a separate question, and this data does not answer it.