Ten of seventeen entries played no games
Every number below comes from the public data contract: register.json, methodik.json and epochen.json in the snapshot of 2026-09-05, with the register ID named beside it. Counts and sums are taken over the entries T-0054 to T-0070 and their stages; no other quantity is derived. Fingerprints, commits, machine names and authorship fields are in the export but are not reproduced here.
Between 2 and 5 September the register gained seventeen entries. Seven of them played 174,984 games; the other ten played none. Four in all record a measured gain.
Between 2 and 5 September the register went from T-0054 to T-0070: seventeen new entries in four days. Seven of them played games. The other ten played none at all.
Those ten are not missing data. They were settled another way.
The five paths in
An entry earns its verdict on one of several paths, and the register names which (register.json, methodik.json):
| Path | Entries | Games | What the entry asserts |
|---|---|---|---|
| games — SPRT | 7 | 174,984 | this build plays better, or does not, against its predecessor |
| correctness | 4 | 0 | every difference from the predecessor is listed and accounted for |
| epoch anchor | 3 | 0 | conditions changed here; do not compare across this line |
| identity | 2 | 0 | this change produced a byte-identical binary |
| takeover | 1 | 0 | this is the running predecessor, adopted as the starting point |
Of the seven that played games, four passed and three were rejected. So of seventeen entries added in four days, four record a measured gain — and none of the other thirteen claims one.
What the entries without games do
They are not smaller versions of a strength test. They answer a different question, and they answer it in a way a strength test cannot.
Identity. T-0057 and T-0063 each proposed the same thing: that adding a comment line to the engine source leaves the compiled binary unchanged. The register records the verdict as binaer_sha256-gleich — the fingerprints of the two binaries matched — and both entries report a total duration of zero seconds. Games are the wrong instrument for this question. They could at best give slow, indirect evidence that two builds play alike, and playing alike is not what is being claimed: the claim is that the compiled file is the same file. An equal fingerprint settles that directly, in no time at all.
That the same hypothesis appears twice, once in epoch 2 and once in epoch 3, is worth noticing. It is the identity path being checked after the conditions changed, rather than assumed to have survived them.
Correctness. T-0058 through T-0061 took the correctness path, which the published method describes as an independent differential gate (methodik.json). The gate compares a candidate against its predecessor over a fixed corpus of positions and produces a manifest of those where the two behave differently. What it establishes is equality across that corpus, which is narrower than equality everywhere. Three of the four report an empty manifest: bounding an evaluation value to a fixed range, checking the plausibility of a field when reading a position, and saturating counters at their maximum each left play untouched.
The fourth did not. T-0059 corrects the counting of knights in the draw rule, and its entry records behavioural equality except for three positions in the manifest. It passed anyway — and that is the path working, not failing. A correction is meant to change behaviour, in exactly the cases that were wrong before. What the gate asks is not whether anything changed, but whether everything that changed is listed. An entry claiming to fix a rule while reporting an empty manifest would be the one worth a second look.
Epoch anchors. T-0056, T-0062 and T-0070 mark the three points where measuring conditions changed — a toolchain change, then twice a change in how measurements are made (epochen.json), and they are the subject of an earlier entry. T-0070 is dated 5 September and opens epoch 4, which at the time of this snapshot contains one entry. An anchor makes no claim about play at all; it draws a line and says that results either side of it belong to different series.
Takeover. T-0054 is the running predecessor, entered as the starting point of the series. It publishes no stage either.
A count that catches some of it
The register carries a benchmark node count with sixteen of these seventeen entries — all but T-0054, which publishes no stage at all. For fifteen of the sixteen it is 3,537. For one, T-0068, it is 3,223, and T-0068 is the entry whose hypothesis was about changing what the search prunes.
That is a tripwire, and it is worth saying at once how little it shows. It fired once. It did not fire for the other five candidates that played games, and those were not idle changes: their hypotheses concern the transposition table, the evaluation and the ordering of the search. So an unchanged count is not evidence of an unchanged engine — the benchmark is one fixed position set, it sees some kinds of change and not others, and its silence proves nothing.
The identity path does not need it either: an equal fingerprint is the stronger statement and the count adds nothing to it. What the count offers is a check that costs no games and can still fail, as it did once here.
Why the distinction is the point
A register that grew by seventeen entries in four days sounds like four days of getting stronger. It was not. Ten entries were about keeping the record honest — proving a change inert, proving a binary unchanged, marking where conditions moved. Three more were candidates that were measured and rejected.
If the paths were not separated, activity would read as progress, and a busy week would look like a strong one.
What this does not say
It does not say the four accepted candidates add up to anything. Each was measured against its own predecessor under its own conditions, three of the seventeen entries mark a change in those conditions, and the register does not carry a figure for the series as a whole.
It does not say the entries without games were free. They cost review, building and gate runs; what the register reports as zero there is games played, and it puts no figure on the effort.
And it does not say this mix is typical. Four days, one register, one project — and a stretch in which four of seventeen entries happened to be measured gains.