Three thousand games, none lost on time
This entry does not draw on the public data contract. Every number in it is a project measurement, not in the register: the runs carry no register ID, appear in none of the eight files under /daten/, and received no verdict from any gate. Counts and end reasons are taken from the project's calibration record as delivered; the rule in the second-to-last section is quoted from that record as a proposal. Counts, end reasons and the details under Provenance are taken from that record as delivered. Where it states that something is not available — a physical core count, a closer description of the virtualisation, independent evidence for the claimed emulation of the time source, a separate crash counter — this entry says so rather than supplying one. The heading reports what these runs showed and is not a statement about the measuring environment in general.
A project measurement, not in the register: one build played itself three thousand times and no game ended on the clock. Under an assumption of independent games with a constant rate of loss on time in the regime measured, the runs bound how often that can happen; about playing strength they say nothing.
The runs described here are a project measurement, not in the register. They were made on the project’s own workstation, carry no register ID, appear in none of the eight published files, and received no verdict from any gate. A reader cannot check them the way the rest of this blog can be checked, and the sections below mark where that matters.
Method
Time controls are enforced by a clock. If the clock of the measuring environment misbehaves, games can end on time for reasons unrelated to how either side played, and a local pre-check built on time losses would be unreliable. The calibration was run to measure whether that occurs under the conditions in use.
The build was played against itself: the same binary on both sides, so neither side holds a structural advantage from it. The usual advantage of the first move remains. Two series were run one after the other, at a time control of 8 seconds per side plus an increment of eight hundredths of a second per move.
Provenance
The record as delivered fixes the following, and says plainly where a figure is not available.
The base was the build following T-0067, with the same bound binary on both sides. That is a reference to a register entry, not a register verdict about these series. The six-way series ran on 5 September 2026 from 08:25:23 to 10:05:25 UTC, 1,000 colour-swapped pairs making 2,000 games; the three-way series from 10:33:18 to 12:13:57 UTC, 500 pairs making 1,000 games. The timestamps mark the start of the first game and the end of the last. In both series the time control was 8 seconds of base time per side and game, with eight hundredths of a second credited after each executed legal move.
The machine is given as 32 logical processors, x86-64, virtualised. That figure comes from a processor count and a logged affinity list; it is not an evidenced count of physical cores, and no physical core count is available. No closer description of the virtualisation is available.
Two things the record does not carry, and this entry therefore does not assert. The run was prompted by a claim that the time source is emulated, and the record holds no independent evidence for that claim; what is evidenced is a virtualised environment and its logged clock source. The figure for the cost of a single time query is recorded as a note inherited from the commission, not as something measured in these sources. There is also no separate crash counter, which is why the combined category is the only thing available to read.
One detail of measurement bears on what “on time” means here. The driver measures elapsed time from sending the search command to reading the move back, including scheduling and communication, rather than the engine’s processor time alone. A loss on time is recorded when the moving side’s remaining time, reduced by that measured duration, is negative before the increment is credited.
Results
Project measurement, not in the register:
| Concurrent games | Games | Ended on time |
|---|---|---|
| 6 | 2,000 | 0 |
| 3 | 1,000 | 0 |
The end reasons are recorded in full, and are reproduced here because a category reading zero is only informative if the categories account for every game. The sums establish that no game is missing from the tally. They do not establish that games were classified correctly: a game ending on time and recorded under another reason would leave the totals intact.
Project measurement, not in the register:
| Ending | 6 concurrent | 3 concurrent |
|---|---|---|
| checkmate | 1,609 | 804 |
| insufficient material | 188 | 91 |
| threefold repetition | 136 | 63 |
| fifty moves | 46 | 30 |
| stalemate | 21 | 12 |
| on time | 0 | 0 |
| time or engine error | 0 | 0 |
| illegal move | 0 | 0 |
| total | 2,000 | 1,000 |
Both columns sum to their game count.
What the observation bounds
No game in these runs ended on time. It does not follow that the clock is sound, and the two statements should not be substituted for one another.
What zero in a fixed number of games does support is a bound, and the bound holds under an assumption. Treating the games as independent with a constant rate of loss on time, and with no such ending observed in the 2,000 games of the six-way series, the customary one-sided bound at ninety-five per cent confidence puts the underlying rate at no more than about three in two thousand games. The assumption is not exactly met: the games were played as colour-swapped pairs, six at a time, so neither independence nor a constant rate follows from the design. The figure is therefore a conditional model calculation, not an approximation for these runs; how far its coverage holds without those assumptions is not established. It applies to the regime that was measured, and it says nothing about other time controls, other machines or other concurrency settings.
The series ran sequentially rather than side by side, so background load was not held constant between them, and the additional logging that made the measurement possible is itself load that ordinary runs do not carry. The six-way and three-way series are therefore two observations under differing conditions, not a controlled comparison of concurrency.
One category is limited in what it records. “Time or engine error” is a combined category, not a crash counter. A zero there records that neither kind of report occurred; it does not record that a separate crash count was kept and came out empty.
The proposed rule
The calibration was made to support a threshold for a local pre-check that would run before a candidate reaches the register. The project record describes the following as a proposal and does not state that it has been put into effect. It concerns candidate runs, in which the two sides are different builds — unlike the calibration above, where both sides were the same build.
As proposed, a candidate in the calibrated local six-way regime would be permitted no losses on time, with no tolerance. One documented candidate loss on time would be a red result, as would illegal moves and reported time or engine errors. Every ending on time would have to carry its evidence: game, colour, engine, move, and the remaining time and duration of that move in seconds. If the losing side could not be identified, or the error log were absent, the result would be yellow and nothing would be released. A documented loss on time on the base side would be recorded separately and would not count as a candidate loss. Zero losses on time would not by itself be a pass: the other red conditions would continue to apply, and the local evidence would have to be complete.
The measurement supports one half of that threshold and not the other. It shows that zero losses on time is attainable in this regime: 2,000 games produced none. Both sides there were the same build, and no candidate was examined, so nothing follows from it about how often a correct candidate would meet the same threshold. It does not show that zero is required either: nothing here establishes what a single loss on time would imply about a candidate. Choosing red over yellow-with-investigation, and no tolerance over some, is a decision about how much risk to accept at an early gate, not a result read off the data.
What this does not say
It does not say the rule is in use. The record describes a proposal. Whether it was adopted, and what it did afterwards, is not part of this measurement.
It does not say a local pre-check substitutes for anything in the register. It would run earlier and replace no gate; a green local result would be no verdict and would produce no entry.
It does not cover runs under other conditions. Both series set one build against itself, each at one concurrency setting, at one time control. A run with a different opponent, a different concurrency setting or a different build is outside what these two series bound, and an ending on time in such a run would not contradict them.
It does not say anything about playing strength. Both sides were the same build, which is what makes the series suitable for checking the apparatus and unsuitable for ranking anything. No Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. figure is derived from it.
And it does not say the clock is sound in general. Three thousand games at one time control, on one machine, on one day, at two concurrency settings, with the bound above applying only to the larger of the two series. Faults rarer than that bound are not ruled out by these runs.