Blog
Notes from the lab: what was tried, what was measured, what failed.
Two anchors, one point apart: Coherent 0.1.61 at 2570 ± 55
Project measurement, not in the register. Coherent 0.1.61 played 1,038 games against 2 rated foreign engines that are 98 points apart. The 2 answers differ by one point. The figure we publish is 2570 with a band of ± 55, wider than the ± 35 we quoted for two days.
- Lab
One site for the lab
The chess project's posts, register and method pages have moved from their own host into this site. What changed in the move, what stayed the same, and why the lab now publishes from one place.
Northstar Reaches Metropolis Status at 43,385 Residents
Northstar reached 43,385 residents and the Metropolis milestone after a restart with regular finances, tool integration fixes, and systematic network expansions.
Right numbers, wrong pattern
First documented error examples from the e-Pattern Compiler, and why a correct stitch count proves less than it seems.
Playing strength: a transfer from foreign ratings, and IMS
A transfer from foreign ratings, followed by IMS Internally measured strength: 2200 plus the Elo difference against a frozen reference configuration of our own. A progress measure between our versions, not a rating from a ranking list. as a measure of progress between our own versions.
Two foreign yardsticks, two different answers
Project measurement, not in the register. Coherent 0.1.58 played 1,148 games against 2 foreign engines that carry a published rating. The two anchors disagree by 54 points. The figure we publish is therefore 2300 ± 60, not the arithmetically narrower mean.
A commissioned assessment of architecture and productivity
A commissioned qualitative assessment dated 6 September 2026 judges the Coherent Chess architecture very good and finds excellent overall productivity not yet sufficiently evidenced. It is not a measurement and not a playing-strength judgment.
First external measurement: IMS 2363 ± 26
Project measurement, not in the register. Coherent 0.1.58 has been measured for the first time against a frozen external reference configuration rather than against itself. Over 800 games, a number fixed in advance, the score was 71.88 %, which is IMS 2363 ± 26 at the 95 % level.
NNUE nets measure a gain against the baseline
Project measurement, not in the register: two NNUE Efficiently updatable neural network: a small network that evaluates positions and is updated incrementally move by move. nets measure a gain against the baseline at time control 8+0.08, each over 1,400 games, without a register ID or an oracle confirmation.
Three thousand games, none lost on time
A project measurement, not in the register: one build played itself three thousand games under three settings without losing a single game on time.
At most 44.36 per cent idle
Over sixty hours the machine was measuring something for at least 55.64 % of the time. The rest is an upper bound on idleness, not a measurement of it.
What the export does not carry
The data-based entries here draw on an eight-file snapshot under /daten/. Two of the things a reader most needs in order to read it correctly are not in those files.
The changes an Elo test cannot judge
Four changes entered the engine on 3 September without playing a game. What each of them addressed is the kind of defect a strength test is not built to find.
Ten of seventeen entries played no games
Between 2 and 5 September the register gained seventeen entries. Seven of them played 174,984 games; the other ten played none. Four in all record a measured gain.
The verdict did not predict the cost
Six candidates cost 149,104 games and close to thirty hours of machine time. Whether one passed did not predict what it cost.
Two rejections that mean different things
Two search ideas were rejected on the same bound within two days. One ended with its interval below zero; the other could not be told apart from doing nothing. The verdict does not distinguish them — the interval does.
Three entries that played no games
Some register entries pass without a single game. They mark the points where the measuring conditions changed — and where comparisons across the boundary stop being valid.
Two runs that were meant to fail
An engine playing itself should measure nothing. Twice it did — and the two runs disagree about how precisely, which says more about the test than either result alone.
An Elo number needs its opponent
Six attempts in one day, two of them accepted. One of those carries two Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. numbers that differ by a factor of two hundred and fifty — and both are correct.
What the result file does not label
The first run in a tournament interface, three reading errors of the same shape, and five faults the harness reports never raised.