Two anchors, one point apart: Coherent 0.1.61 at 2570 ± 55
Project measurement, not in the register. Every figure is taken from the anchoring report of 10 September 2026 and the handoff of 9 September 2026. None of the figures is taken from the public data contract.
Project measurement, not in the register. Coherent 0.1.61 played 1,038 games against 2 foreign engines that carry a published rating, at the thinking time for which those ratings hold. The 2 anchors are 98 rating points apart. The 2 answers they give are one point apart.
Coherent
0.1.61is at roughly 2570 on the scale of the public CCRL Computer Chess Rating Lists, a public ranking of chess engines from games under standard conditions. blitz list, with a band of ± 55.
Of the 2,192 engines rated there, 1,447 stand above it and 745 below.
0.1.61, the figure published from them, and the earlier figure for 0.1.58 for comparison. The vertical lines mark the listed ratings of the 2 anchors. The ± of the 2 anchorings are the half-widths of nominal ninety-5 per cent intervals from the games; the published band of ± 55 also carries the transfer error measured separately in the round robin.The largest step this project has measured
The previous version, 0.1.58, was anchored at 2300 ± 60 — 2 foreign yardsticks that gave 2 different answers. The step is therefore roughly +270 points, by a wide margin the largest single step measured in this project. One change accounts for it: the hand-written position evaluation was replaced by NNUE Efficiently updatable neural network: a small network that evaluates positions and is updated incrementally move by move. inference, Generation 3.
The remarkable part is not the size of the step but that it arrived intact. Against its own predecessor the new evaluation had measured +247.0 ± 9.5 and +252.5 ± 13.1 internally, with a fixed number of games and no early stop, and a search change — IIR at non-PV Principal variation: the line of best moves the search currently expects. nodes — had measured +16.22 ± 8.31. Together that is roughly +265. Externally the full amount is still there. Internal gains usually shrink once a program meets foreign opponents. Here they did not.
The prediction was written down first
The sounding that opened this measurement was a test, not an interpretation after the fact, because 2 hypotheses and their expected score rates were written down before the games were in. The timestamp is 07:38 UTC on 8 September 2026, at which point 4 of 120 games of the sounding had finished.
- H1 — the 2 scales map onto each other 1:1. Then 2300 + 16 + 250 is about 2566.
- H2 — our own scale is stretched by a factor of 3.1. Then 2300 + 5 + 81 is about 2386.
The factor of 3.1 was not invented for the occasion. Between 2 earlier anchors, Blunder 7.1.0 and Drofa 2.0.0, we had measured 109.6 ± 30.1 points where the list says 32 — 5.05 σ out of 442 games. If that ratio were a property of our conditions rather than of those 2 engines, every figure derived from foreign ratings would be too high, the published 2300 ± 60 included.
The 2 hypotheses predict score rates that are far enough apart for 24 games per anchor to separate them:
| Anchor | H1 predicts | H2 predicts | observed in the sounding |
|---|---|---|---|
| Chal 1.3.2 | 63 % | 39 % | 60.4 % |
| Leorik 2.1 | 50 % | 26 % | 52.1 % |
| Lynx 1.3.0 | 38 % | 18 % | 37.5 % |
| 4ku 2.0 | 29 % | 12 % | 35.4 % |
| Inanis 1.1.0 | 24 % | 10 % | 22.9 % |
0.1.61 against each anchor, 24 games per anchor, against the 2 rates predicted in writing before the games were counted. The predictions are the ones recorded at 07:38 UTC on 8 September 2026.The measuring agent's own expectation, written down in the same file, was a value between the two, closer to H2, for 3 reasons: that the factor came from a single pair of engines, one of which crashes; that part of the stretching comes from the draw rate and therefore hangs on book and thinking time; and that the internal +250 was measured on books that deliberately amplify differences, which would push the same way. All 3 reasons were wrong. H1 held across all 5 anchors, and the plain arithmetic that the measuring agent had itself called naive came out within 5 points of the result.
That is the point of writing predictions down. The figure is not credible because it is pleasant. It is credible because it was predicted in advance, against the expectation of the person measuring.
5 anchors, 24 games each
The sounding used 5 foreign engines spread over a ladder of 293 rating points, 24 games against each. From every score rate a value for Coherent follows, and those 5 values are derived independently of one another:
| Anchor | listed | score rate of 0.1.61 | anchored value |
|---|---|---|---|
| Chal 1.3.2 | 2470 | 60.4 % | 2543 |
| Leorik 2.1 | 2568 | 52.1 % | 2582 |
| Lynx 1.3.0 | 2653 | 37.5 % | 2563 |
| 4ku 2.0 | 2723 | 35.4 % | 2619 |
| Inanis 1.1.0 | 2763 | 22.9 % | 2552 |
5 independently derived values scatter by 76 points over an anchor ladder of 293. At 24 games each these are coarse numbers, but they agree, and that agreement over such a range is the first evidence that the 2 scales fit together in this region.
2 long runs, one point apart
2 anchorings followed, long enough to matter:
| Anchor | listed | games | score rate | Elo difference | anchored |
|---|---|---|---|---|---|
| Leorik 2.1 | 2568 ± 18 | 438 | 50.23 % | +1.59 ± 29.6 | 2570 ± 35 |
| Chal 1.3.2 | 2470 ± 17 | 600 | 64.17 % | +101.21 ± 24.0 | 2571 ± 29 |
The run against Leorik 2.1 is 219 colour-swapped pairs: 169 wins, 102 draws, 167 losses. The 2 anchors are 98 rating points apart and produce numbers that differ by 1 point, z = 0.04. Weighted together they give 2570.6 with a statistical half-width of ± 22. There were 0 losses on time in all 1,038 games.
It is worth saying what that agreement is worth, because the same project has seen the opposite. With 0.1.58 the 2 anchors disagreed by 54 points, and an arithmetically neat 2303 ± 23 had to become an honest 2300 ± 60. 2 anchors that agree are not a formality; they are the check that could have failed and did not.
The Leorik run ended at 438 games instead of 600 because the host of the measuring machine restarted and took the guest system with it. That does not bias the result. An early stop biases a measurement only when the decision to stop depends on the interim score; this one depended on the host, which knows nothing about chess. The rule is worth stating in general: it is not the abortion that is dangerous, it is the abortion that looks at the score first.
Why ± 55 and not ± 22
The ± 22 is the error from the games. It is not the whole uncertainty, and we found that out in the same series of measurements.
The 5 anchors played a round robin among themselves — 10 pairings, 60 games each, 600 games in total, without Coherent. It answers a question that every anchoring silently assumes: how well does a published rating transfer to somebody else's conditions?
| Engine | listed | measured by us | residual |
|---|---|---|---|
| Lynx 1.3.0 | 2653 | 2688 | +35 |
| Chal 1.3.2 | 2470 | 2499 | +29 |
| Leorik 2.1 | 2568 | 2584 | +16 |
| Inanis 1.1.0 | 2763 | 2740 | −23 |
| 4ku 2.0 | 2723 | 2667 | −56 |
A single published rating can be dozens of points off under foreign conditions. The scatter of those deviations is about 37 points once the tournament noise of the round robin is taken out. Average 2 anchors and ± 51 of that remains at the 95 % level. Together with the ± 22 from the games, that is ± 55.
This is wider than the ± 35 we quoted for two days. It is not a worse number, it is a more honest one: we did not know beforehand how reliable a single foreign anchor is, and then we measured it.
Is playing strength a number at all?
The round robin also settles a question that comes before every rating: whether one number per engine can describe this field at all. Style differences can produce genuine non-transitivity, where A beats B, B beats C and C beats A, and then no ladder exists to hang anything on. The evaluation plan demanded that this be checked first, before any slope was computed, on the grounds that reporting a slope that does not exist would be the worse mistake.
A one-dimensional model fits: chi-square 9.42 at 6 degrees of freedom, 1.57 per degree of freedom. No genuine non-transitivity is found. That was not clear in advance.
Then the slope of our measured ratings against the listed ones. Over all 5 anchors it came out at 0.769 ± 0.071, 3.25 σ below 1 — apparently a compression by 23 %, and in the opposite direction to the stretching that H2 had suspected. The leave-one-out check overturns it:
| left out | slope | chi-square per degree of freedom |
|---|---|---|
| none | 0.769 ± 0.071 | 1.57 |
| Chal 1.3.2 | 0.805 ± 0.139 | 0.83 |
| Lynx 1.3.0 | 0.708 ± 0.080 | 1.50 |
| 4ku 2.0 | 0.936 ± 0.089 | 0.37 |
| Leorik 2.1 | model no longer fits | 2.07 |
| Inanis 1.1.0 | model no longer fits | 2.93 |
Without 4ku 2.0 the slope is compatible with 1 and the fit becomes excellent. All 5 leave-one-out results are recorded here, not only the convenient one.
The apparent compression therefore hangs entirely on one engine. A real effect of our conditions could not do that: it would hit all pairings evenly and would sit in the slope, not in the residual of a single engine. That distinction was written into the plan before the games were played, which is why it can be used now.
2.0, whose residual of −56 produces the whole apparent compression.Two things follow that are worth keeping apart. The scale as a whole carries: a published rating is good information about the scale. The individual engine does not always: a published rating is uncertain information about that one program. And a piece of luck, verified after the fact — Leorik 2.1, the anchor Coherent's number hangs on, has the smallest residual of all 5 at +16.
2.0 moves the slope onto 1; leaving out Leorik 2.1 or Inanis 1.1.0 breaks the one-dimensional model, and no slope is reported for those 2 rows.How it was measured
| Time control | 2 minutes plus 1 second per move — the thinking time of the ranking list used |
|---|---|
| Opening book | 8moves_v3, 34,700 positions, colour-swapped pairs |
| Hash | 64 MB per side |
| Concurrent games | 4 |
| Adjudication | none |
| Procedure | fixed number of games, no early stop on an interim score, no SPRT Sequential probability ratio test. Games are played until the accumulated evidence reaches one of two bounds; then the test stops. |
| Tournament manager | fastchess 1.8.2 |
| Machine | 16 cores, Linux, dedicated measurement machine |
| Losses on time | 0 in 1,038 games |
The measured build is commit b4ad0521, built with the switches of the project makefile. Its identity is proven rather than assumed: bench returns 3406 nodes, identical to what the independent test stands report at the same commit, on a different compiler and a different machine.
Both anchors are the official binaries of their authors, not builds of our own. That is not convenience. Among the candidates we rejected there was one whose only build path in the published tag was a debug build with an address sanitiser. It would have compiled cleanly, announced itself properly as an engine and passed as an anchor carrying its published rating — at a fraction of its playing strength. Coherent would have looked brilliant. Hence the rule: never build an anchor yourself when the author ships a binary.
All ratings used here were looked up by us in the list on 9 September 2026, not taken from somebody's summary.
What remains open
All of this comes from one opening book. The 5 agreeing values show that this book at 2+1 maps onto the conditions of the list. They do not show that the scale as such is book-independent. The orchestrator measured the same duel with a narrow and a wide book and got +10.3 against +28.6 Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. with an otherwise identical setup — a factor of 2.79. The reading we hold to is that the book acts on the resolution of small differences, not on the unit of the axis, and at distances of 100 to 300 points the real difference in strength dominates. That reading is plausible and it is not proven.
One finding is still unchecked. Between Blunder 7.1.0 and Drofa 2.0.0 we measured 109.6 ± 30.1 points where the list says 32, out of 442 games, 5.05 σ. Neither of those 2 engines plays in the round robin, so the round robin cannot decide whether that was the scale or whether it was Blunder. The experiment that would close the question is prepared and pre-registered: 2 bridges, over Chal 1.3.2 and over Leorik 2.1, 200 games per pairing, with 2 bridges rather than one so that a bridge engine that is itself off shows up instead of shifting everything invisibly. Its handling of aborted games is fixed in advance — crashes counted as losses, and crashes excluded, with both lines reported and a sentence saying which question each line answers. If both lines give the same residual, instability is not the explanation and the list is wrong about Blunder's strength itself. That run was not started, on the operator's instruction.
Engine ratings are not FIDE ratings. 2570 here does not mean grandmaster strength. These are different populations, and engines play against each other.
IMS Internally measured strength: 2200 plus the Elo difference against a frozen reference configuration of our own. A progress measure between our versions, not a rating from a ranking list. stays what it was. The internal scale measures against a frozen reference configuration of our own and remains a measure of progress between our own versions, nothing more.
When we measure again
A new external anchoring is worth doing once the accumulated internal gain exceeds the width of the measurement, so roughly 50 Elo. Single search constants worth +10 to +15 vanish inside the band; measuring them from outside would produce numbers without information. The next anchoring waits until there is a step large enough for a foreign yardstick to see it.
The engines Chal 1.3.2, Leorik 2.1, Lynx 1.3.0, 4ku 2.0, Inanis 1.1.0, Blunder 7.1.0 and Drofa 2.0.0 are named by word mark and version because the measurement conditions require it; no connection to their authors and no endorsement by them exists.