Two anchors, one point apart: Coherent 0.1.61 at 2570 ± 55

Project measurement, not in the register. Every figure is taken from the anchoring report of 10 September 2026 and the handoff of 9 September 2026. None of the figures is taken from the public data contract.

Project measurement, not in the register. Coherent 0.1.61 played 1,038 games against 2 foreign engines that carry a published rating, at the thinking time for which those ratings hold. The 2 anchors are 98 rating points apart. The 2 answers they give are one point apart.

Coherent 0.1.61 is at roughly 2570 on the scale of the public CCRL Computer Chess Rating Lists, a public ranking of chess engines from games under standard conditions. blitz list, with a band of ± 55.

Of the 2,192 engines rated there, 1,447 stand above it and 745 below.

Project measurement, not in the register. 4 horizontal intervals on one scale of anchored strength: 0.1.61 through Leorik 2.1 at 2570 with a half-width of 35, 0.1.61 through Chal 1.3.2 at 2571 with a half-width of 29, the published figure for 0.1.61 at 2570 with a band of 55, and the earlier figure for 0.1.58 at 2300 with a band of 60. 2 vertical lines mark the listed ratings of the 2 anchors, 2470 and 2568.
Project measurement, not in the register: no register ID, no assignable register epoch, and no oracle confirmation. The 2 anchorings of 0.1.61, the figure published from them, and the earlier figure for 0.1.58 for comparison. The vertical lines mark the listed ratings of the 2 anchors. The ± of the 2 anchorings are the half-widths of nominal ninety-5 per cent intervals from the games; the published band of ± 55 also carries the transfer error measured separately in the round robin.

The largest step this project has measured

The previous version, 0.1.58, was anchored at 2300 ± 60 — 2 foreign yardsticks that gave 2 different answers. The step is therefore roughly +270 points, by a wide margin the largest single step measured in this project. One change accounts for it: the hand-written position evaluation was replaced by NNUE Efficiently updatable neural network: a small network that evaluates positions and is updated incrementally move by move. inference, Generation 3.

The remarkable part is not the size of the step but that it arrived intact. Against its own predecessor the new evaluation had measured +247.0 ± 9.5 and +252.5 ± 13.1 internally, with a fixed number of games and no early stop, and a search change — IIR at non-PV Principal variation: the line of best moves the search currently expects. nodes — had measured +16.22 ± 8.31. Together that is roughly +265. Externally the full amount is still there. Internal gains usually shrink once a program meets foreign opponents. Here they did not.

The prediction was written down first

The sounding that opened this measurement was a test, not an interpretation after the fact, because 2 hypotheses and their expected score rates were written down before the games were in. The timestamp is 07:38 UTC on 8 September 2026, at which point 4 of 120 games of the sounding had finished.

The factor of 3.1 was not invented for the occasion. Between 2 earlier anchors, Blunder 7.1.0 and Drofa 2.0.0, we had measured 109.6 ± 30.1 points where the list says 32 — 5.05 σ out of 442 games. If that ratio were a property of our conditions rather than of those 2 engines, every figure derived from foreign ratings would be too high, the published 2300 ± 60 included.

The 2 hypotheses predict score rates that are far enough apart for 24 games per anchor to separate them:

AnchorH1 predictsH2 predictsobserved in the sounding
Chal 1.3.263 %39 %60.4 %
Leorik 2.150 %26 %52.1 %
Lynx 1.3.038 %18 %37.5 %
4ku 2.029 %12 %35.4 %
Inanis 1.1.024 %10 %22.9 %
Project measurement, not in the register. 5 rows, one per anchor engine, on a common axis of score rate. Each row carries the score rate predicted under H1, the score rate predicted under H2, and the score rate observed in the sounding of 24 games. In every row the observed rate lies at the H1 prediction and far from the H2 prediction.
Project measurement, not in the register: no register ID, no assignable register epoch, and no oracle confirmation. Score rates of Coherent 0.1.61 against each anchor, 24 games per anchor, against the 2 rates predicted in writing before the games were counted. The predictions are the ones recorded at 07:38 UTC on 8 September 2026.

The measuring agent's own expectation, written down in the same file, was a value between the two, closer to H2, for 3 reasons: that the factor came from a single pair of engines, one of which crashes; that part of the stretching comes from the draw rate and therefore hangs on book and thinking time; and that the internal +250 was measured on books that deliberately amplify differences, which would push the same way. All 3 reasons were wrong. H1 held across all 5 anchors, and the plain arithmetic that the measuring agent had itself called naive came out within 5 points of the result.

That is the point of writing predictions down. The figure is not credible because it is pleasant. It is credible because it was predicted in advance, against the expectation of the person measuring.

5 anchors, 24 games each

The sounding used 5 foreign engines spread over a ladder of 293 rating points, 24 games against each. From every score rate a value for Coherent follows, and those 5 values are derived independently of one another:

Anchorlistedscore rate of 0.1.61anchored value
Chal 1.3.2247060.4 %2543
Leorik 2.1256852.1 %2582
Lynx 1.3.0265337.5 %2563
4ku 2.0272335.4 %2619
Inanis 1.1.0276322.9 %2552

5 independently derived values scatter by 76 points over an anchor ladder of 293. At 24 games each these are coarse numbers, but they agree, and that agreement over such a range is the first evidence that the 2 scales fit together in this region.

2 long runs, one point apart

2 anchorings followed, long enough to matter:

Anchorlistedgamesscore rateElo differenceanchored
Leorik 2.12568 ± 1843850.23 %+1.59 ± 29.62570 ± 35
Chal 1.3.22470 ± 1760064.17 %+101.21 ± 24.02571 ± 29

The run against Leorik 2.1 is 219 colour-swapped pairs: 169 wins, 102 draws, 167 losses. The 2 anchors are 98 rating points apart and produce numbers that differ by 1 point, z = 0.04. Weighted together they give 2570.6 with a statistical half-width of ± 22. There were 0 losses on time in all 1,038 games.

It is worth saying what that agreement is worth, because the same project has seen the opposite. With 0.1.58 the 2 anchors disagreed by 54 points, and an arithmetically neat 2303 ± 23 had to become an honest 2300 ± 60. 2 anchors that agree are not a formality; they are the check that could have failed and did not.

The Leorik run ended at 438 games instead of 600 because the host of the measuring machine restarted and took the guest system with it. That does not bias the result. An early stop biases a measurement only when the decision to stop depends on the interim score; this one depended on the host, which knows nothing about chess. The rule is worth stating in general: it is not the abortion that is dangerous, it is the abortion that looks at the score first.

Why ± 55 and not ± 22

The ± 22 is the error from the games. It is not the whole uncertainty, and we found that out in the same series of measurements.

The 5 anchors played a round robin among themselves — 10 pairings, 60 games each, 600 games in total, without Coherent. It answers a question that every anchoring silently assumes: how well does a published rating transfer to somebody else's conditions?

Enginelistedmeasured by usresidual
Lynx 1.3.026532688+35
Chal 1.3.224702499+29
Leorik 2.125682584+16
Inanis 1.1.027632740−23
4ku 2.027232667−56

A single published rating can be dozens of points off under foreign conditions. The scatter of those deviations is about 37 points once the tournament noise of the round robin is taken out. Average 2 anchors and ± 51 of that remains at the 95 % level. Together with the ± 22 from the games, that is ± 55.

This is wider than the ± 35 we quoted for two days. It is not a worse number, it is a more honest one: we did not know beforehand how reliable a single foreign anchor is, and then we measured it.

Is playing strength a number at all?

The round robin also settles a question that comes before every rating: whether one number per engine can describe this field at all. Style differences can produce genuine non-transitivity, where A beats B, B beats C and C beats A, and then no ladder exists to hang anything on. The evaluation plan demanded that this be checked first, before any slope was computed, on the grounds that reporting a slope that does not exist would be the worse mistake.

A one-dimensional model fits: chi-square 9.42 at 6 degrees of freedom, 1.57 per degree of freedom. No genuine non-transitivity is found. That was not clear in advance.

Then the slope of our measured ratings against the listed ones. Over all 5 anchors it came out at 0.769 ± 0.071, 3.25 σ below 1 — apparently a compression by 23 %, and in the opposite direction to the stretching that H2 had suspected. The leave-one-out check overturns it:

left outslopechi-square per degree of freedom
none0.769 ± 0.0711.57
Chal 1.3.20.805 ± 0.1390.83
Lynx 1.3.00.708 ± 0.0801.50
4ku 2.00.936 ± 0.0890.37
Leorik 2.1model no longer fits2.07
Inanis 1.1.0model no longer fits2.93

Without 4ku 2.0 the slope is compatible with 1 and the fit becomes excellent. All 5 leave-one-out results are recorded here, not only the convenient one.

The apparent compression therefore hangs entirely on one engine. A real effect of our conditions could not do that: it would hit all pairings evenly and would sit in the slope, not in the residual of a single engine. That distinction was written into the plan before the games were played, which is why it can be used now.

Project measurement, not in the register. Scatter plot of 5 anchor engines: the rating listed on the CCRL blitz list on the horizontal axis, the rating we measured in the round robin on the vertical axis, with the listed half-widths as horizontal bars, a dashed identity line, and a line carrying the fitted slope of 0.769.
Project measurement, not in the register: no register ID, no assignable register epoch, and no oracle confirmation. Listed rating against the rating measured by us for the 5 anchors, from a round robin of 600 games without Coherent. Horizontal bars are the listed half-widths. The dashed line is identity; the solid line carries the fitted slope of 0.769 and is drawn through the centre of the 5 points. The point below the identity line on the right is 4ku 2.0, whose residual of −56 produces the whole apparent compression.

Two things follow that are worth keeping apart. The scale as a whole carries: a published rating is good information about the scale. The individual engine does not always: a published rating is uncertain information about that one program. And a piece of luck, verified after the fact — Leorik 2.1, the anchor Coherent's number hangs on, has the smallest residual of all 5 at +16.

Project measurement, not in the register. 6 rows on a common axis of the fitted slope, with a dashed line at 1: with no engine left out 0.769 ± 0.071, without Chal 1.3.2 0.805 ± 0.139, without Lynx 1.3.0 0.708 ± 0.080, without 4ku 2.0 0.936 ± 0.089, and without Leorik 2.1 or Inanis 1.1.0 no slope because the model no longer fits.
Project measurement, not in the register: no register ID, no assignable register epoch, and no oracle confirmation. The leave-one-out slopes of the scale fit. Leaving out 4ku 2.0 moves the slope onto 1; leaving out Leorik 2.1 or Inanis 1.1.0 breaks the one-dimensional model, and no slope is reported for those 2 rows.

How it was measured

Time control2 minutes plus 1 second per move — the thinking time of the ranking list used
Opening book8moves_v3, 34,700 positions, colour-swapped pairs
Hash64 MB per side
Concurrent games4
Adjudicationnone
Procedurefixed number of games, no early stop on an interim score, no SPRT Sequential probability ratio test. Games are played until the accumulated evidence reaches one of two bounds; then the test stops.
Tournament managerfastchess 1.8.2
Machine16 cores, Linux, dedicated measurement machine
Losses on time0 in 1,038 games

The measured build is commit b4ad0521, built with the switches of the project makefile. Its identity is proven rather than assumed: bench returns 3406 nodes, identical to what the independent test stands report at the same commit, on a different compiler and a different machine.

Both anchors are the official binaries of their authors, not builds of our own. That is not convenience. Among the candidates we rejected there was one whose only build path in the published tag was a debug build with an address sanitiser. It would have compiled cleanly, announced itself properly as an engine and passed as an anchor carrying its published rating — at a fraction of its playing strength. Coherent would have looked brilliant. Hence the rule: never build an anchor yourself when the author ships a binary.

All ratings used here were looked up by us in the list on 9 September 2026, not taken from somebody's summary.

What remains open

All of this comes from one opening book. The 5 agreeing values show that this book at 2+1 maps onto the conditions of the list. They do not show that the scale as such is book-independent. The orchestrator measured the same duel with a narrow and a wide book and got +10.3 against +28.6 Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. with an otherwise identical setup — a factor of 2.79. The reading we hold to is that the book acts on the resolution of small differences, not on the unit of the axis, and at distances of 100 to 300 points the real difference in strength dominates. That reading is plausible and it is not proven.

One finding is still unchecked. Between Blunder 7.1.0 and Drofa 2.0.0 we measured 109.6 ± 30.1 points where the list says 32, out of 442 games, 5.05 σ. Neither of those 2 engines plays in the round robin, so the round robin cannot decide whether that was the scale or whether it was Blunder. The experiment that would close the question is prepared and pre-registered: 2 bridges, over Chal 1.3.2 and over Leorik 2.1, 200 games per pairing, with 2 bridges rather than one so that a bridge engine that is itself off shows up instead of shifting everything invisibly. Its handling of aborted games is fixed in advance — crashes counted as losses, and crashes excluded, with both lines reported and a sentence saying which question each line answers. If both lines give the same residual, instability is not the explanation and the list is wrong about Blunder's strength itself. That run was not started, on the operator's instruction.

Engine ratings are not FIDE ratings. 2570 here does not mean grandmaster strength. These are different populations, and engines play against each other.

IMS Internally measured strength: 2200 plus the Elo difference against a frozen reference configuration of our own. A progress measure between our versions, not a rating from a ranking list. stays what it was. The internal scale measures against a frozen reference configuration of our own and remains a measure of progress between our own versions, nothing more.

When we measure again

A new external anchoring is worth doing once the accumulated internal gain exceeds the width of the measurement, so roughly 50 Elo. Single search constants worth +10 to +15 vanish inside the band; measuring them from outside would produce numbers without information. The next anchoring waits until there is a step large enough for a foreign yardstick to see it.

The engines Chal 1.3.2, Leorik 2.1, Lynx 1.3.0, 4ku 2.0, Inanis 1.1.0, Blunder 7.1.0 and Drofa 2.0.0 are named by word mark and version because the measurement conditions require it; no connection to their authors and no endorsement by them exists.