Two foreign yardsticks, two different answers

Project measurement, not in the register. Every figure is taken from the Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. anchoring data delivery of 7 September 2026. None of the figures is taken from the public data contract.

Project measurement, not in the register. Coherent 0.1.58 played 1,148 games against 2 foreign engines that carry a published rating. The two anchors disagree by 54 points. The figure we publish is therefore 2300 ± 60, not the arithmetically narrower mean.

Project measurement, not in the register. Until now Coherent has only been measured against itself. The number for that is IMS Internally measured strength: 2200 plus the Elo difference against a frozen reference configuration of our own. A progress measure between our versions, not a rating from a ranking list., internally measured strength. It says how the engine fares against a frozen reference configuration of its own — and nothing else. Anyone who wants to know where the program stands in the field needs opponents that somebody else built and that somebody else rated.

That has now happened: 1,148 games by Coherent 0.1.58 against 2 foreign engines with a published rating. What came out is not a clean number but a contradiction — and the contradiction is the interesting part.

The two measurements

The same setup in both: a time control of 2 minutes plus 1 second per move — the thinking time for which the ranking list used publishes its values. Colour-swapped pairs from a fixed opening book, a fixed number of games, no early stop on an interim score, no adjudication of positions by an arbiter, 64 MB of hash per side.

Opponentpublished ratinggamesscore ratemeasured differencefrom this: Coherent
Blunder 7.1.02388 ± 1854841.88 %−56.93 ± 28.022331 ± 33
Drofa 2.0.02420 ± 1960030.50 %−143.07 ± 26.282277 ± 32

All error figures are half-widths at the 95 % level.

Project measurement, not in the register. Three horizontal intervals on one scale of anchored strength: through Blunder 2331 with a half-width of 33, through Drofa 2277 with a half-width of 32, and the published figure 2300 with a band of 60. Two vertical lines mark the published ratings of the opponents, 2388 and 2420. The ± figures are the half-widths of nominal ninety-five per cent intervals. The two anchorings are not independent estimates of the same value.
Project measurement, not in the register: no register ID, no assignable register epoch, and no oracle confirmation. Two anchorings of the same engine through 2 different opponents, together with the figure published from them. The vertical lines mark the published ratings of the two opponents. Error bars are the half-widths of nominal ninety-five per cent intervals. The two anchorings are not independent estimates of the same value: a shared transfer error would move both in the same direction and would not cancel when they are averaged. 1,148 games at 2 minutes plus 1 second per move; ± figures are half-widths of nominal ninety-five per cent intervals; Vertical lines: the published ratings of the two opponents. The two anchorings disagree; a shared transfer error would not cancel when they are averaged.

The two answers lie 54 points apart. That is 2.28 standard deviations; the two-sided probability of seeing it by chance is 0.022. Too rare to wave away.

What we say — and what we do not say

The honest statement is this:

Coherent 0.1.58 is at roughly 2300, with a band of ± 60 — carried over from 2 opponents that contradict each other.

Arithmetically the two measurements would give a mean of 2303 with ± 23. We do not write that number down. It only arises if one assumes the two anchors agree — and that is precisely what they do not. The range from 2277 to 2331 would be too narrow as well: those are 2 point values that each carry about ± 32 of their own, not the edges of what is plausible.

And even ± 60 does not cover everything. Whatever might be equally wrong in both runs — our hardware, our book, our conditions against those of the ranking list — moves both numbers in the same direction and does not cancel when they are averaged. An error that none of our measurements can see remains possible.

A word on the form: this reads like a place in a ranking, but it is not one. It is a transfer over 2 opponents, not a tournament in a grown field.

What it could be

Here it gets uncomfortable. The contradiction is clearest when the two opponents are held against each other rather than against us: the ranking list puts Drofa 32 points above Blunder. By the detour through our games we measure 86 points between those same two engines.

2 explanations fit the same numbers equally well:

  1. The foreign ratings do not transfer. A rating holds for the conditions under which it arose — different hardware, a different field of opponents, different openings. That it does not transfer point for point would be the normal case, not a mishap.
  2. Coherent suits one opponent better than the other. Playing strength is not really a number on a straight line. Between 2 programs there can be matters of style that no ranking list captures, without either rating being wrong.

We would have preferred the first explanation — it shows us in a better light. But it does not follow from the data. On the contrary: of the spread with which we judge the contradiction at all, roughly 2 thirds comes from our own games and only 1 third from the published ratings. Anyone looking to assign blame here has to start at home.

The experiment that settles it

The question can be answered, and without us: let the two anchors play each other directly, in our setup, with our book, at the same thinking time.

That run is queued and starts as soon as the machine is free. We will report the result — including the case where it is inconvenient.

A mishap we are not playing down

The run against Blunder ended at game 549 instead of 600 because the opponent crashed and the tournament manager stopped there. The aborted game is not counted.

That is more than a footnote, because it hits precisely the run that comes out more favourably for us. A crash in a lost position that is not counted pulls our score rate up — that is, exactly towards the higher anchor of 2331. We had 0 losses on time across all 1,148 games, but that says nothing about how crashes are handled. Anyone giving this run less weight lands closer to 2280 than to 2300.

We mention it because an aborted series of measurements is more dangerous than a missing one: it looks complete.

And what does that mean for IMS?

The internal number stands at 2363 ± 26 and therefore lies above both anchors. The reason was predicted: IMS hangs on an artificially throttled reference configuration whose setting we adopted as the zero point. An earlier bridge measurement had shown that this throttling is not linear — between 2 settings that should have been 300 points apart, we measured 133.

The suspicion is confirmed: the IMS scale sits too high. By how much exactly, 2 anchors that disagree cannot say. As a measure of progress between 2 versions of our own it remains usable — that is what it was built for, and for that the zero point does not matter. As a statement about strength in the field it was too optimistic.

Measurement conditions

Time control2 minutes plus 1 second per move
Opening book8moves_v3, 34,700 positions, colour-swapped pairs
Hash64 MB per side
Concurrent games4
Adjudicationnone
Procedurefixed number of games, no early stop
Machine16 cores, dedicated measurement machine
Losses on time0
Coherent crashes0

The engines Blunder 7.1.0 and Drofa 2.0.0 are named by word mark and version because the measurement conditions require it; no connection to their authors and no endorsement by them exists.

The statistics in this entry were recomputed by an independent model. An earlier version gave the range between the two anchors as the result and blamed the contradiction on the foreign ratings. Neither of those stands here any more, and the number of games has been corrected since.