Two foreign yardsticks, two different answers
Project measurement, not in the register. Every figure is taken from the Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. anchoring data delivery of 7 September 2026. None of the figures is taken from the public data contract.
Project measurement, not in the register. Coherent 0.1.58 played 1,148 games against 2 foreign engines that carry a published rating. The two anchors disagree by 54 points. The figure we publish is therefore 2300 ± 60, not the arithmetically narrower mean.
Project measurement, not in the register. Until now Coherent has only been measured against itself. The number for that is IMS Internally measured strength: 2200 plus the Elo difference against a frozen reference configuration of our own. A progress measure between our versions, not a rating from a ranking list., internally measured strength. It says how the engine fares against a frozen reference configuration of its own — and nothing else. Anyone who wants to know where the program stands in the field needs opponents that somebody else built and that somebody else rated.
That has now happened: 1,148 games by Coherent 0.1.58 against 2 foreign engines with a published rating. What came out is not a clean number but a contradiction — and the contradiction is the interesting part.
The two measurements
The same setup in both: a time control of 2 minutes plus 1 second per move — the thinking time for which the ranking list used publishes its values. Colour-swapped pairs from a fixed opening book, a fixed number of games, no early stop on an interim score, no adjudication of positions by an arbiter, 64 MB of hash per side.
| Opponent | published rating | games | score rate | measured difference | from this: Coherent |
|---|---|---|---|---|---|
Blunder 7.1.0 | 2388 ± 18 | 548 | 41.88 % | −56.93 ± 28.02 | 2331 ± 33 |
Drofa 2.0.0 | 2420 ± 19 | 600 | 30.50 % | −143.07 ± 26.28 | 2277 ± 32 |
All error figures are half-widths at the 95 % level.
The two answers lie 54 points apart. That is 2.28 standard deviations; the two-sided probability of seeing it by chance is 0.022. Too rare to wave away.
What we say — and what we do not say
The honest statement is this:
Coherent
0.1.58is at roughly 2300, with a band of ± 60 — carried over from 2 opponents that contradict each other.
Arithmetically the two measurements would give a mean of 2303 with ± 23. We do not write that number down. It only arises if one assumes the two anchors agree — and that is precisely what they do not. The range from 2277 to 2331 would be too narrow as well: those are 2 point values that each carry about ± 32 of their own, not the edges of what is plausible.
And even ± 60 does not cover everything. Whatever might be equally wrong in both runs — our hardware, our book, our conditions against those of the ranking list — moves both numbers in the same direction and does not cancel when they are averaged. An error that none of our measurements can see remains possible.
A word on the form: this reads like a place in a ranking, but it is not one. It is a transfer over 2 opponents, not a tournament in a grown field.
What it could be
Here it gets uncomfortable. The contradiction is clearest when the two opponents are held against each other rather than against us: the ranking list puts Drofa 32 points above Blunder. By the detour through our games we measure 86 points between those same two engines.
2 explanations fit the same numbers equally well:
- The foreign ratings do not transfer. A rating holds for the conditions under which it arose — different hardware, a different field of opponents, different openings. That it does not transfer point for point would be the normal case, not a mishap.
- Coherent suits one opponent better than the other. Playing strength is not really a number on a straight line. Between 2 programs there can be matters of style that no ranking list captures, without either rating being wrong.
We would have preferred the first explanation — it shows us in a better light. But it does not follow from the data. On the contrary: of the spread with which we judge the contradiction at all, roughly 2 thirds comes from our own games and only 1 third from the published ratings. Anyone looking to assign blame here has to start at home.
The experiment that settles it
The question can be answered, and without us: let the two anchors play each other directly, in our setup, with our book, at the same thinking time.
- If a difference around 32 comes out, the ranking list carries over to us, and the rest is a matter of style between Coherent and the two.
- If a difference around 86 comes out, our setup contradicts the list.
That run is queued and starts as soon as the machine is free. We will report the result — including the case where it is inconvenient.
A mishap we are not playing down
The run against Blunder ended at game 549 instead of 600 because the opponent crashed and the tournament manager stopped there. The aborted game is not counted.
That is more than a footnote, because it hits precisely the run that comes out more favourably for us. A crash in a lost position that is not counted pulls our score rate up — that is, exactly towards the higher anchor of 2331. We had 0 losses on time across all 1,148 games, but that says nothing about how crashes are handled. Anyone giving this run less weight lands closer to 2280 than to 2300.
We mention it because an aborted series of measurements is more dangerous than a missing one: it looks complete.
And what does that mean for IMS?
The internal number stands at 2363 ± 26 and therefore lies above both anchors. The reason was predicted: IMS hangs on an artificially throttled reference configuration whose setting we adopted as the zero point. An earlier bridge measurement had shown that this throttling is not linear — between 2 settings that should have been 300 points apart, we measured 133.
The suspicion is confirmed: the IMS scale sits too high. By how much exactly, 2 anchors that disagree cannot say. As a measure of progress between 2 versions of our own it remains usable — that is what it was built for, and for that the zero point does not matter. As a statement about strength in the field it was too optimistic.
Measurement conditions
| Time control | 2 minutes plus 1 second per move |
|---|---|
| Opening book | 8moves_v3, 34,700 positions, colour-swapped pairs |
| Hash | 64 MB per side |
| Concurrent games | 4 |
| Adjudication | none |
| Procedure | fixed number of games, no early stop |
| Machine | 16 cores, dedicated measurement machine |
| Losses on time | 0 |
| Coherent crashes | 0 |
The engines Blunder 7.1.0 and Drofa 2.0.0 are named by word mark and version because the measurement conditions require it; no connection to their authors and no endorsement by them exists.
The statistics in this entry were recomputed by an independent model. An earlier version gave the range between the two anchors as the result and blamed the contradiction on the foreign ratings. Neither of those stands here any more, and the number of games has been corrected since.