Playing strength

Coherent 0.1.61 transferred to 2570 ± 55 through 2 foreign engines that agree, then IMS Internally measured strength: 2200 plus the Elo difference against a frozen reference configuration of our own. A progress measure between our versions, not a rating from a ranking list. as a measure of progress between our own versions. Project measurement, not in the register.

Transferred playing strength

Coherent 0.1.61: approximately 2570 ± 55 — transferred over 2 opponents that are 98 rating points apart and give answers 1 point apart.

published rating + measured difference

Of the 2,192 engines rated on the CCRL Computer Chess Rating Lists, a public ranking of chess engines from games under standard conditions. blitz list, 1,447 stand above this figure and 745 below. The step from 0.1.58 is roughly +270 points, the largest single step measured in this project; it comes from replacing the hand-written position evaluation with NNUE Efficiently updatable neural network: a small network that evaluates positions and is updated incrementally move by move. inference.

The two anchorings of 0.1.61

OpponentPublished ratingGamesScore rateMeasured differenceAnchored Coherent
Leorik 2.12568 ± 1843850.23 %+1.59 ± 29.62570 ± 35
Chal 1.3.22470 ± 1760064.17 %+101.21 ± 24.02571 ± 29

All error figures are half-widths at the 95% level. The run against Leorik 2.1 is 219 colour-swapped pairs, 169 / 102 / 167; it ended at 438 of 600 games because the host of the measuring machine restarted. An early stop biases a measurement only when the decision to stop depends on the interim score; this one did not. There were 0 losses on time in all 1,038 games.

Weighted together the two anchorings give 2570.6 with a statistical half-width of ± 22; z between them is 0.04.

Project measurement, not in the register. Four horizontal intervals on one scale of anchored strength: 0.1.61 through Leorik 2.1 at 2570 with a half-width of 35, 0.1.61 through Chal 1.3.2 at 2571 with a half-width of 29, the published figure for 0.1.61 at 2570 with a band of 55, and the earlier figure for 0.1.58 at 2300 with a band of 60. Two vertical lines mark the listed ratings of the two anchors, 2470 and 2568.
Project measurement, not in the register. The two anchorings of 0.1.61, the figure published from them, and the previous measurement of 0.1.58 for comparison. Vertical lines: the listed ratings of the two anchors. The ± of the anchorings are half-widths of nominal 95 % intervals from the games; the published ± 55 also carries the transfer error measured in the round robin.

Why the published band is ± 55 and not ± 22

The 5 anchors played a round robin among themselves, 10 pairings of 60 games, 600 games in total, without Coherent. It shows how well a published rating transfers to our conditions:

EnginePublished ratingMeasured by usResidual
Lynx 1.3.026532688+35
Chal 1.3.224702499+29
Leorik 2.125682584+16
Inanis 1.1.027632740−23
4ku 2.027232667−56

The scatter of these deviations is about 37 points once the tournament noise is taken out; averaged over 2 anchors, ± 51 of it remains at the 95 % level. Together with the ± 22 from the games that is ± 55. Leorik 2.1, the anchor the figure hangs on, has the smallest residual of the five.

A one-dimensional model fits the round robin: chi-square 9.42 at 6 degrees of freedom, 1.57 per degree of freedom, no genuine non-transitivity. The slope of our measured ratings against the listed ones is 0.769 ± 0.071 over all five, and 0.936 ± 0.089 without 4ku 2.0, where the fit becomes excellent at 0.37 per degree of freedom. The apparent compression therefore sits in one engine, not in the scale: the two scales map 1:1, and 2570 stands without conversion.

Measurement conditions of the anchorings of 0.1.61

PropertyValue
Time control120+1
Opening book8moves_v3 · 34,700 positions, colour-swapped pairs
Hash (MB per side)64
Concurrent games4
Adjudicationnone
Procedurefixed number of games, no early stop, no SPRT Sequential probability ratio test. Games are played until the accumulated evidence reaches one of two bounds; then the test stops.
Tournament managerfastchess 1.8.2
Machine16 cores, Linux, dedicated measurement machine
Measured buildcommit b4ad0521, bench 3406 nodes
Losses on time0 in 1,038 games

Both anchors are the official binaries of their authors. All published ratings were looked up by us in the list on 2026-09-09.

Anchoring report of 2026-09-10 · Measurements, prediction, scale and limitations

This is a transfer, not a place in a ranking. Engine ratings are not FIDE ratings: 2570 here does not mean grandmaster strength.

The engines Chal 1.3.2, Leorik 2.1, Lynx 1.3.0, 4ku 2.0, Inanis 1.1.0, Blunder 7.1.0 and Drofa 2.0.0 are named by word mark and version because the measurement conditions require it; no connection to their authors and no endorsement by them exists.

Previous measurement: Coherent 0.1.58

Coherent 0.1.58: approximately 2300 ± 60 — transferred over 2 opponents that disagree with each other.

published rating + measured difference

The two measurements

OpponentPublished ratingGamesScore rateMeasured differenceAnchored Coherent
Blunder 7.1.02388 ± 1854841.88 %−56.93 ± 28.022331 ± 33
Drofa 2.0.02420 ± 1960030.50 %−143.07 ± 26.282277 ± 32

All error figures are half-widths at the 95% level.

The two answers lie 54 points apart: 2.28 standard deviations; the two-sided chance probability is 0.022.

This is a transfer, not a place in a ranking. A shared error in transferring the ratings could shift both anchors in the same direction and is not covered by the published band.

Blunder 7.1.0 against Drofa 2.0.0 directly in our own setup: completed.

The two anchors play each other directly in our setup

PropertyValue
Games442
Pairs221
Pentanomial (Blunder)68 / 49 / 65 / 28 / 11
Score rate (Blunder)34.73 %
Measured difference109.6 ± 30.1
Difference from the ranking list77.6

Explanation 1 confirmed: published ratings do not transfer to our conditions. The measured gap is more than three times the one in the ranking list.

Measured in two parts (200 and 242 games, different book positions); part 2 was ended early on the operator's instruction, the games played remain valid. Merged at pair level, not across the two Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. estimators.

Blunder 7.1.0: The run against Blunder ended at game 549 because the opponent crashed and the tournament manager stopped there. The aborted game is not counted.

Measurement conditions of the anchorings of 0.1.58

PropertyValue
Time control120+1
Opening book8moves_v3 · 34,700 positions
Hash (MB per side)64
Concurrent games4
Adjudicationnone
Procedurefixed number of games, no early stop

Data delivery created on 2026-09-07 · Measurements, interpretation and limitations

IMS says how the engine fares against a frozen reference configuration of its own — and nothing else. Anyone who wants to know where the program stands in the field needs opponents that somebody else built and that somebody else rated.

IMS lies above both anchors of 0.1.58. As a measure of progress between versions of our own it remains usable; as a statement about strength in the field it was too optimistic.

There is no conversion between IMS and Elo, no shared axis and no shared chart with two scales.

IMS — internal progress

IMS = 2200 + Elo difference against R1

IMS is a difference against a frozen reference configuration, not a rating from a ranking list.

The zero point is calibrated to the setting of this configuration, not to a public pool; comparisons with public ranking lists are therefore not permitted.

The reference configuration plays throttled and non-deterministically, so a repetition can easily deviate.

Measurement plot

0.1.58 · 2026-09-06 · IMS 2363 ±26 · nominal 95%
Project measurement, not in the register. Error bands show measurement error relative to the reference, not the unknown scale error. Point: IMS of the measured build; Error band: ± half-width, nominal 95%; fixed game count, no SPRT.

Today this is one measurement with its error band, not a trend; the curve appears only as further builds are measured.

Measurements

Measurements and conditions — project measurement, not in the register

BuildDateIMSHalf-widthGamesW/D/LPentanomialReference
0.1.582026-09-062363±26 (nominal 95%)800 (fixed game count, no SPRT)558 / 34 / 20828 / 7 / 145 / 27 / 193R1

W/D/L: wins / draws / losses; pentanomial Games are played in pairs with swapped colours; the five counts are the pairs scoring 0, ½, 1, 1½ and 2 points.: game pairs scoring 0 / ½ / 1 / 1½ / 2 points; conditions per reference: see Reference configuration.

Reference configuration

Frozen reference R1

PropertyValue
EngineStockfish 19
Buildsf_19, Linux x86-64 universal
OptionsUCI_LimitStrength=true UCI_Elo=2200
IMS2200
Time control10+0.1
Book8moves_v3 · 34,700 positions
Hash (MB)64
Threads1
Adjudicationnone
Frozen2026-09-05

The reference configuration plays with limited strength and is not deterministic; a repeat measurement may differ slightly.

The authors of the engine used as the reference did not know of this measurement, took no part in it, and did not consent to it. The program was used solely as freely available software under its licence.

Methodology

How it is made

Updated: