Playing strength
Coherent 0.1.61 transferred to 2570 ± 55 through 2 foreign engines that agree, then IMS Internally measured strength: 2200 plus the Elo difference against a frozen reference configuration of our own. A progress measure between our versions, not a rating from a ranking list. as a measure of progress between our own versions. Project measurement, not in the register.
Transferred playing strength
Coherent 0.1.61: approximately 2570 ± 55 — transferred over 2 opponents that are 98 rating points apart and give answers 1 point apart.
published rating + measured difference
Of the 2,192 engines rated on the CCRL Computer Chess Rating Lists, a public ranking of chess engines from games under standard conditions. blitz list, 1,447 stand above this figure and 745 below. The step from 0.1.58 is roughly +270 points, the largest single step measured in this project; it comes from replacing the hand-written position evaluation with NNUE Efficiently updatable neural network: a small network that evaluates positions and is updated incrementally move by move. inference.
The two anchorings of 0.1.61
| Opponent | Published rating | Games | Score rate | Measured difference | Anchored Coherent |
|---|---|---|---|---|---|
| Leorik 2.1 | 2568 ± 18 | 438 | 50.23 % | +1.59 ± 29.6 | 2570 ± 35 |
| Chal 1.3.2 | 2470 ± 17 | 600 | 64.17 % | +101.21 ± 24.0 | 2571 ± 29 |
All error figures are half-widths at the 95% level. The run against Leorik 2.1 is 219 colour-swapped pairs, 169 / 102 / 167; it ended at 438 of 600 games because the host of the measuring machine restarted. An early stop biases a measurement only when the decision to stop depends on the interim score; this one did not. There were 0 losses on time in all 1,038 games.
Weighted together the two anchorings give 2570.6 with a statistical half-width of ± 22; z between them is 0.04.
Why the published band is ± 55 and not ± 22
The 5 anchors played a round robin among themselves, 10 pairings of 60 games, 600 games in total, without Coherent. It shows how well a published rating transfers to our conditions:
| Engine | Published rating | Measured by us | Residual |
|---|---|---|---|
| Lynx 1.3.0 | 2653 | 2688 | +35 |
| Chal 1.3.2 | 2470 | 2499 | +29 |
| Leorik 2.1 | 2568 | 2584 | +16 |
| Inanis 1.1.0 | 2763 | 2740 | −23 |
| 4ku 2.0 | 2723 | 2667 | −56 |
The scatter of these deviations is about 37 points once the tournament noise is taken out; averaged over 2 anchors, ± 51 of it remains at the 95 % level. Together with the ± 22 from the games that is ± 55. Leorik 2.1, the anchor the figure hangs on, has the smallest residual of the five.
A one-dimensional model fits the round robin: chi-square 9.42 at 6 degrees of freedom, 1.57 per degree of freedom, no genuine non-transitivity. The slope of our measured ratings against the listed ones is 0.769 ± 0.071 over all five, and 0.936 ± 0.089 without 4ku 2.0, where the fit becomes excellent at 0.37 per degree of freedom. The apparent compression therefore sits in one engine, not in the scale: the two scales map 1:1, and 2570 stands without conversion.
Measurement conditions of the anchorings of 0.1.61
| Property | Value |
|---|---|
| Time control | 120+1 |
| Opening book | 8moves_v3 · 34,700 positions, colour-swapped pairs |
| Hash (MB per side) | 64 |
| Concurrent games | 4 |
| Adjudication | none |
| Procedure | fixed number of games, no early stop, no SPRT Sequential probability ratio test. Games are played until the accumulated evidence reaches one of two bounds; then the test stops. |
| Tournament manager | fastchess 1.8.2 |
| Machine | 16 cores, Linux, dedicated measurement machine |
| Measured build | commit b4ad0521, bench 3406 nodes |
| Losses on time | 0 in 1,038 games |
Both anchors are the official binaries of their authors. All published ratings were looked up by us in the list on 2026-09-09.
Anchoring report of 2026-09-10 · Measurements, prediction, scale and limitations
This is a transfer, not a place in a ranking. Engine ratings are not FIDE ratings: 2570 here does not mean grandmaster strength.
The engines Chal 1.3.2, Leorik 2.1, Lynx 1.3.0, 4ku 2.0, Inanis 1.1.0, Blunder 7.1.0 and Drofa 2.0.0 are named by word mark and version because the measurement conditions require it; no connection to their authors and no endorsement by them exists.
Previous measurement: Coherent 0.1.58
Coherent 0.1.58: approximately 2300 ± 60 — transferred over 2 opponents that disagree with each other.
published rating + measured difference
The two measurements
| Opponent | Published rating | Games | Score rate | Measured difference | Anchored Coherent |
|---|---|---|---|---|---|
| Blunder 7.1.0 | 2388 ± 18 | 548 | 41.88 % | −56.93 ± 28.02 | 2331 ± 33 |
| Drofa 2.0.0 | 2420 ± 19 | 600 | 30.50 % | −143.07 ± 26.28 | 2277 ± 32 |
All error figures are half-widths at the 95% level.
The two answers lie 54 points apart: 2.28 standard deviations; the two-sided chance probability is 0.022.
This is a transfer, not a place in a ranking. A shared error in transferring the ratings could shift both anchors in the same direction and is not covered by the published band.
Blunder 7.1.0 against Drofa 2.0.0 directly in our own setup: completed.
The two anchors play each other directly in our setup
| Property | Value |
|---|---|
| Games | 442 |
| Pairs | 221 |
| Pentanomial (Blunder) | 68 / 49 / 65 / 28 / 11 |
| Score rate (Blunder) | 34.73 % |
| Measured difference | 109.6 ± 30.1 |
| Difference from the ranking list | 77.6 |
Explanation 1 confirmed: published ratings do not transfer to our conditions. The measured gap is more than three times the one in the ranking list.
Measured in two parts (200 and 242 games, different book positions); part 2 was ended early on the operator's instruction, the games played remain valid. Merged at pair level, not across the two Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. estimators.
Blunder 7.1.0: The run against Blunder ended at game 549 because the opponent crashed and the tournament manager stopped there. The aborted game is not counted.
Measurement conditions of the anchorings of 0.1.58
| Property | Value |
|---|---|
| Time control | 120+1 |
| Opening book | 8moves_v3 · 34,700 positions |
| Hash (MB per side) | 64 |
| Concurrent games | 4 |
| Adjudication | none |
| Procedure | fixed number of games, no early stop |
Data delivery created on 2026-09-07 · Measurements, interpretation and limitations
IMS says how the engine fares against a frozen reference configuration of its own — and nothing else. Anyone who wants to know where the program stands in the field needs opponents that somebody else built and that somebody else rated.
IMS lies above both anchors of 0.1.58. As a measure of progress between versions of our own it remains usable; as a statement about strength in the field it was too optimistic.
There is no conversion between IMS and Elo, no shared axis and no shared chart with two scales.
IMS — internal progress
IMS = 2200 + Elo difference against R1
IMS is a difference against a frozen reference configuration, not a rating from a ranking list.
The zero point is calibrated to the setting of this configuration, not to a public pool; comparisons with public ranking lists are therefore not permitted.
The reference configuration plays throttled and non-deterministically, so a repetition can easily deviate.
Measurement plot
Today this is one measurement with its error band, not a trend; the curve appears only as further builds are measured.
Measurements
Measurements and conditions — project measurement, not in the register
| Build | Date | IMS | Half-width | Games | W/D/L | Pentanomial | Reference |
|---|---|---|---|---|---|---|---|
| 0.1.58 | 2026-09-06 | 2363 | ±26 (nominal 95%) | 800 (fixed game count, no SPRT) | 558 / 34 / 208 | 28 / 7 / 145 / 27 / 193 | R1 |
W/D/L: wins / draws / losses; pentanomial Games are played in pairs with swapped colours; the five counts are the pairs scoring 0, ½, 1, 1½ and 2 points.: game pairs scoring 0 / ½ / 1 / 1½ / 2 points; conditions per reference: see Reference configuration.
Reference configuration
Frozen reference R1
| Property | Value |
|---|---|
| Engine | Stockfish 19 |
| Build | sf_19, Linux x86-64 universal |
| Options | UCI_LimitStrength=true UCI_Elo=2200 |
| IMS | 2200 |
| Time control | 10+0.1 |
| Book | 8moves_v3 · 34,700 positions |
| Hash (MB) | 64 |
| Threads | 1 |
| Adjudication | none |
| Frozen | 2026-09-05 |
The reference configuration plays with limited strength and is not deterministic; a repeat measurement may differ slightly.
The authors of the engine used as the reference did not know of this measurement, took no part in it, and did not consent to it. The program was used solely as freely available software under its licence.
Methodology
Updated: