First external measurement: IMS 2363 ± 26
Project measurement, not in the register. Every figure is taken from the internal measurement report of 6 September 2026, the binding IMS Internally measured strength: 2200 plus the Elo difference against a frozen reference configuration of our own. A progress measure between our versions, not a rating from a ranking list. scale definition of 5 September 2026, and the accompanying data delivery. None of the figures is taken from the public data contract.
Project measurement, not in the register. Coherent 0.1.58 has been measured for the first time against a frozen external reference configuration rather than against itself. Over 800 games, a number fixed in advance, the score was 71.88 %, which is IMS 2363 ± 26 at the 95 % level. IMS is a difference against that reference, not a rating from a public ranking list.
Project measurement, not in the register. Until this measurement, Coherent had only an internal yardstick: each candidate against its predecessor, and in earlier entries a long stage against a fixed early build of the engine itself. A comparison against the predecessor says whether a change is stronger than the last one. It says nothing about where the engine stands. On 5 and 6 September 2026, version 0.1.58 played 800 games, a number fixed in advance, against a frozen reference configuration. The score was 71.88 %, which is IMS 2363 ± 26 at the 95 % level. It is the first strength figure for Coherent that comes from a match against an external opponent.
The result
Over 800 games, a number fixed in advance, Coherent 0.1.58 scored 575.0 points. All 800 games ended normally. There were no timeouts and no crashes.
Project measurement, not in the register.
IMS (Coherent 0.1.58) | 2363 ± 26 (95 %) |
|---|---|
| Score against R1 | 71.88 % (575.0 of 800) |
| Difference against R1 (Elo units) | +162.99 ± 26.11 |
| Games | 800, fixed in advance |
| Wins / losses / draws | 558 / 208 / 34 |
| Pentanomial | 28, 7, 145, 27, 193 |
| Terminations | 800 of 800 normal; 0 timeouts; 0 crashes |
The error band is the pentanomial Games are played in pairs with swapped colours; the five counts are the pairs scoring 0, ½, 1, 1½ and 2 points. half-width, nominally 95 %. Where no error band is stated, the sources do not give one.
The reference configuration R1 was frozen on 5 September 2026.
Project measurement, not in the register.
| Reference | R1, frozen 2026-09-05 |
|---|---|
| Engine | Stockfish 19, official release |
| Options | strength limiting on, strength setting 2200, hash 64 MB, one search thread |
| Time control | 10+0.1 (ten seconds plus a tenth of a second per move) |
| Opening book | 8moves_v3, 34,700 positions, random order, both colours per position |
| Adjudication / own book / tablebases | none |
| Concurrent games | 4 |
R1 plays throttled and is not deterministic, so a repeat of the same match can differ slightly.
What IMS measures
IMS stands for Internal Measured Strength — the same abbreviation in every language, so that multilingual editions do not drift apart. The unit is Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating.. The quantity is not Elo. IMS of Coherent is 2200 plus the Elo-difference against the frozen reference configuration R1.
A score of 50 % is IMS 2200. The 71.88 % scored here is a difference of +162.99 ± 26.11 Elo. IMS is 2200 plus 162.99, rounded to 2363; the half-width 26.11 is rounded to 26.
Only the difference between Coherent and R1 is measured. A run of 800 games takes no position on Coherent’s strength in any public pool. It records where Coherent stands relative to R1. Exactly one thing is assumed: that R1 lies at 2200. That zero is a convention, taken from the reference’s strength-limit setting. A measured rating in a public pool it is not.
There are two uncertainties, and they behave differently.
Project measurement, not in the register.
| Question | Shrinks through | Status | |
|---|---|---|---|
| Measurement error | Where does Coherent stand relative to R1? | more games against R1 | ±26 (800 games, 95 %) |
| Scale error | Does R1 really lie at 2200? | only opponents with a published rating | unknown |
The measurement error shrinks with every further game against R1. The scale error is unknown and does not shrink, however many games are played against R1. Mixing the two treats ±26 as the accuracy of the figure 2363. It is not.
The clock decides the value of a measurement
Two properties of the machine were checked before the match, because both can devalue a time measurement without leaving a trace.
A timestamp costs under 30 nanoseconds on the measurement machine. Two runs measured 13.8 and 22.0 nanoseconds; the citable figure is the bound, not the more favourable of the two. The same call on the operator’s workstation costs 29,356 nanoseconds. Both figures come from the same clock, so they can be compared. What was measured is the price of one call, not how often an engine makes it; how much search time that costs in a game does not follow from these figures. Nothing in the match output would show it either way. That finding has devalued earlier tournament figures of the project. Both machines are guests. The one used here reads the processor’s own timestamp counter. The workstation is a guest too, and other clocks on it are cheap: a coarse tick counter and a system-time call each cost one to two nanoseconds there. Why that one counter is so expensive was not established. The hardware counter behind it was not identified, and how much the virtualisation contributes is open. That it is emulated is the obvious reading, and no more than that. What was measured is the price of the call.
Hybrid cores, and what a guest cannot see
The measurement machine is a guest under Hyper-V with 16 logical processors on 8 cores, two threads per core, one socket, running Ubuntu 26.04.1 LTS. The processor is a 13th Gen Intel Core i9-13980HX.
That model is a hybrid design, with fast cores and power-saving cores. From inside the guest the split is invisible: the kernel exposes no core-type information, so how many cores of which kind the host has cannot be measured from here. To name a split would be to quote a specification, not to report a measurement.
It also does not matter. What matters is whether the guest’s own processors behave alike, because if the two engines of one game sat on different kinds, both would get the same clock time but one would get roughly twice the search out of it. So that was measured. The same fixed-depth search, bound to each of the 16 processors in turn, produced the same node count of 454,482, at speeds between 698,129 and 889,397 nodes per second — a spread of 27 per cent, and no second cluster at half speed. Power-saving cores among them would have produced one.
That the guest therefore sits on the fast cores is a plausible inference, and the two threads per core point the same way. It was not observed directly.
Self-check against the same binary
Before the external match, Coherent played itself under the same conditions, identical binary.
Project measurement, not in the register.
| Run | Games | Score |
|---|---|---|
| 1 | 100 | 55.0 % |
| 2 | 300 | 48.17 % |
| Combined | 400 | 49.88 % |
The expected score is 50 %. The first run at 55.0 % is described as lying within the error band of a fair match. Had 55.0 % been a real bias, every later measurement would have been shifted by the same amount. That is why the larger run was played. Over 400 games the score is 49.88 %, against an expectation of 50 %. That is the observed rate; it does not prove that no bias exists. All 400 games ended normally, with no timeouts and no crashes.
How many games, and how many of them are independent
A test that stops once the result looks favourable overestimates systematically. For a point estimate the game count is fixed in advance, and then every game is played.
A preliminary run of 40 games scored 53.75 %. The 800 games scored 71.875 %. A gap of 18.125 percentage points is wide, and it was checked rather than assumed.
The 800-game run is flat within itself. In blocks of 100 games the score runs 70.5, 73.5, 72.5, 71.0, 72.5, 75.5, 70.0 and 69.5 per cent — all within the spread of a block of this size, where the standard error is 5.0 points. That figure treats each block as 100 independent games; counted in pairs, as the paragraph below argues they must be, it would be larger. Nothing drifted. The first 40 games of this run scored 71.25 per cent, not 53.75. The difference lies between the two setups, not inside the measurement.
The decisive quantity had been counted wrongly. The preliminary run played 40 games, but as 20 openings with both colours each — 20 independent units, not 40. On that basis the standard error is 11.18 points rather than 7.91, and the gap of 18.125 points is 1.62 standard errors. That is ordinary. Counted as 40 independent samples the same gap looks like 2.29 standard errors, and one starts looking for a cause the numbers do not call for.
These standard errors are the conservative approximation from the square root of 0.25 over the number of units, not a pentanomial estimate. They serve the comparison between the right and the wrong count. They are not a significance test.
Where openings are played in pairs, the pairs are what count.
A curve of one point
The public IMS series against R1 begins with this point. A curve from one point is not a trend. It is a measurement with an error band. The curve exists only once further versions have been measured, and then it shows exactly what IMS is for: progress across versions, against an unchanged reference.
What the figure does not say
IMS is a difference against a frozen reference configuration, not a rating from a ranking list. The zero point is calibrated to the setting of this configuration, not to a public pool; comparisons with public ranking lists are therefore not permitted. The reference configuration plays throttled and non-deterministically, so a repetition can easily deviate.
The three sentences above are translated; the German wording is binding.
The authors of the engine used as the reference did not know of this measurement, took no part in it, and did not consent to it. The program was used solely as freely available software under its licence.
What comes next
The next measurements are matches against open-source engines for which a published rating exists. Those matches take place at the time control of that list. At a different time control the published rating would not transfer. Only that yields a figure an outsider can check. It then stands as its own curve beside IMS and is not converted into IMS.