Run T-0020 failed
· Evidence: SPRT Sequential probability ratio test. Games are played until the accumulated evidence reaches one of two bounds; then the test stops.
Hypothesis
The engine should play more strongly because it will preserve winnable minor-piece endings, recognize repetitions despite irrelevant en-passant fields, and keep halfmove- and path-dependent draw scores from contaminating the transposition table.
Stages
| Field | short stage |
|---|---|
| Verdict | failed |
| Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. | −20.67 ± 10.40 |
| nElo Normalised Elo: the Elo difference divided by the spread of the game results, so that the SPRT bounds mean the same at different time controls. | −27.76 ± 13.93 |
| Games | 2,390 |
| Wins / draws / losses | 849 / 550 / 991 |
| pentanomial Games are played in pairs with swapped colours; the five counts are the pairs scoring 0, ½, 1, 1½ and 2 points. | 142 / 220 / 565 / 174 / 94 |
| LLR Log-likelihood ratio, the running evidence of an SPRT. It starts at 0; a stage passes at the upper bound (about +2.94) and fails at the lower one (about −2.94). | −2.98 |
| Bounds | −2.94 … +2.94 |
| SPRT Sequential probability ratio test. Games are played until the accumulated evidence reaches one of two bounds; then the test stops. | [0.0, 5.0] |
| Error rates α / β | 0.05 / 0.05 |
| Model | normalized |
| time control Base time plus increment per move in seconds, e.g. 8+0.08: eight seconds per game plus 0.08 seconds per move. | 8+0.08 |
| Book | A |
| Opponent | vs predecessor |
| Duration | 12 min |
Provenance
- Date
- 2026-08-29
- Started
- 2026-08-29 05:03:58 UTC
- Duration
- 12 min
- Author
- Agent B
- Bench nodes
- 5,073
- Commit
e0fce07f54c762efe78e3599a8656f15cc2f4927- SHA-256 A cryptographic hash. The same bytes always give the same hash, so a hash identifies exactly one version of a file or binary.
204b6fa28105b771b7ab0e13d52c41fce407aa00ce01559446bc59b1400593f9