Run T-0019 failed
· Evidence: SPRT Sequential probability ratio test. Games are played until the accumulated evidence reaches one of two bounds; then the test stops.
Hypothesis
Depth-preferred replacement with generation aging should improve move selection in longer games by preserving deeper transposition results from shallower collisions while allowing stale entries to yield immediately.
Stages
| Field | short stage |
|---|---|
| Verdict | failed |
| Elo Strength difference to the opponent, estimated from the games, with a 95-percent half-width where the oracle reported one. Never an absolute rating. | −0.58 ± 2.76 |
| nElo Normalised Elo: the Elo difference divided by the spread of the game results, so that the SPRT bounds mean the same at different time controls. | −1.01 ± 4.78 |
| Games | 20,294 |
| Wins / draws / losses | 7,728 / 4,804 / 7,762 |
| pentanomial Games are played in pairs with swapped colours; the five counts are the pairs scoring 0, ½, 1, 1½ and 2 points. | 566 / 1165 / 6693 / 1183 / 540 |
| LLR Log-likelihood ratio, the running evidence of an SPRT. It starts at 0; a stage passes at the upper bound (about +2.94) and fails at the lower one (about −2.94). | −2.95 |
| Bounds | −2.94 … +2.94 |
| SPRT Sequential probability ratio test. Games are played until the accumulated evidence reaches one of two bounds; then the test stops. | [0.0, 5.0] |
| Error rates α / β | 0.05 / 0.05 |
| Model | normalized |
| time control Base time plus increment per move in seconds, e.g. 8+0.08: eight seconds per game plus 0.08 seconds per move. | 8+0.08 |
| Book | A |
| Opponent | vs predecessor |
| Duration | – |
Provenance
- Date
- 2026-08-29
- Author
- Agent B
- Bench nodes
- 5,073
- Commit
8f5c6f7aab6cca6931fb2efcdae90dacf94d5f89- SHA-256 A cryptographic hash. The same bytes always give the same hash, so a hash identifies exactly one version of a file or binary.
7d24dfc65cd7b16edfcb3bc95db4eb26bc241d60bb0468ac75f09f305b28b85e