Chess language-model zoo, ONNX with KV cache
Small chess language models from the Hub, converted to ONNX with a KV cache so a game can be played incrementally on-device. Used by CrispChess as its "bot zoo". All credit for the models goes to their authors; this repository only converts them. Only models whose licence permits redistribution are included; each keeps its original licence.
| Folder | Upstream | Author | Licence | Human moves matched* | Size |
|---|---|---|---|---|---|
| chess_llama_68m | bharathrajcl/chess_llama_68m | bharathrajcl | Apache-2.0 | 41% | 136 MB |
| ChessSLM-PM | FlameF0X/ChessSLM-PM | FlameF0X | Apache-2.0 | 34% | 101 MB |
| amdchess-v9 | nlpguy/amdchess-v9 | nlpguy | Apache-2.0 | 32% | 269 MB |
| grandpythia-200k-70m | mlabonne/grandpythia-200k-70m | Maxime Labonne | Apache-2.0 | 28% | 141 MB |
| dialochess | DedeProGames/dialochess | DedeProGames | MIT | 26% | 328 MB |
| smolchess-v2 | nlpguy/smolchess-v2 | nlpguy | Apache-2.0 | 21% | 327 MB |
| Chesser-248K-Mini | DedeProGames/Chesser-248K-Mini | DedeProGames | Apache-2.0 | 20% | 497 MB |
| chessformer | nsarrazin/chessformer | nsarrazin | MIT | not measured (UCI) | 474 MB |
* Share of 90 positions from CC0 Lichess games (the maia1/maia5/maia9 bots'
opponents) where the model's most likely legal move is the move actually played, with
the game given in PGN spacing (1. e4 e5 2. Nf3). Chance is about 3%. The same models
given only 1., as in the Chess LLM Arena, match 1-9%. chessformer reads UCI
(e2e4 e7e5), one token per move.
Apache-2.0 text: LICENSE-APACHE-2.0.txt. MIT models: the notice
"Copyright (c) the model's author, released under the MIT License" applies, with the
standard MIT permission text. Each folder keeps the upstream model card
(UPSTREAM_README.md) and config (upstream_config.json).
Graph
model_kv_fp16.onnx (opset 17, IR 8). Weights are stored as fp16 and cast to fp32, so
computation is fp32.
- inputs:
input_idsint64[b, t];past_key_i,past_value_ifloat32[b, kv_heads, p, head_dim]for each layeri(p = 0 on the first call) - outputs:
logits[b, vocab]for the last position only;present_key_i,present_value_i[b, kv_heads, p + t, head_dim]
Feed the game once, then only new tokens; score candidate moves by extending a copy of the cache (batch = number of branches). Each graph was checked against the PyTorch model's full forward pass (incremental prefill and steps, batched branches, prompts of 5-150 tokens): max logit difference 1e-5 to 2e-4 (Pythia 1e-2, rotary rounding); fp16 storage adds up to ~0.15. Exported by export_chess_lm_kv.py.