Chess language-model zoo, ONNX with KV cache

Small chess language models from the Hub, converted to ONNX with a KV cache so a game can be played incrementally on-device. Used by CrispChess as its "bot zoo". All credit for the models goes to their authors; this repository only converts them. Only models whose licence permits redistribution are included; each keeps its original licence.

Folder Upstream Author Licence Human moves matched* Size
chess_llama_68m bharathrajcl/chess_llama_68m bharathrajcl Apache-2.0 41% 136 MB
ChessSLM-PM FlameF0X/ChessSLM-PM FlameF0X Apache-2.0 34% 101 MB
amdchess-v9 nlpguy/amdchess-v9 nlpguy Apache-2.0 32% 269 MB
grandpythia-200k-70m mlabonne/grandpythia-200k-70m Maxime Labonne Apache-2.0 28% 141 MB
dialochess DedeProGames/dialochess DedeProGames MIT 26% 328 MB
smolchess-v2 nlpguy/smolchess-v2 nlpguy Apache-2.0 21% 327 MB
Chesser-248K-Mini DedeProGames/Chesser-248K-Mini DedeProGames Apache-2.0 20% 497 MB
chessformer nsarrazin/chessformer nsarrazin MIT not measured (UCI) 474 MB

* Share of 90 positions from CC0 Lichess games (the maia1/maia5/maia9 bots' opponents) where the model's most likely legal move is the move actually played, with the game given in PGN spacing (1. e4 e5 2. Nf3). Chance is about 3%. The same models given only 1., as in the Chess LLM Arena, match 1-9%. chessformer reads UCI (e2e4 e7e5), one token per move.

Apache-2.0 text: LICENSE-APACHE-2.0.txt. MIT models: the notice "Copyright (c) the model's author, released under the MIT License" applies, with the standard MIT permission text. Each folder keeps the upstream model card (UPSTREAM_README.md) and config (upstream_config.json).

Graph

model_kv_fp16.onnx (opset 17, IR 8). Weights are stored as fp16 and cast to fp32, so computation is fp32.

  • inputs: input_ids int64 [b, t]; past_key_i, past_value_i float32 [b, kv_heads, p, head_dim] for each layer i (p = 0 on the first call)
  • outputs: logits [b, vocab] for the last position only; present_key_i, present_value_i [b, kv_heads, p + t, head_dim]

Feed the game once, then only new tokens; score candidate moves by extending a copy of the cache (batch = number of branches). Each graph was checked against the PyTorch model's full forward pass (incremental prefill and steps, batched branches, prompts of 5-150 tokens): max logit difference 1e-5 to 2e-4 (Pythia 1e-2, rotary rounding); fp16 storage adds up to ~0.15. Exported by export_chess_lm_kv.py.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support