Kev-0.8B for kev-rs (mistral.rs)

Export of jaredpalmer/kev-0.8b in the layout kev-rs (the Kev server in the espetro/mistral.rs fork) loads directly:

  • model/: Qwen/Qwen3.5-0.8B-Base with the Kev LoRA merged in bf16, as a plain HF checkpoint
  • head.safetensors: the Kev pointer head (q, k)
  • kev.json: base, head_dim, temperature, dtype, special-token ids

Produced by kev-rs/scripts/export_checkpoint.py --dtype bf16. Parity against the bf16 torch reference (kev-rs parity, smoke-v1 development set): 0/40 argmax flips, max |dp| 0.018. An fp32 export lives on the fp32 branch of this repo.

kev-rs serve --checkpoint espetro/kev-0.8b-mistralrs --run jaredpalmer/kev-0.8b
# or quantized in-situ at load:
kev-rs serve --checkpoint espetro/kev-0.8b-mistralrs --run jaredpalmer/kev-0.8b --isq 8
# mixed: GDN recurrent path 8-bit, attention + MLP 6-bit (same parity, ~25% less weight)
kev-rs serve --checkpoint espetro/kev-0.8b-mistralrs --run jaredpalmer/kev-0.8b --topology topo.yml

Same Apache-2.0 license as the source checkpoint and base model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for espetro/kev-0.8b-mistralrs

Finetuned
(126)
this model