Kev 9B v1.0 โ€” packed GGUF

Kev-9B is a Jev-style "System One" decision model: a Qwen3.5-9B backbone plus a trained pointer head that returns calibrated probabilities for typed questions instead of generating text.

Built from Kev v1.0 โ€” jaredpalmer/kev-9b @ db029f0 (LoRA adapter + head.pt) merged onto Qwen/Qwen3.5-9B-Base @ 68c46c4b with espetro/llama.cpp's tools/kev pipeline (kev_v10_merge.py -> convert_hf_to_gguf --no-mtp -> kev_head.py -> kev_pack.py). The pointer head and calibration temperature (2.1936) are baked in as dec.head_* tensors + kev.* metadata.

Files

file size use
kev-9b-q8_0.gguf 9.7 GB recommended for inference
kev-9b-f16.gguf 18 GB re-quantization source / zero-drift reference

Use with the Kev-enabled llama.cpp fork

llama-server -hf espetro/kev-9b-gguf:Q8_0
# -> serves TypeSafe /v1/systemone + /studio automatically

curl localhost:8080/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Shoes arrived two weeks late and in the wrong size.",
  "questions": {
    "department": {"type": "choice", "instructions": "Which team handles this?",
                   "criteria": {"returns": "Exchanges, refunds", "shipping": "Delays, lost packages"}}
  }
}'
llama-decide -hf espetro/kev-9b-gguf:Q8_0 --json request.json

The files still load in stock llama.cpp as an ordinary Qwen3.5 model โ€” the kev.* metadata and head tensors are simply ignored there.

In the browser

espetro.github.io/llama.cpp runs the 0.8B demo quant in llama.cpp WASM. For a smaller local download of this model, see espetro/kev-9b-demo-gguf.

Source & license

Downloads last month
440
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support