Kev-9B v1.0 (demo quant) - GGUF

Packed Kev "System One" decision model: Qwen3.5-9B backbone + trained pointer head, quantized for fast downloads. Answers typed questions (noul, choice, score) about a state with calibrated probabilities - no text generation.

This is a demo quantization. ~5.6 GB download for quick trials. Probabilities can drift vs the full-precision model and confidence is approximate - it shows how System One answers work, not what a production deployment should ship. For the best-calibrated artifact use espetro/kev-9b-gguf (q8_0 + f16).

Built with the Kev fork of llama.cpp: q4_k_m + a task-specific importance matrix (the token_embd override is already the default at 9B). Measured vs packed F16 on a 17-question probe set: 0 argmax flips, mean |dp| 0.023, max |dp| 0.136.

Use it

# server - TypeSafe /v1/systemone API
llama-server -hf espetro/kev-9b-demo-gguf

# CLI
llama-decide -hf espetro/kev-9b-demo-gguf --json request.json

Or try it in your browser right now - it runs fully client-side via WebAssembly:

https://espetro.github.io/llama.cpp/?model=https://huggingface.co/espetro/kev-9b-demo-gguf/resolve/main/kev-9b-demo-q4km-im.gguf

Source

Weights: jaredpalmer/kev-9b v1.0 adapter merged onto Qwen/Qwen3.5-9B-Base, packed with tools/kev/kev_pack.py, then quantized (llama-imatrix on a decision-input corpus, llama-quantize --imatrix ... q4_k_m). Kev/Jev research: https://github.com/jaredpalmer/kev

Downloads last month
284
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support