Plumb-4B

A 4B decision model: give it evidence, a question and 2-16 options, and it returns a probability for every option from one forward pass, with no generated tokens. Probabilities use a single temperature (T = 2.07) fitted on held-out decisions. Trained on hard decisions written and double-checked by Qwen3.8-27B, with the questions it got wrong or was unsure of weighted up.

Split n Accuracy ECE p50 latency
easy 48 1.000 0.056 28 ms
standard 72 0.972 0.152 28 ms
hard 111 0.802 0.094 78 ms

On JevBench hard, Plumb-4B scores 89/111 (80.2%) versus 82/111 (73.9%) for JevK5 v0.2: 8 items fixed, 1 regressed. Easy and standard accuracy are unchanged.

Public JevBench items, JevBench's own runner, RTX 4080 Super. No JevBench item was used for training, tuning, checkpoint selection or calibration. Aggregate results on the public JevBench set did inform later training-recipe decisions; see https://github.com/crh225/plumb for the full disclosure, audits and per-item results.

Training data: crh225/plumb-decisions. GGUF for llama.cpp and Ollama: crh225/plumb-4b-GGUF.

Use it

It runs with the open jevk5 runtime (Apache-2.0), which reads the temperature from jevk5_config.json in this repo. It needs a CUDA GPU with about 10 GB free (bf16).

pip install "jevk5[fast] @ git+https://github.com/allebee/jevk5@v0.2.0"

From Python, one decision per call:

from jevk5 import JevK5

model = JevK5("crh225/plumb-4b")

model.decide(
    "Refunds need a receipt and a purchase within 30 days. "
    "The customer bought 12 days ago and has no receipt.",
    {"type": "noul", "instructions": "Is a refund permitted under the policy?"},
)
# {'type': 'noul', 'noul': <probability of true>, 'confidence': ..., 'input_tokens': ...}

model.decide(
    "Order #7120 shows delivered to No. 17; the customer lives at No. 71.",
    {"type": "choice", "instructions": "What happened to the parcel?",
     "criteria": ["delivered", "misdelivered", "unknown"]},
)
# {'type': 'choice', 'choice': ..., 'probabilities': {...}, 'confidence': ..., 'input_tokens': ...}

Question types: noul (true/false), choice (2-16 options, a list or a {key: description} map) and score (ordinal levels). Every answer is a full probability over the options, from one forward pass, with no generated tokens.

As a service, with a TypeSafe-style request shape (several questions about one state per request):

jevk5-serve --model crh225/plumb-4b --port 8090
curl -s localhost:8090/v1/systemone -d '{"state": "I was billed twice, please refund",
  "questions": {"refund": {"type": "noul", "instructions": "Asks for money back?"}}}'

v5.2: reading answers for JevBench v1.5

JevBench v1.5 counts a yes/no answer whose P(yes) lies between 0.2 and 0.8 as an abstention, scored wrong, and grades score questions on the expected level of the returned distribution. v5.2 keeps these weights unchanged (revision 55de0378) and changes only how answers are read, with Plumb's own server from github.com/crh225/plumb @ 0e6d558:

git clone https://github.com/crh225/plumb && cd plumb && git checkout 0e6d558d3ea3002c300ce121db4d3e8df6ad792a
pip install -r serve/requirements.txt          # torch from your CUDA environment
hf download crh225/plumb-4b --revision 55de037801a8a9b9de3db5c0e16cef86210c2186 --local-dir plumb-4b
python serve/plumb_server.py --model ./plumb-4b --port 8090 --one-read --temperature 2.07 \
    --noul-commit --score-temperature 1.2

Plumb applies a deterministic output transform for JevBench v1.5 yes/no decisions: predictions whose probability falls inside the benchmark's non-credit band are committed to the nearest boundary, while predictions already outside the band are unchanged. Score questions are read at a lower temperature (1.2) than choice questions (2.07). Model weights are unchanged. Calibration is evaluated on the transformed probabilities. Both settings were chosen on Plumb's own development sets (fit half), measured on their judge half, and frozen before any JevBench run; the public items were run once afterwards as a report.

Measured with the v1.5 rules, same weights before and after:

yes/no, chance-corrected yes/no abstentions score, chance-corrected
Own dev sets, judge half 61.8 → 77.2 12.6% → 0% 71.4 → 75.5
Held-out lockbox (556) 83.0 → 88.7 3.8% → 0% 88.6 → 89.4
Held-out lockbox (3,002) 69.1 → 81.1 9.3% → 0% 77.8 → 81.5
JevBench public, 231 (report only) 37.8 → 81.1 24.3% → 0% 82.3 → 89.5

Choice answers are unchanged.

Credits and licence

Apache-2.0. Fine-tuned from JevK5 v0.2 (Apache-2.0), itself built on Qwen3.5-4B (Apache-2.0), with a readout from SemIf (MIT); see NOTICE.

Downloads last month
744
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for crh225/plumb-4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(2)
this model
Finetunes
1 model
Quantizations
2 models

Dataset used to train crh225/plumb-4b