Kodiak small โ€” research preview v2

Research preview. An early model, shared while we build in public. It's fast and often right, and it also makes mistakes (see Known weaknesses). Code, docs, evaluation and the whole build story: https://github.com/grizzlypeaksoftware/kodiak ยท Demo: https://huggingface.co/spaces/comgen42/kodiak-demo ยท Feedback and failure cases welcome as GitHub issues. Previous preview: kodiak-small-r1-preview.

Kodiak is an encoder-only "System One" decision model by Cortex Agent LLC: a state (text, list of texts, or JSON) plus typed questions in, calibrated answers out, in one forward pass. Choice answers are always one of your labels, scores stay inside your range, and every question can come back as "not answerable from this state" with its own probability.

# pip install "kodiak-s1[infer] @ git+https://github.com/grizzlypeaksoftware/kodiak"
from kodiak_s1.hub import Kodiak

kodiak = Kodiak.from_pretrained("cortex-agent-llc/kodiak-small-v2-preview")   # GPU if available, else CPU
kodiak.decide(
    {"order_id": "A-1042", "status": "delivered", "message": "The box arrived crushed and the lamp is broken."},
    [{"type": "choice", "id": "intent", "text": "What does the customer want?",
      "labels": ["refund or replacement", "delivery status", "cancel order", "product question"]},
     {"type": "choice", "id": "carrier", "text": "Which carrier delivered it?", "labels": ["UPS", "FedEx", "USPS"]}],
)
# -> intent: "refund or replacement" (0.61); carrier: abstains (p_null 0.96) because the state never says.
#    Honest note: this fix is fragile. Drop the order_id and v2 leans "delivery status" (0.56); the r1 preview got it wrong both ways (0.53, 0.85).

What's new in v2

Trained on the public datasets plus Generator v2 synthetic data (9,137 training examples): a 971-document-type taxonomy, ~45% real web passages, questions labeled stated / inferred / unanswerable, and two AI critics. Compared with the same model trained on the v1 synthetic data at equal size, averaged over three training runs each:

v1 data v2 data (this model's recipe)
Wrong refusals on never-seen tasks (forced โˆ’ plain accuracy) 4.3 pts 1.7 pts
When it abstains, how often it's right 84% 92%
Calibration error (ECE) 0.038 0.029
Never-seen tasks, forced to answer 72.1% 72.0% (no change)
Familiar tasks 81.4% 81.7%

So v2 answers inference questions instead of refusing them and its "can't tell" is more trustworthy; it is not better at never-seen classification tasks. Details: GENERATOR_V2.md ยง11, decision D29.

This checkpoint

  • Run b-small-s1-B-v2 (seed 0, chosen among three seeds by validation loss, not the eval set), step 6000. ModernBERT-base backbone (Apache-2.0), 152M parameters. Calibration temperatures baked in; default abstain threshold 0.70 (override per request with null_threshold).
  • Frozen eval set v0.1: overall 0.795, familiar tasks 0.813, never-seen tasks 0.703 (0.721 forced), ECE 0.028, abstain precision 0.92.
  • Speed: ~8 ms per request on a GPU, ~80 ms on 8 ARM CPU cores.
  • Includes handler.py for Hugging Face Inference Endpoints (Deploy โ†’ Inference Endpoints; CPU is enough).
  • Includes model.onnx (fp32, calibration baked in) for self-hosting without Python: the Node.js + Docker server in server/ answers in ~30 ms per request on 8 CPU threads. Download this repo into a folder and point KODIAK_MODEL at it.

Known weaknesses

  • Fragile intent calls: small changes to the input can flip close calls (see the example above); check confidence and escalate low ones.
  • Judgment scores on new scales can be badly off: it rates "I was charged twice and nobody answers my emails!" as barely urgent.
  • Never-seen category-inference tasks (e.g. occupation from a biography): roughly on par with open zero-shot classifiers, well behind large LLMs.
  • Agent tool routing when the right tool is only implied; states longer than 512 tokens are truncated; English only.
  • Not for high-stakes decisions about people without human review. Bias in Bios (occupation) is a held-out task with known gender bias.
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.2B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cortex-agent-llc/kodiak-small-v2-preview

Quantized
(78)
this model

Space using cortex-agent-llc/kodiak-small-v2-preview 1