Kefuro-30B

Weights are not public. The code is open (Apache-2.0). Everything needed to run Kefuro on top of the public NVIDIA base model, train your own adapter on your own tickets, and evaluate it is in this repository.

Kefuro-30B is a customer service model for ecommerce, built on NVIDIA Nemotron 3.5 Lightning 30B-A3B (30B-parameter hybrid Mamba-Transformer mixture-of-experts, about 3B active parameters per token) and post-trained on real support outcomes.

Base model NVIDIA Nemotron 3.5 Lightning 30B-A3B (OpenMDW-1.1)
Task Ecommerce support replies and per-store intent classification
Languages English, Spanish, French, German, Italian, Japanese
Serving vLLM (BF16 or NVFP4), one RTX PRO 6000 or H100 per replica

What we have measured so far (internal preview, not a release claim)

On our own held-out data for store-specific intent classification (173 blind test items across 212 store catalogs):

Accuracy
Base model, same prompt 57.2%
Kefuro-30B intent adapter (preview) 89.6%
Kefuro-30B, auto-handled only (model confidence ≥ 0.9, 82% of traffic) 97.2%

Low-confidence requests are routed to a human or a fallback instead of being answered automatically. These numbers will be re-measured on the public Kefuro-Bench, with human-verified labels, before any weights are released.

Full evaluation and performance report: EVAL.md (intent accuracy, confidence-gated automation, Kefuro-Bench, throughput on RTX PRO 6000, training cost, Nemotron-H engineering notes).

Use it without our weights

git clone https://huggingface.co/realset/Kefuro-30B && cd Kefuro-30B
pip install -r requirements.txt

# 1) Serve the public NVIDIA base model (1× RTX PRO 6000 / H100 is enough; NVFP4 needs Blackwell)
vllm serve nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 --served-model-name kefuro-base \
  --trust-remote-code --max-model-len 16384 --enable-lora --max-lora-rank 16

# 2) Run the intent / ticket-tagging service (OpenAI-compatible, drop-in for tool-call classifiers)
KEFURO_LORA_URL=http://localhost:8000/v1 KEFURO_LORA_MODEL=kefuro-base \
  uvicorn cs_model.intent_service:app --port 8820

The service accepts the same chat/completions request you already send to a hosted LLM for classification (a forced tool call whose argument is intent or tag, with the catalog in the prompt) and returns the same shape, plus a kefuro field with the route and confidence. Requests whose confidence is below KEFURO_CONF_THRESHOLD (default 0.9) are marked needs_review, or forwarded to KEFURO_FALLBACK_URL if you set one.

Train your own adapter

Put your labeled tickets in data/intent/ (train.jsonl, validation.jsonl, test_for_inference.jsonl, labels.json; one chat per line with a system message holding the store's intent catalog, the customer message as user, and {"intent": "<label>"} as assistant). Then on a GPU box:

pip install --no-build-isolation causal-conv1d mamba-ssm      # fused Mamba kernels, required for long catalogs
python training/intent/train_lora.py data/intent models/my-adapter --epochs 2
curl -X POST http://localhost:8000/v1/load_lora_adapter -d '{"lora_name":"my-adapter","lora_path":"models/my-adapter"}'
python training/intent/eval_gen.py http://localhost:8000/v1 my-adapter results/eval.json

Notes from our runs: LoRA must target q/k/v/o_proj and the Mamba in_proj only (PEFT rejects out_proj/conv1d on Nemotron-H); computing logits only for the answer tokens keeps peak memory near 62 GB on a 96 GB card; vLLM can hot-load the adapter (VLLM_ALLOW_RUNTIME_LORA_UPDATING=True).

What else is here

Path What it does
cs_model/intent_service.py Intent / tagging service with confidence-based routing (single model) or Gate + LoRA cross-check (dual)
cs_model/pipeline.py, gates.py, generator.py Reply pipeline: hard rules → decision gate → generation → quality gate → escalate
cs_model/hard_rules.py Multilingual must-escalate rules (safety, legal, account security, urgent operations, human requests)
data/scrub.py, data/build_datasets.py PII scrubbing and SFT / DPO / decision-label dataset builders
training/ LoRA SFT for Nemotron, Laya-based gate fine-tuning (RLCD + temperature calibration), translation for multilingual data
bench/ Kefuro-Bench v0 (synthetic, 6 languages) and its deterministic scorer
deploy/ Rent a GPU on Vast.ai and start vLLM, throughput benchmark

Run the tests with pytest -q tests bench/tests.

License

Code in this repository: Apache-2.0 (see LICENSE). The base model is NVIDIA Nemotron 3.5 Lightning under OpenMDW-1.1; your use of it is governed by that license.

Attribution

Kefuro-30B is derived from NVIDIA Nemotron 3.5 Lightning, which is released by NVIDIA under the OpenMDW-1.1 license. "NVIDIA" and "Nemotron" are trademarks of NVIDIA Corporation; Kefuro is not affiliated with or endorsed by NVIDIA.

Contact: support@flatkey.ai · https://kefuro.com

Downloads last month
106
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for realset/Kefuro-30B

Finetuned
(19)
this model