Kefuro-30B
Weights are not public. The code is open (Apache-2.0). Everything needed to run Kefuro on top of the public NVIDIA base model, train your own adapter on your own tickets, and evaluate it is in this repository.
Kefuro-30B is a customer service model for ecommerce, built on NVIDIA Nemotron 3.5 Lightning 30B-A3B (30B-parameter hybrid Mamba-Transformer mixture-of-experts, about 3B active parameters per token) and post-trained on real support outcomes.
| Base model | NVIDIA Nemotron 3.5 Lightning 30B-A3B (OpenMDW-1.1) |
| Task | Ecommerce support replies and per-store intent classification |
| Languages | English, Spanish, French, German, Italian, Japanese |
| Serving | vLLM (BF16 or NVFP4), one RTX PRO 6000 or H100 per replica |
What we have measured so far (internal preview, not a release claim)
On our own held-out data for store-specific intent classification (173 blind test items across 212 store catalogs):
| Accuracy | |
|---|---|
| Base model, same prompt | 57.2% |
| Kefuro-30B intent adapter (preview) | 89.6% |
| Kefuro-30B, auto-handled only (model confidence ≥ 0.9, 82% of traffic) | 97.2% |
Low-confidence requests are routed to a human or a fallback instead of being answered automatically. These numbers will be re-measured on the public Kefuro-Bench, with human-verified labels, before any weights are released.
Full evaluation and performance report: EVAL.md (intent accuracy, confidence-gated automation, Kefuro-Bench, throughput on RTX PRO 6000, training cost, Nemotron-H engineering notes).
Use it without our weights
git clone https://huggingface.co/realset/Kefuro-30B && cd Kefuro-30B
pip install -r requirements.txt
# 1) Serve the public NVIDIA base model (1× RTX PRO 6000 / H100 is enough; NVFP4 needs Blackwell)
vllm serve nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 --served-model-name kefuro-base \
--trust-remote-code --max-model-len 16384 --enable-lora --max-lora-rank 16
# 2) Run the intent / ticket-tagging service (OpenAI-compatible, drop-in for tool-call classifiers)
KEFURO_LORA_URL=http://localhost:8000/v1 KEFURO_LORA_MODEL=kefuro-base \
uvicorn cs_model.intent_service:app --port 8820
The service accepts the same chat/completions request you already send to a hosted LLM for classification (a forced tool call whose argument is intent or tag, with the catalog in the prompt) and returns the same shape, plus a kefuro field with the route and confidence. Requests whose confidence is below KEFURO_CONF_THRESHOLD (default 0.9) are marked needs_review, or forwarded to KEFURO_FALLBACK_URL if you set one.
Train your own adapter
Put your labeled tickets in data/intent/ (train.jsonl, validation.jsonl, test_for_inference.jsonl, labels.json; one chat per line with a system message holding the store's intent catalog, the customer message as user, and {"intent": "<label>"} as assistant). Then on a GPU box:
pip install --no-build-isolation causal-conv1d mamba-ssm # fused Mamba kernels, required for long catalogs
python training/intent/train_lora.py data/intent models/my-adapter --epochs 2
curl -X POST http://localhost:8000/v1/load_lora_adapter -d '{"lora_name":"my-adapter","lora_path":"models/my-adapter"}'
python training/intent/eval_gen.py http://localhost:8000/v1 my-adapter results/eval.json
Notes from our runs: LoRA must target q/k/v/o_proj and the Mamba in_proj only (PEFT rejects out_proj/conv1d on Nemotron-H); computing logits only for the answer tokens keeps peak memory near 62 GB on a 96 GB card; vLLM can hot-load the adapter (VLLM_ALLOW_RUNTIME_LORA_UPDATING=True).
What else is here
| Path | What it does |
|---|---|
cs_model/intent_service.py |
Intent / tagging service with confidence-based routing (single model) or Gate + LoRA cross-check (dual) |
cs_model/pipeline.py, gates.py, generator.py |
Reply pipeline: hard rules → decision gate → generation → quality gate → escalate |
cs_model/hard_rules.py |
Multilingual must-escalate rules (safety, legal, account security, urgent operations, human requests) |
data/scrub.py, data/build_datasets.py |
PII scrubbing and SFT / DPO / decision-label dataset builders |
training/ |
LoRA SFT for Nemotron, Laya-based gate fine-tuning (RLCD + temperature calibration), translation for multilingual data |
bench/ |
Kefuro-Bench v0 (synthetic, 6 languages) and its deterministic scorer |
deploy/ |
Rent a GPU on Vast.ai and start vLLM, throughput benchmark |
Run the tests with pytest -q tests bench/tests.
License
Code in this repository: Apache-2.0 (see LICENSE). The base model is NVIDIA Nemotron 3.5 Lightning under OpenMDW-1.1; your use of it is governed by that license.
Attribution
Kefuro-30B is derived from NVIDIA Nemotron 3.5 Lightning, which is released by NVIDIA under the OpenMDW-1.1 license. "NVIDIA" and "Nemotron" are trademarks of NVIDIA Corporation; Kefuro is not affiliated with or endorsed by NVIDIA.
Contact: support@flatkey.ai · https://kefuro.com
- Downloads last month
- 106