Instructions to use torchcast-ai/torchcast-decision-27b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use torchcast-ai/torchcast-decision-27b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="torchcast-ai/torchcast-decision-27b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("torchcast-ai/torchcast-decision-27b") model = AutoModelForMultimodalLM.from_pretrained("torchcast-ai/torchcast-decision-27b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use torchcast-ai/torchcast-decision-27b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "torchcast-ai/torchcast-decision-27b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "torchcast-ai/torchcast-decision-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/torchcast-ai/torchcast-decision-27b
- SGLang
How to use torchcast-ai/torchcast-decision-27b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "torchcast-ai/torchcast-decision-27b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "torchcast-ai/torchcast-decision-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "torchcast-ai/torchcast-decision-27b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "torchcast-ai/torchcast-decision-27b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use torchcast-ai/torchcast-decision-27b with Docker Model Runner:
docker model run hf.co/torchcast-ai/torchcast-decision-27b
Torchcast Decision 27B

Decision Index 0.2.1: 65.00, +7.09 over Jev 1.13. One yes/no question in 33 ms on one H100.
Torchcast Decision 27B answers typed questions about a state: pick one of several options (choice), yes or no
(noul), or a rating on an ordinal scale (score). Every question comes back with a probability for each option, and
all questions of a request are answered together, in a single forward pass for typical requests. Requests and responses use the TypeSafe-compatible
/v1/systemone format, so clients written for the choice, noul and score questions of Jev (TypeSafe AI) should
work without changes. Context up to 262,144 tokens (256K); the
evaluations below are text-only.
Results
| Torchcast Decision 27B | |
|---|---|
| Decision Index 0.2.1 | 65.00 |
| Latency, one yes/no question | 33.0 ms |
| Latency, one request with three questions | 49.0 ms |
| Input | text or JSON, up to 262,144 tokens |
Decision Index: the official kit (apolinario/decision-index, release
0.2.1, commit 87d4650) on all 155,390 suite rows, coverage 1.0, scored by the kit. The model runs on one H100
80GB; rows were answered in bf16 by our batched runner (eval/di_batch.py in the GitHub repository), with the suite
split between two such GPUs, each running a full copy of the model, to halve the wall-clock time. Latency is
end to end over HTTP on one H100 80GB in bf16, one request at a time, mean of 100 after 20 warm-up requests
(p95 34.6 ms and 52.1 ms); the three-question request is the ticket example under Usage.


Torchcast Decision 27B is ahead of Jev 1.13 on 31 of the 38 index benchmarks. The largest gains are on the home appliance simulator (tool use), POP909-CL, aspect-sentiment extraction (ACOS), stance (VAST), grade-school math, fact verification and contract reasoning. Jev 1.13 remains clearly ahead on hard knowledge benchmarks (GPQA Diamond, MMLU-Pro, BBH), and the Knowledge area of Torchcast Decision 27B is 5.8 points below it.
Compared with other decision models
| Model | Decision Index 0.2.1 | Source |
|---|---|---|
| Torchcast Decision 27B | 65.00 | this card |
| StartLux-Decision-27B | 63.88 | its model card |
| Clef | 61.21 | self-reported by Cloudflare, 2026-10-01 |
| Jev 1.13 | 57.91 | public board |
| Surogate Rune 26B-A4B v3 | 57.44 | public board |
| Decider chat · Gemma-4-31B | 57.33 | public board |
| pplx-decider-v1-27b | 56.40 | public board |
| Winnow-12B | 50.02 | public board |
Our score is author-run with the official kit and is not on the public board; board entries were run by the board's maintainers. Public board: Decision Index 0.2.1, snapshot of 2026-09-28. Latency is not compared across models because the published figures use different hardware and paths. On the same pipeline, the base StartLux-Decision-27B scores 63.99, within 0.11 of its published 63.88.
Fast inference
The folder ships the inference package torchcast_decision/, which is the fast path:
requirements.txtinstalls the fast kernels,flash-linear-attentionandcausal-conv1d, andpython -m torchcast_decision.check .confirms they are active. Without them transformers falls back to a much slower path, and the server refuses to start on a GPU.- The questions of a typical request share one forward pass; a very long state is read once in chunks first, and a choice with more than 26 options takes extra rounds. The server records CUDA graphs at start-up and replays them for short requests: on one H100 a single yes/no question takes 33.0 ms end to end and the three-question ticket request 49.0 ms.
- For bulk work,
decide_batchbatches the questions of many requests together.
Serve the model with the included package as shown under Usage. torchcast_decision/ is Apache-2.0 and
derived from StartLux-Decision's inference code; see its NOTICE.
Usage
Linux with an NVIDIA GPU with at least 64 GB of memory (54.7 GB of bf16 weights).
hf download torchcast-ai/torchcast-decision-27b --local-dir torchcast-decision-27b
cd torchcast-decision-27b
pip install -r requirements.txt # causal-conv1d may need --no-build-isolation
python -m torchcast_decision.check . # must print "fast kernels: active"
python -m torchcast_decision.server --model . --port 8090
curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
"state": {"ticket": "I was charged twice for order #4411 and the app still shows it as unpaid."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Payments, refunds and invoices",
"shipping": "Delivery and tracking",
"technical": "App, login and account problems"}},
"urgent": {"type": "noul", "instructions": "Should this ticket be answered today?"},
"severity": {"type": "score", "instructions": "How severe is the impact?",
"criteria": ["cosmetic", "annoying", "blocks the customer"]}
}
}'
The response (abridged) carries a probability for every option:
{"answers": {
"team": {"type": "choice", "choice": "billing", "confidence": 0.948,
"probabilities": {"billing": 0.965, "shipping": 0.001, "technical": 0.034}},
"urgent": {"type": "noul", "noul": 0.706},
"severity": {"type": "score", "score": 1.71, "confidence": 0.565,
"probabilities": {"0": 0.016, "1": 0.258, "2": 0.726}}},
"usage": {"input_tokens": 290, "output_tokens": 0},
"model": "torchcast-decision-27b"}
Or in Python, from the same folder:
from torchcast_decision import TorchcastDecision
m = TorchcastDecision(".")
answers, usage = m.decide(state, questions)
many = m.decide_batch([(state, questions), ...])
confidence follows TypeSafe's definitions (for a choice, (p_max − 1/n) / (1 − 1/n)); the top probability is in
probabilities. Serving and evaluation code:
Torchcast-AI/torchcast-decision-27b.
Training
A LoRA fine-tune of StartLux-Decision-27B, merged into the weights. Training inputs come from train and dev splits of
public datasets and benchmark-format decision data built from them, with the datasets' gold labels as targets. Some
sources carry non-commercial, share-alike or
research-only terms; see LICENSE.
Benchmark exposure
Training inputs came only from train and dev splits of public datasets; no Decision Index test rows were used, and all training data was screened against every test row. As expected for these datasets, about 1% of training rows share a source document or question template with a test item, and one CLINC150 utterance appears in both splits of that dataset. ForecastBench training rows all resolve before the test resolution window (which starts 2026-07-01). Some are earlier instances of recurring question series that also appear in the test set; each of those resolves before the earliest forecast date of the test questions on the same event, and prediction-market questions that appear in the test set were excluded.
Intended use
Research and non-commercial use. Not intended for safety-critical, security, medical, legal, financial or employment
decisions without human review; probabilities can be miscalibrated on new domains. No warranty (see LICENSE).
Limitations
- The scores are single-run point estimates without uncertainty intervals, and Decision Index results guided checkpoint selection, so they are not an untouched holdout.
- The Language area is slightly below the base model (73.81 against 74.52 on the same pipeline), and the Knowledge area trails Jev 1.13.
- Evaluated on English text only; image inputs are supported by the inherited vision tower but were not evaluated here, and probabilities can be unreliable on new domains.
- The base model's multi-token-prediction head is not included; the weights are for decisions, not speculative decoding.
License
Model weights: CC BY-NC 4.0, inherited from StartLux-Decision-27B; non-commercial use only, with attribution.
Commercial use is not permitted under this license. It would require separate permission from StartLux Labs
(contact@startlux.com) for StartLux-Decision-27B and from Torchcast AI for this fine-tune, and may also be restricted
by the terms of the training data. StartLux-Decision-27B
is a modified version of a model released by Alibaba Cloud under the Apache License 2.0; files taken unchanged from
it, such as the tokenizer files, remain under that license. torchcast_decision/ is Apache-2.0 code derived from
StartLux-Decision's inference code (see torchcast_decision/NOTICE). See LICENSE and NOTICE, which carry
StartLux's license and notices.
Credit Torchcast AI for this checkpoint, StartLux Labs for StartLux-Decision, and the Qwen team at Alibaba Cloud for the underlying model. This project is independent of, and not endorsed by, StartLux Labs, TypeSafe AI, Alibaba Cloud, Cloudflare, the Decision Index maintainers, or any other model provider named here. Product names are trademarks of their respective owners and are used only to identify those products.
- Downloads last month
- 652