Spark-X2.5-1.7B-GGUF

Spark-X2.5 is a compact, general-purpose language model for conversation, writing, translation, reasoning, coding, tool use, and agentic workflows. It uses a hybrid attention architecture, supports a native context length of up to 1M tokens, and covers more than 200 languages. For its architecture, training methods, benchmark results, fine-tuning, and citation, see the Spark-X2.5-1.7B.

Local Deployment

llama.cpp

Native spark2_5 support requires llama.cpp b10828 or later.

Run the model in the CLI:

llama cli -m /path/to/Spark-X2.5-1.7B-GGUF

Launch the OpenAI-compatible API server:

llama serve -m /path/to/Spark-X2.5-1.7B-GGUF

Ollama

Download Ollama

Native spark2_5 support requires Ollama v0.34.1 or later.

ollama run SparkLLM/Spark-X2.5-1.7B

LM Studio

Download LM Studio

Native spark2_5 support requires runtime version 2.34.0 or later.

Unsloth

Download Unsloth Studio

The latest Unsloth Studio supports native Spark-X2.5 GGUF inference.

License

Released under the Apache License 2.0.

Downloads last month
54,190
GGUF
Model size
2B params
Architecture
spark2_5
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Model tree for XHToken/Spark-X2.5-1.7B-GGUF

Quantized
(8)
this model
Quantizations
1 model

Collection including XHToken/Spark-X2.5-1.7B-GGUF