Text Generation
MLX
Safetensors
qwen3_5
omlx
oq
oq5e
mtp
speculative-decoding
text-only
tool-use
conversational
5-bit
Instructions to use xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp
An Apple MLX enhanced oQ5 (oQ5e) quantization of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, built with oMLX using 128 calibration samples at sequence length 512, strict importance-matrix coverage, BF16 compute, and group size 64.
Key Configuration
- Modality: Pure Text-Only (LLM). Vision tower weights and multimodal processor configs have been cleanly omitted, saving ~1.5 GB of unified memory.
- Speculative Decoding (MTP): Includes the official 15-tensor Multi-Token Prediction draft head (
model-mtp.safetensors, grafted fromQwen/Qwen3.5-9B). When loaded in oMLX withmtp_enabled: true, speculative decoding achieves ~1.4x–1.6x faster token generation on text tasks.
Validation
Validated locally before upload:
- Strict model and MTP module instantiation through the oMLX runtime
- Global 5-bit affine quantization metadata with group size 64
- Complete strict oQe importance-matrix application
- Tokenizer SHA-256 identity verified against official source
- Official chat template renders OpenAI-style function schemas and thinking blocks cleanly
- Deterministic generation smoke test verified
Usage with oMLX
To enable speculative decoding acceleration:
// ~/.omlx/model_settings.json
"MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp": {
"mtp_enabled": true,
"mtp_num_draft_tokens": 3,
"temperature": 0.6,
"repetition_penalty": 1.05
}
License
MIT, following the upstream model.
- Downloads last month
- 277
Model size
9B params
Tensor type
U32
·
BF16 ·
Hardware compatibility
Log In to add your hardware
5-bit
Model tree for xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp
Base model
Qwen/Qwen3.5-9B-Base Finetuned
Qwen/Qwen3.5-9B Finetuned
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B