MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp

An Apple MLX enhanced oQ5 (oQ5e) quantization of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B, built with oMLX using 128 calibration samples at sequence length 512, strict importance-matrix coverage, BF16 compute, and group size 64.

Key Configuration

  • Modality: Pure Text-Only (LLM). Vision tower weights and multimodal processor configs have been cleanly omitted, saving ~1.5 GB of unified memory.
  • Speculative Decoding (MTP): Includes the official 15-tensor Multi-Token Prediction draft head (model-mtp.safetensors, grafted from Qwen/Qwen3.5-9B). When loaded in oMLX with mtp_enabled: true, speculative decoding achieves ~1.4x–1.6x faster token generation on text tasks.

Validation

Validated locally before upload:

  • Strict model and MTP module instantiation through the oMLX runtime
  • Global 5-bit affine quantization metadata with group size 64
  • Complete strict oQe importance-matrix application
  • Tokenizer SHA-256 identity verified against official source
  • Official chat template renders OpenAI-style function schemas and thinking blocks cleanly
  • Deterministic generation smoke test verified

Usage with oMLX

To enable speculative decoding acceleration:

// ~/.omlx/model_settings.json
"MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp": {
  "mtp_enabled": true,
  "mtp_num_draft_tokens": 3,
  "temperature": 0.6,
  "repetition_penalty": 1.05
}

License

MIT, following the upstream model.

Downloads last month
277
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xunlinkx/MiMo-V2.6-Distill-Qwen-9B-Text-oQ5e-mtp

Finetuned
Qwen/Qwen3.5-9B
Quantized
(56)
this model