MTPLX: the fastest way to run Qwen 3.8 on a Mac. Native multi-token-prediction speculative decoding on Apple Silicon, two to three times the speed of plain decoding, exact at any temperature.

Qwen 3.8 Flash-Next Optimized Speed

Dynamic 4-bit quant with 8-bit attention. Higher quality and slightly slower. Recommended.

Qwen's 125B-A6B Flash-Next preview, the Qwen4-generation architecture with GDN hybrid MoE, Qwen Sparse Attention, and the 51B-parameter n-gram memory, running natively on MTPLX from day 0, with its native multi-token-prediction head drafting through MTPLX's speculative path. This is the recommended build: the Qwen Sparse Attention projections are kept at 8-bit, so the attention pathway that steers long contexts keeps its precision. For the absolute fastest build, pick Bare Speed.

The 32 GB n-gram embedding table streams from SSD by default, so the model fits a 96 GB+ Apple Silicon Mac with headroom, the weights stay resident, the table does not have to.

Measured on MTPLX 2.11.3 (16 September 2026)

This is the Qwen3.8-Flash-Next MLX pack for MTPLX, the fastest way to run Qwen 3.8 Flash Next on a Mac. MacBook Pro M5 Max with 128 GB, fans verified at maximum, sampled at the model's own settings (temperature 1.0, top-p 0.95, top-k 20). Conditions and sources for every row: mtplx.com/benchmarks.

Run tok/s
One OpenCode request on the Optimized Speed pack: 1,301 tokens generated, 18,539-token prompt with 18,364 tokens served from cache, MTP depth 3 125.8
9k-token code prompt, 1,500 tokens generated, seeded sampler, thinking off (62.5 on MTPLX 2.11.2) 79.3
109k-token OpenCode turn, mean of two runs 61.8
200k-token OpenCode turn, warm, mean of two runs 50.3
Full 45k to 56k-token generations (Flappy Bird at effort xhigh), whole turn 66.8

Exactness on this release: a thousand four-token draws from the fast path match a thousand from the plain path within the plain path's own noise at every joint length, at temperature 1, top-p 0.95, top-k 20. A 96,760-token conversation restored from the session cache in 8 ms. 261,120-token prompts decode. Details: MTPLX 2.11.3 release notes.

By mlx-serve's own M5 Max table (their repository at v26.9.3, docs/mtp-acceptance-port.md), MTPLX exact decoding measured 102.0 tok/s on a short prompt and 90.9 tok/s at about 16K tokens against their exact default at 92.5 and 82.4. The comparison page: MTPLX vs mlx-serve.

Runs on Apple Silicon Macs with 96 GB of unified memory or more: MacBook Pro M4 Max and M5 Max with 128 GB, Mac Studio M3 Ultra, M4 Max and M5 Max. Guide: Run Qwen 3.8 Flash Next on a Mac.

Speeds

Measured on an M5 Max, fans verified at max, single stream, real server (mtplx serve), official Qwen 3.8 sampling (temperature 1.0, top-p 0.95, top-k 20, sampled, not greedy).

Run tok/s
Coding task, MTP speculative decode (the default) 73.5
Same task, plain autoregressive 43.8

That is a 1.7x speculative multiplier through the product serve path, on sampled output that follows the model's own distribution.

How it is built

  • MoE experts and dense matrices at 4-bit with 64-weight groups; the Qwen Sparse Attention projections promoted to 8-bit (the quality edge over Bare Speed).
  • The GDN convolution and recurrent-state parameters, every norm, the QSA indexer, and the MTP head stay 16-bit.
  • The n-gram embedding table ships as a separate ngram-table.safetensors sidecar that MTPLX streams from SSD (resident is opt-in on very large machines). The vision tower is preserved in the weights.
Download 115.1 GB (includes the 32 GB n-gram table)
Resident weights (n-gram on SSD) ~83 GB + working set
Recommended Macs 96 GB+ unified memory
Context window 262,144 tokens
MTP depth adaptive, ceiling 3
Sampling temperature 1.0, top-p 0.95, top-k 20 (the official Qwen 3.8 contract)

The serving contract ships inside mtplx_runtime.json. MTPLX reads it on load. Drafts are accepted with the probability-ratio rule plus residual resampling, so the output follows the model's own distribution at any temperature.

Use it

Mac app: download at mtplx.com, pick "Qwen 3.8 Flash-Next Optimized Speed".

Command line:

pip install mtplx
mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed

Sibling: Bare Speed (flat 4-bit, the quickest build).

Base model: Qwen/Qwen3.8-Flash-Next (Qwen Community License; the upstream model card is preserved in this repo as README-upstream-qwen.md).

Downloads last month
15,696
Safetensors
Model size
126B params
Tensor type
U32
·
BF16
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed

Quantized
(266)
this model