Qwopus3.6-27B-Fusion — NVFP4 + FP8
Mixed-precision quantization · BF16 MTP drafter preserved · text-only
A mixed-precision NVFP4 + FP8 quantization of
KyleHessling1/Qwopus3.6-27B-Fusion-BF16,
built to run on two consumer 16 GB Blackwell cards with the BF16 MTP (multi-token-prediction)
drafter preserved so speculative decoding still works.
~23.4 GiB of weights. Serves at 135,168-token context under tensor-parallel-size 2 on
2× RTX 5060 Ti 16 GB. Do note that uncached graphs may cause OOM on first load, but should recover on bounce back.
Gated on real hardware, 2026-08-01. Every number in Measured Results below was measured on this checkpoint under TP=2 — none are inherited from the parent or from a sibling quant. The blocking NVFP4 TP=2 coherence probe passed.
What this is
| Format | compressed-tensors, format: mixed-precision (nvfp4-pack-quantized + float-quantized) |
| Weights | 243 layers FP8 · 157 layers NVFP4 · 106 tensors BF16 |
| Activations | Also quantized — NVFP4 layers are W4A4, FP8 layers are W8A8. This is not a weight-only quant (see note below). |
| Method | RTN (round-to-nearest) with calibrated static scales — not GPTQ/AWQ/AutoRound-tuned |
| Writer | llm-compressor 0.12.0, single oneshot() pass |
| Calibration | lemon07r/pile-calibration-v5, 512 samples × 2048 tokens |
| MTP module | BF16, preserved — grafted back post-oneshot (see Reproducing below) |
| Vision tower | Absent by design — text-only; serve with --language-model-only |
| Size | ~23.4 GiB |
What's quantized where, and why
The per-layer split is not uniform — it comes from a sensitivity scan of every decoder Linear
measured at each candidate scheme, then a budget knapsack against a hard 16 GB-per-card ceiling:
| Group | Precision | Layers |
|---|---|---|
mlp.gate_proj / up_proj |
NVFP4 (W4A4, group 16) | 0–32 |
mlp.gate_proj / up_proj |
FP8 (channel W / dynamic-token A) | 33–63 |
mlp.down_proj |
NVFP4 | 0–52 |
mlp.down_proj |
FP8 | 53–63 |
linear_attn.out_proj (DeltaNet) |
NVFP4 | front blocks |
linear_attn.out_proj |
FP8 | back blocks |
linear_attn.in_proj_qkv / in_proj_z |
FP8 | all |
self_attn.q/k/v/o_proj (full-attn layers) |
FP8 | 3, 7, 11 … 63 (16 layers) |
self_attn.* on layer 0 |
BF16 | 0 — in no quantization group |
lm_head, embed_tokens, all mtp.*, DeltaNet in_proj_a/b, conv1d |
BF16 | — |
Three choices are deliberate and worth calling out:
- All 16
o_projare FP8, not NVFP4. The automated ranking, which scores layers by local reconstruction error (MSE against the BF16 layer output), puto_projon layers 3–47 in the NVFP4 group. That was overridden by hand, following the earlier INT-family recipes on this architecture whose map was derived with a KL-Fisher criterion and kept these projections at higher precision. The two criteria disagree here; the override sides with the one that had already been validated by serving. Cost: ~0.16 GiB by the recipe's own estimate. - Layer 0's attention block is left entirely in BF16. It is the first full-attention layer and falls outside every group's layer list — an artifact of the recipe's front-boundary handling that was kept because it is cheap (one block) and the resulting map is the one that passed gates.
- The DeltaNet recurrent path (
in_proj_a/in_proj_b) is never quantized. State-space parameters are numerically unforgiving; they stay BF16 in every recipe on this family.
Activations are quantized, not just weights. NVFP4 layers run W4A4 (4-bit weights and 4-bit activations, per-group scales computed dynamically at runtime with a static per-tensor global scale); FP8 layers run W8A8 (static per-channel weight scales, dynamic per-token activation scales). This is what makes the checkpoint fast on Blackwell FP4/FP8 tensor cores rather than merely small — but it also means quality depends on activations resembling the calibration distribution. If you need robustness on far-out-of-distribution inputs, a weight-only quant (
W4A16/NVFP4A16) is the safer family, at the cost of the tensor-core speedup.
Measured results
Measured on 2× RTX 5060 Ti 16 GB, TP=2, vLLM 0.23.1rc1.dev861, 2026-08-01.
| Gate | Result |
|---|---|
| TP=2 generation coherence (blocking) | PASS — coherent, zero !-runs, zero NaN. The generated nth_prime function was executed and verified correct (n=1..10, and n=100 → 541) |
| Deterministic accuracy set | 6/6 exact (arithmetic, rate, code-trace, logic, percentage, GSM-style), all finish=stop |
| Tool calling | 3/3 phases pass — single-turn call, multi-turn tool-result replay, repetition canary. 8/8 calls parsed at temp 0.7 and temp 0.2 |
| Long-context needle | EXACT at 129,473 prompt tokens (40 % depth) in 129 s |
| MTP acceptance (K=3) | 66.4 % (5171/7791 draft tokens); per-position 80.5 / 64.7 / 53.9 %; acceptance length 2.99 |
| Single-stream throughput | 56.2 tok/s |
| 5-stream aggregate | 113.0 tok/s, all five coherent |
| KV pool at boot | 177,408 tokens, 1.31× concurrency at 135,168 max-len |
| OOM / NaN | 0 / 0 |
For scale, the same recipe applied to a sibling 27B finetune on the same box measures 47.5 tok/s single-stream, 85.3 tok/s at 5 streams, and an exact needle at 128,978 in 140 s. This checkpoint is faster on every one of those, but they are different models — read it as "no regression from quantization", not as a quality comparison.
Benchmarks are deliberately absent. No HumanEval/MBPP/GSM8K numbers are reproduced here.
The upstream card's scores were measured on a different artifact (a Q4_K_M GGUF), and quoting
them for these weights would be a claim nobody has tested. If you benchmark this quant, please open
a discussion and the numbers will be added with credit.
Serving it (vLLM)
Built for and tested on 2× RTX 5060 Ti 16 GB (Blackwell, sm_120), PCIe, no NVLink, CUDA 13.
vllm serve /path/to/Qwopus3.6-27B-Fusion-NVFP4-FP8 \
--served-model-name local-model \
--tensor-parallel-size 2 \
--max-model-len 135168 \
--gpu-memory-utilization 0.92 \
--dtype float16 \
--disable-custom-all-reduce \
--max-num-seqs 5 \
--max-num-batched-tokens 4128 \
--kv-cache-dtype turboquant_3bit_nc \
--reasoning-parser qwen3 \
--tool-call-parser hermes --enable-auto-tool-choice \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--default-chat-template-kwargs '{"preserve_thinking":true}' \
--compilation-config '{"cudagraph_mode":"PIECEWISE","cudagraph_capture_sizes":[1,2,4,8,12,16,20]}' \
--language-model-only --trust-remote-code --no-enable-flashinfer-autotune
Flags that are load-bearing, not taste:
--language-model-only— required. The vision tower is absent, but without this flag vLLM still treats the checkpoint as multimodal and provisions a phantom encoder cache, costing enough load-time VRAM to OOM the BF16 MTP drafter.config.json'slanguage_model_onlyis not read; only the CLI flag works.--dtype float16— the TurboQuant attention backend is fp16-only.--max-num-batched-tokens 4128— the Mamba/DeltaNet block-size minimum on this family.cudagraph_mode: PIECEWISE— a full decode cudagraph corrupts the MTP verify path (produces1111…garbage). Piecewise is fine.--enforce-eageris not needed.--kv-cache-dtype turboquant_3bit_nc— optional but recommended; it is what buys the 135K context. Drop tofp8_e4m3if your vLLM build lacks TurboQuant, and lower--max-model-len.--tool-call-parser hermes— required, and not the one you'd copy from a Qwen3.6 profile. See the chat-template section immediately below; the wrong parser fails silently.
⚠ The chat template here differs from the source — read this if you use tools
This repo ships a one-character fix to chat_template.jinja. The original is preserved
alongside it as chat_template.jinja.orig. Nothing else about the template was touched.
The source template instructs the model to emit tool calls in this shape:
{"name": <function-name>, "arguments": <args-json-object>} ← note: no quotes
…while the branch that renders previously-issued tool calls quotes the name correctly
({"name": " + name + "). The model does what the spec in its system prompt shows, so on a
first-turn tool call it emits invalid JSON:
{"name": search_repo}, "arguments": {"repo": "vllm-project/vllm", ...}}
Unquoted name, stray brace — no parser accepts it. On follow-up turns, where a correctly rendered example is already in the conversation, the model imitates that and emits valid JSON. That asymmetry is exactly what we measured.
The fix is to quote the placeholder in the spec line:
{"name": "<function-name>", "arguments": <args-json-object>}
Measured effect (8 seeds per cell, single-turn tool call, --tool-call-parser hermes):
| temp 0.2 | temp 0.7 | |
|---|---|---|
| source template | 1/8 (12 %) | 3/8 (37 %) |
| fixed template | 8/8 (100 %) | 8/8 (100 %) |
Lineage
Qwen/Qwen3.6-27B (base architecture + license)
└── Unsloth/Qwen3.6-27B (repackaged base; MTP module fix)
├── Jackrong/Qwopus3.6-27B-v2 (reasoning finetune) ──┐
└── Jackrong/Qwopus3.6-27B-Coder (code finetune, │ ← descends from v2
itself downstream of v2) │
▼
KyleHessling1/Qwopus3.6-27B-Fusion-BF16
W(L) = v2 + α(L)·(Coder − v2), α: 0.12 → 0.48 by depth
│
▼
this repo — NVFP4 + FP8, MTP preserved, text-only
Acknowledgments
None of the ideas in this quant are mine. The recipe is assembled from techniques other people worked out and published — I ran a pipeline, measured the result, and wrote down what happened. The section below is not a formality; these are the people whose public work this is built from.
People to follow
The Kaitchup — quants, and the Substack that explains why a method works instead of just shipping a file. The AutoRound write-ups in particular are where a lot of this thinking started. Probably the single best place to learn practical quantization. Go follow.
rdtand — PrismaAURA. Mixed-precision-per-layer quants that proved out the idea this whole recipe rests on: not every layer deserves the same number of bits. The predecessor experiments on this box were PrismaAURA-informed.
Minachist — the front/back higher-precision idea that this recipe's central structure is lifted from: protect the early and late blocks, spend the cheap bits in the middle. Their mixed int4/int8 AutoRound models are also what demonstrated that vLLM will serve a per-layer mixed-width checkpoint at all — without that existence proof there was no reason to build one. Every boundary in the precision table above is a descendant of this.
sakamakismile — the NVFP4 + text-only idea.
Qwen3.6-27B-Text-NVFP4-MTPis what proved a vision-strippedQwen3_5ForConditionalGenerationwill actually load in vLLM. This checkpoint is text-only because they showed it was possible.unsloth — NVFP4 quants and a steady stream of ideas, plus the repackaged base and MTP fix that this model's own lineage is built on. Their shipped
Qwen3.6-27B-NVFP4was the reference for what a correct config for this family looks like.cpatonn — AWQ/INT4 quants across this exact model family (including both Qwopus parents). A reference point for what good quants of these models look like.
Lorbus — AutoRound INT4 work that was the daily driver here for a long stretch and the baseline later recipes had to beat.
crushleorey — NVFP4 of
Qwopus3.6-27B-v2, the reasoning parent.Jackrong — built both Qwopus parent finetunes.
Qwen team, Alibaba Cloud —
Qwen3.6-27B: the architecture (theqwen35hybrid of DeltaNet linear attention with periodic full attention), the weights everything downstream derives from, the MTP design, and the license this release inherits. Nothing here exists without them.Jackrong — the two parent finetunes,
Qwopus3.6-27B-v2andQwopus3.6-27B-Coder(see above).KyleHessling1 — the Fusion merge: the layer-weighted delta-scale recipe, the measurement of parent lineage that motivated it, and the BF16 release that made this quant possible at all. His card also credits Grok 4.5 (xAI) for surfacing the parent→descendant lineage from weight geometry, and Fable for the earlier DUS and DARE-TIES approaches that were tried first.
unsloth — the repackaged base this lineage builds on; the checkpoint still carries
unsloth_fixed_mtp, i.e. their MTP-module fix rides along into these weights (see above).
The quantization stack
- vLLM project — three separate things, all essential:
llm-compressor(wrote this checkpoint),compressed-tensors(the mixed-precision format that can carry an NVFP4 group and an FP8 group in one file), and vLLM itself, which loads and serves it. - NVIDIA — the NVFP4 format, the Blackwell FP4 tensor cores that make W4A4 fast rather than merely small, and CUTLASS, whose FP4 GEMM kernels this checkpoint dispatches to.
- FlashInfer — the attention and FP4 kernel layer vLLM uses on this path.
- Intel — AutoRound — not a writer here, but the
measurement engine: the per-layer sensitivity scan that produced this recipe uses AutoRound's
own codec registry (
nv_fp4,fp8_sym,int_sym) atiters=0so the scan measures with the exact scale semantics a real quant would start from. - TurboQuant — the 3-bit KV cache scheme that buys the long context: Amir Zandieh, Insu Han, Majid Daliri, Amin Karbasi (Google Research / NYU, ICLR 2026), building on DRIVE and EDEN by Vargaftik et al.
- Dao-AILab (
causal-conv1d) and fla-org (flash-linear-attention) — the fast paths for the DeltaNet layers that make up most of this model. - PyTorch — everything above stands on it.
Data & infrastructure
- Hugging Face —
transformers,safetensors,huggingface_hub,datasets, and the Hub that hosts every model in the lineage diagram above. Thesafetensorsformat in particular is what makes surgery like the MTP graft below safe to do at all. - lemon07r — the calibration set used for the static scales, itself derived from Bartowski's v5 semantic mix (code + chat + docs + math, quality-filtered), which draws on EleutherAI's The Pile.
Tooling around it
- Claude (Anthropic) — design and implementation partner for that pipeline: the sensitivity scout, the physics-constrained recipe menu, the MTP-restore step, and the acceptance-gate discipline that decides whether an artifact like this is allowed to ship.
Reproducing this
Two non-obvious steps stand between "run the quantizer" and "a checkpoint vLLM will load".
1. llm-compressor loads this family through the text-only class. The saved artifact therefore
(a) drops the BF16 MTP module and the vision tower, (b) renames tensors
model.language_model.* → model.*, and (c) writes a bare text config that vLLM's Qwen3.5
integration rejects outright. A post-pass must rename the tensors back, graft mtp.* BF16 from the
source, and rebuild the full config.
2. The recipe's mtp ignore patterns are silently dropped. llm-compressor resolves the ignore
list to concrete module names of the model it saw — the text-only one, which has no MTP. If you
don't re-add re:^mtp\..* and re:.*\.mtp\..* to the saved quantization_config, vLLM builds the
drafter with quantized schemes and then chokes loading the grafted BF16 weights.
# The recipe shape (abridged) — one oneshot pass, two config groups + an ignore list
QuantizationModifier(
ignore=["lm_head", "re:.*embed_tokens.*", "re:^mtp\\..*", "re:.*\\.mtp\\..*",
"re:.*visual.*", "re:.*\\.linear_attn\\.in_proj_a$",
"re:.*\\.linear_attn\\.in_proj_b$", "re:.*conv1d.*"],
config_groups={
"group_nvfp4": {...}, # W4A4, type float, tensor_group, group_size 16, dynamic local
"group_fp8": {...}, # W8 channel static, A8 token dynamic
},
)
Limitations & known issues
- Text-only. The vision tower is not in this checkpoint. If you need the multimodal path, use the BF16 source. (The upstream card also notes vision is untested on the merged weights.)
- RTN, not tuned. Expect it to sit below a GPTQ/AWQ/AutoRound-tuned quant of the same size. The trade is that it takes ~45 minutes instead of ~15 hours, and it's format-native for vLLM.
- NVFP4 at TP=2 has a NaN failure history on consumer Blackwell. That's why the blocking coherence probe exists; see Measured Results.
- Inherited, not re-measured: the MTP
K=3setting and the sampling defaults come from a sibling artifact.Kis a serving flag — re-measure it for your workload. - Untested outside this hardware. Verified only on 2× RTX 5060 Ti under TP=2. Single-GPU (24 GB+) should work but has not been tried.
- Research preview, like its parent. No additional safety alignment was performed; quantization does not preserve safety behavior in any guaranteed way.
License
Inherited from the base model: usage is governed by the Qwen license, as with the BF16 source and both parent finetunes. No training data is redistributed here — this repo contains quantized weights derived from a publicly published checkpoint.
Disclaimer
Shared for educational and research purposes only, as is, with no warranty of any kind. I take no responsibility whatsoever for how it's used or for anything it produces — if you run it, the outcome is yours.
- Downloads last month
- 50
Model tree for tasticleeze/Qwopus3.6-27B-Fusion-NVFP4-FP8
Base model
Qwen/Qwen3.6-27B