GLM-4.7-Flash-Hauhau-Aggressive-W6A8

Native compressed-tensors W6A8 checkpoint published by Xananthium. Credit and upstream terms: HauhauCS/GLM-4.7-Flash-Uncensored-HauhauCS-Aggressive. The source model card is preserved in README_ORIGINAL.md. Quantization provenance and retained high-precision components are recorded in conversion-receipt.json.

Six bits on a 3090? Hell yes. >:)

I wanted W6A8, so that is what we built: six-bit stored weights, dynamic eight-bit activations, and vLLM doing the serving. Humming packs the INT6 values across INT32 word boundaries and unpacks them for the 3090's supported INT8 matrix math. The cards do not need a native six-bit Tensor Core instruction. The extra unpacking and scale work is part of the cost, and the speed tests include it.

The six-bit weights use symmetric groups of 128. Embeddings, routers, norms, vision, and designated sensitive components keep higher precision. BF16 remains the surrounding model dtype; eligible quantized linears actually use INT8 activations, with silent activation fallback disabled. TurboQuant Q4 is a separate KV-cache choice when the architecture supports it.

Accuracy matters too. A smaller checkpoint gets us memory back, but the model still has to follow the damn instructions, call tools, and get the security answers right. That is why the results include the knowledge quiz, security-policy checks, long-context recall, real Claude Code tool use, and concurrency. Failed checks stay in the report.

The measured numbers for this checkpoint are below. A six-bit speed gain over an eight-bit baseline is unmeasured unless a matched comparison report is attached.

Measured on two RTX 3090 24 GiB cards, vLLM 0.31.0

Test Result
CyberMetric-80-v1, primary 2048-token output budget 76/80
CyberMetric-80 after capped-response retries at 4096 tokens 78/80
Synthetic security policy checks, first answer 23/43
Synthetic long-context replay 53/54
Synthetic exact-answer reasoning, three repeated samples 18/24
Reasoning after capped-response retries at 4096 tokens 18/24
Cases correct in every repeat 6/8
Ordinary workflow recovery after an initial tool failure 6/6
Extra attempts at a failed tool approach 4
Short-context decode 122.5335 tokens/s
Long-context decode 56.2375 tokens/s
Four simultaneous requests, aggregate 297.6825 tokens/s

Reports in results/ contain the exact per-model profile, budgets, formatting metrics, failed checks and client integration outcomes. Counting successful test programs does not mean all accuracy checks passed. Token throughput includes reasoning tokens; it is not visible-answer throughput. The policy assessment is synthetic, and lengthy reasoning can exhaust the output budget. Eight independent long prefixes were replayed; this was not a fresh eight-turn conversation. CyberMetric is a small public knowledge quiz that may overlap training data, and does not establish practical security competence or a broad model ranking. Questions and private reasoning traces are not redistributed.

vLLM loading

hf download Xananthium/GLM-4.7-Flash-Hauhau-Aggressive-W6A8 --local-dir ./checkpoint
pip install --no-deps ./checkpoint/runtime/humming-compile-compat
LD_LIBRARY_PATH="$VIRTUAL_ENV"/lib/python3.12/site-packages/nvidia/cu13/lib \
LOCAL_HUMMING_COMPILE_COMPAT=1 \
VLLM_HUMMING_INPUT_QUANT_CONFIG='{"dtype":"int8","group_size":0,"allow_fallback":false}' \
vllm serve ./checkpoint --quantization humming --linear-backend humming --moe-backend humming --tensor-parallel-size 2 --dtype bfloat16 --trust-remote-code --enable-auto-tool-choice --tool-call-parser glm47 --reasoning-parser glm47 --kv-cache-dtype bfloat16 --max-model-len 147456

The loading recipe requires the architecture and quantization support in vLLM 0.31.0. Consult results/serving-profile.json for tested sampling and memory settings. Model-side INT8 activations apply only to eligible quantized linears; embeddings, routers, norms and designated sensitive components retain higher precision. KV cache precision is a separate setting. GLM MLA does not use TurboQuant in this version. Retained MTP weights do not imply speculative decoding was tested.

The complete comparison: The 3090 Plan

Seven contenders. Two RTX 3090s. The Council kept the failed checks in the record. The full report, charts, and evaluation scripts compare cybersecurity knowledge, reasoning, consistency, recovery after failed tools, long context, and concurrency. A Hugging Face mirror contains the same report.

The owner selected Nemotron Heretic W8A8 for deployment after reviewing the results. The earlier automated quality ranking remains visible and recommended REDCELL; the deployment choice does not change any measured score.

This checkpoint's complete final-round record is in results/full-round-evaluation.json. The public report explains the output-budget retries, architecture-specific cache settings, and limits of the small local comparison. No Kali workflows were run.

Downloads last month
29
Safetensors
Model size
30B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Xananthium/GLM-4.7-Flash-Hauhau-Aggressive-W6A8

Quantized
(2)
this model