Request access to Warlock GGUF

This model has had its refusal behavior removed. Requests are reviewed manually.

By requesting access you agree to use this model only for lawful research and evaluation, and to follow the GLM-5.3 and Warlock license terms.

Log in or Sign Up to review the conditions and access this model content.

Warlock GGUF

GGUF quantizations of audnai/penclaw-GLM-5.3-abliterated ("Warlock"), an abliterated version of zai-org/GLM-5.3. Converted with llama.cpp from the BF16 weights at revision 027ae98a6f9dc37093b26759d7b8dbbfed9a97f8.

What was changed from GLM-5.3

It's a direct weight edit with no fine-tuning. Compared tensor by tensor against the official weights, layers 2 to 49 had every matrix that writes into the residual stream (attention o_proj, the dense MLP down_proj, and the shared and routed experts' down_proj) replaced with W - 3 r (r^T W), where r is that layer's refusal direction. Every other tensor matches the official model exactly. Because the strength is 3 rather than 1, the refusal component is reversed and doubled, not just removed. The model will answer requests the original refuses, and may also skip safety caveats in harmless contexts.

Files

Each quant lives in its own folder, split into parts. Point llama.cpp at the first part and it loads the rest.

Quant Size Files
IQ4_XS 411 GB IQ4_XS/warlock-4bit-IQ4XS-*.gguf
Q3_K-Q4_K 369 GB Q3_K-Q4_K/warlock-allgpu-Q3K-Q4K-*.gguf
Q8_0 801 GB Q8_0/warlock-Q8_0-*.gguf

Usage

After your access request is approved, set HF_TOKEN to a token from your account, then:

hf download ahnafnafee/Warlock-GLM-5.3-GGUF --include "Q3_K-Q4_K/*.gguf" --local-dir warlock-gguf
llama-server -m warlock-gguf/Q3_K-Q4_K/warlock-allgpu-Q3K-Q4K-00001-of-*.gguf --jinja -c 65536

For another build, download its folder from the table and point llama.cpp at its first part. Each quant needs roughly its file size in combined GPU and system memory. With llama.cpp's default --fit, whatever doesn't fit on the GPUs runs from RAM. This is a thinking model: give it a generation budget of 16k tokens or more, or answers get cut off mid-reasoning.

License

This repo inherits the license terms of GLM-5.3 and of Warlock. Read both before use.

Downloads last month
1
GGUF
Model size
753B params
Architecture
glm-dsa
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ahnafnafee/Warlock-GLM-5.3-GGUF

Quantized
(1)
this model