Instructions to use ahnafnafee/Warlock-GLM-5.3-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ahnafnafee/Warlock-GLM-5.3-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS # Run inference directly in the terminal: llama cli -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS # Run inference directly in the terminal: ./llama-cli -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Use Docker
docker model run hf.co/ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
- LM Studio
- Jan
- Ollama
How to use ahnafnafee/Warlock-GLM-5.3-GGUF with Ollama:
ollama run hf.co/ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
- Unsloth Desktop
- Pi
How to use ahnafnafee/Warlock-GLM-5.3-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ahnafnafee/Warlock-GLM-5.3-GGUF with Docker Model Runner:
docker model run hf.co/ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
- Lemonade
How to use ahnafnafee/Warlock-GLM-5.3-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Run and chat with the model
lemonade run user.Warlock-GLM-5.3-GGUF-IQ4_XS
List all available models
lemonade list
- Hermes Agent
How to use ahnafnafee/Warlock-GLM-5.3-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ahnafnafee/Warlock-GLM-5.3-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ahnafnafee/Warlock-GLM-5.3-GGUF:IQ4_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Request access to Warlock GGUF
This model has had its refusal behavior removed. Requests are reviewed manually.
By requesting access you agree to use this model only for lawful research and evaluation, and to follow the GLM-5.3 and Warlock license terms.
Log in or Sign Up to review the conditions and access this model content.
Warlock GGUF
GGUF quantizations of audnai/penclaw-GLM-5.3-abliterated ("Warlock"), an abliterated
version of zai-org/GLM-5.3. Converted with llama.cpp from the BF16
weights at revision 027ae98a6f9dc37093b26759d7b8dbbfed9a97f8.
What was changed from GLM-5.3
It's a direct weight edit with no fine-tuning. Compared tensor by tensor against the official weights,
layers 2 to 49 had every matrix that writes into the residual stream (attention o_proj, the dense MLP
down_proj, and the shared and routed experts' down_proj) replaced with W - 3 r (r^T W), where r is
that layer's refusal direction. Every other tensor matches the official model exactly. Because the strength
is 3 rather than 1, the refusal component is reversed and doubled, not just removed. The model will answer
requests the original refuses, and may also skip safety caveats in harmless contexts.
Files
Each quant lives in its own folder, split into parts. Point llama.cpp at the first part and it loads the rest.
| Quant | Size | Files |
|---|---|---|
| IQ4_XS | 411 GB | IQ4_XS/warlock-4bit-IQ4XS-*.gguf |
| Q3_K-Q4_K | 369 GB | Q3_K-Q4_K/warlock-allgpu-Q3K-Q4K-*.gguf |
| Q8_0 | 801 GB | Q8_0/warlock-Q8_0-*.gguf |
Usage
After your access request is approved, set HF_TOKEN to a token from your account, then:
hf download ahnafnafee/Warlock-GLM-5.3-GGUF --include "Q3_K-Q4_K/*.gguf" --local-dir warlock-gguf
llama-server -m warlock-gguf/Q3_K-Q4_K/warlock-allgpu-Q3K-Q4K-00001-of-*.gguf --jinja -c 65536
For another build, download its folder from the table and point llama.cpp at its first part.
Each quant needs roughly its file size in combined GPU and system memory. With llama.cpp's default
--fit, whatever doesn't fit on the GPUs runs from RAM. This is a
thinking model: give it a generation budget of 16k tokens or more, or answers get cut off mid-reasoning.
License
This repo inherits the license terms of GLM-5.3 and of Warlock. Read both before use.
- Downloads last month
- 1
4-bit
8-bit
Model tree for ahnafnafee/Warlock-GLM-5.3-GGUF
Base model
audnai/penclaw-GLM-5.3-abliterated