Gemmable 4
Collection
2 items β’ Updated
How to use Mia-AiLab/Gemmable-4-12B-MTP-GGUF with llama.cpp:
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M
docker model run hf.co/Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M
How to use Mia-AiLab/Gemmable-4-12B-MTP-GGUF with Ollama:
ollama run hf.co/Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M
How to use Mia-AiLab/Gemmable-4-12B-MTP-GGUF with Docker Model Runner:
docker model run hf.co/Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M
How to use Mia-AiLab/Gemmable-4-12B-MTP-GGUF with Lemonade:
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Mia-AiLab/Gemmable-4-12B-MTP-GGUF:Q4_K_M
lemonade run user.Gemmable-4-12B-MTP-GGUF-Q4_K_M
lemonade list
Gemmable 4 12B is a GGUF export of Gemma 4 12B fine-tuned on Fable-5 style reasoning and assistant traces.
google/gemma-4-12BStandard load:
llama-server -m "gemmable-4-12b-fp16.gguf"
Speculative / draft-MTP load:
llama-server -m "gemmable-4-12b-Q4_K_M.gguf" \
--spec-draft-model "gemmable-4-12b-Q4_K_M-mtp.gguf" \
--spec-type draft-mtp \
--spec-draft-n-max 4
Use the matching fp16 or quantized main file with its -mtp companion.
(Requires LM Studio with am17an's PR merged or custom llama.cpp runtime. As of 2026-05, mainline LM Studio runtime doesn't yet have draft-mtp for Gemma-4 β track upstream merge.)
gemmable-4-12b-fp16.gguf is the standard fp16 main model.gemmable-4-12b-fp16-mtp.gguf is the matching fp16 assistant / draft file.gemmable-4-12b-Q4_K_M.gguf and gemmable-4-12b-Q4_K_M-mtp.gguf.Gemmable = Gemma + Fable-style tuning.