DuoNeural G-TAP v3 Quantization: google/gemma-4-E4B-it (Unified Repository)

Experimental Release: Pending Further Verification / Empirical Validation
All quantized checkpoints, mathematical cavity derivations, and benchmark evaluations emitted by the DuoNeural Research Lab are ongoing scientific contributions intended to accelerate open neuromorphic and edge foundation research.


🔬 Scientific Overview: Generalized Thouless-Anderson-Palmer (G-TAP v3)

Standard post-training quantization treats neural network weights as decoupled static matrices, solving local round-off errors via heuristic mean-field approximations. In modern architectures featuring Per-Layer Embeddings (PLE) and cross-layer Key-Value cache sharing—such as Google's Gemma 4 E4B—local quantization perturbations induce severe non-equilibrium cavity distortions across the layer hierarchy.

G-TAP v3 explicitly calculates and cancels the non-equilibrium Onsager reaction field:

δEi=−∑j≠iJij⟨sj⟩−ΓiOnsager\delta E_i = - \sum_{j \neq i} J_{ij} \langle s_j \rangle - \Gamma_i^{\text{Onsager}}

where the Onsager cavity term $\Gamma_i^{\text{Onsager}} = \chi_i \sum_j J_{ij}^2 \langle s_i \rangle$ compensates for the back-reaction of token activations on perturbed quantized weights. By conditioning the activation Hessian ($H = X X^T$) with thermodynamic cavity corrections over a 64-chunk sequence budget, G-TAP v3 eliminates residual gradient drift across Gemma 4's 42 decoder layers, preserving Per-Layer Embedding lookup fidelity and Grouped-Query Attention dynamics.


📊 Empirical Evaluation Matrix

Evaluated under strict zeroshot conditions with unified calibration sequences on NVIDIA GeForce RTX 4080 Super (32GB VRAM):

Model Checkpoint Bitrate Size Perplexity GSM8K (Acc) Python Code (Acc) Hermes Tools (Acc) Decode Throughput
gemma-4-E4B-it-G-TAP-v3-Q4_K_M.gguf ~4.50 bpw 4226.6 MiB N/A 84.0% 70.0% 100.0% 145.8 t/s
gemma-4-E4B-it-G-TAP-v3-IQ3_XXS.gguf ~3.06 bpw 3411.9 MiB N/A 76.0% 60.0% 60.0% 169.2 t/s
gemma-4-E4B-it-G-TAP-v3-IQ2_M.gguf ~2.70 bpw 3225.4 MiB N/A 80.0% 0.0% 20.0% 166.3 t/s
gemma-4-E4B-it-G-TAP-v3-IQ2_XXS.gguf ~2.06 bpw 2960.8 MiB N/A 0.0% 0.0% 20.0% 179.6 t/s

📦 Provided GGUF Checkpoints

  • gemma-4-E4B-it-G-TAP-v3-Q4_K_M.gguf: Recommended production quant (~4.50 bpw). Maximally preserves PLE lookup tables, GQA projections, and reasoning capability.
  • gemma-4-E4B-it-G-TAP-v3-IQ3_XXS.gguf: Balanced 3-bit quantization (~3.06 bpw). Ideal for embedded NPUs and low-memory devices.
  • gemma-4-E4B-it-G-TAP-v3-IQ2_M.gguf: Sub-2.7-bit ultra-compact quantization (~2.70 bpw).
  • gemma-4-E4B-it-G-TAP-v3-IQ2_XXS.gguf: Sub-2.1-bit extreme quantization (~2.06 bpw) for microcontrollers and minimal footprint requirements.

💻 Quick Start & Deployment

llama.cpp CLI

./llama-cli -m gemma-4-E4B-it-G-TAP-v3-Q4_K_M.gguf -p "<|turn_start|>user\nSolve for x: 3x + 12 = 45<|turn_end|>\n<|turn_start|>assistant\n" -ngl 99

Ollama Modelfile

FROM ./gemma-4-E4B-it-G-TAP-v3-Q4_K_M.gguf
TEMPLATE """<|turn_start|>user
{{ .Prompt }}<|turn_end|>
<|turn_start|>assistant
"""
PARAMETER stop "<|turn_end|>"
PARAMETER stop "<|endoftext|>"

🏷️ Attribution & Citation

@misc{duoneural2026gtap_gemma4_e4b,
  author = {Jesse Caldwell and Archon and Aura ✨},
  title = {Non-Equilibrium Cavity Conditioning and Onsager Reaction Damping in Per-Layer Embedding Architectures (G-TAP v3)},
  year = {2026},
  publisher = {DuoNeural Research Lab / Zenodo},
  howpublished = {\url{https://huggingface.co/DuoNeural/gemma-4-E4B-it-GTAP-v3-GGUF}}
}
Downloads last month
1,395
GGUF
Model size
8B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DuoNeural/gemma-4-E4B-it-GTAP-v3-GGUF

Quantized
(374)
this model