How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="prithivMLmods/LightOnOCR-3-4B-GGUF")
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("prithivMLmods/LightOnOCR-3-4B-GGUF", device_map="auto")
Quick Links

LightOnOCR-3-4B-GGUF

LightOnOCR-3-4B, developed by lightonai, is the largest and most accurate model in the LightOnOCR-3 family of end-to-end OCR vision-language models, built on the Qwen3.5-4B architecture and released under the Apache 2.0 license. Recommended for most OCR tasks, it serves as a powerful, single-model alternative to complex document understanding pipelines by combining high-quality text transcription (triggered by an empty prompt) with advanced visual understanding features. When used with the grounding prompt, the model outputs labeled bounding boxes for all document elements, generates short descriptions for images, and extracts numerical data from charts into structured HTML tables. Highly versatile and optimized for seamless integration with Transformers and vLLM, it excels at processing complex layouts—including tables, receipts, forms, multi-column documents, and math notation—and performs best when documents are preprocessed at 400 DPI with a target longest dimension of 2048px.

  • Speculative Decoding — No

Model Files

File Name Quant Type File Size File Link Description
LightOnOCR-3-4B.BF16.gguf BF16 8.42 GB Link Full BF16 weights. Highest quality, largest file size.
LightOnOCR-3-4B.Q3_K_L.gguf Q3_K_L 2.42 GB Link Lower quality but usable, good for low RAM availability.
LightOnOCR-3-4B.Q3_K_M.gguf Q3_K_M 2.26 GB Link Low quality.
LightOnOCR-3-4B.Q4_K_M.gguf Q4_K_M 2.71 GB Link Good quality, default size for most use cases, recommended.
LightOnOCR-3-4B.Q4_K_S.gguf Q4_K_S 2.56 GB Link Slightly lower quality with more space savings, recommended.
LightOnOCR-3-4B.Q5_K_M.gguf Q5_K_M 3.07 GB Link High quality, recommended.
LightOnOCR-3-4B.Q5_K_S.gguf Q5_K_S 2.99 GB Link High quality, recommended.
LightOnOCR-3-4B.Q6_K.gguf Q6_K 3.46 GB Link Very high quality, near perfect, recommended.
LightOnOCR-3-4B.Q8_0.gguf Q8_0 4.48 GB Link Extremely high quality, generally unneeded but max available quant.
LightOnOCR-3-4B.mmproj-bf16.gguf mmproj-bf16 676 MB Link Multimodal projection file in BF16 format. Used for vision/language models.
LightOnOCR-3-4B.mmproj-q8_0.gguf mmproj-q8_0 367 MB Link Multimodal projection file in Q8_0 quantization. Smaller size for vision capabilities.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
776
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/LightOnOCR-3-4B-GGUF

Quantized
(11)
this model

Collections including prithivMLmods/LightOnOCR-3-4B-GGUF