jina-ocr-v1 โ€” GGUF for CrispEmbed

GGUF conversion of jinaai/jina-ocr-v1 for CrispEmbed: a 3.4B mixture-of-experts OCR model (~570M active parameters). It uses DeepSeek-OCR's architecture: SAM ViT-B + CLIP-L/14 vision encoder and a DeepSeek-V2 MoE decoder.

Licence: CC BY-NC 4.0 โ€” non-commercial use only. ยฉ Jina AI. For commercial use, contact Jina AI. These files are a format conversion of the original weights and carry the same licence and attribution requirements.

file size
jina-ocr-v1-f16.gguf 6.7 GB
jina-ocr-v1-q8_0.gguf 3.6 GB
jina-ocr-v1-q4_k.gguf 2.3 GB
crispembed -m jina-ocr-v1 --ocr page.png

The GGUF stores the model's own prompt ("Transcribe the provided document image into a clean Markdown format, preserving the natural reading order.", in its <|User|> / <|Assistant|> chat layout), its rope base (1e6), and plain greedy decoding, so no extra flags are needed.

Parity

Checked against transformers (trust_remote_code, the model's own processor) with tools/ci-heavy/jina_ocr_v1.py, on the same 1024 ร— 1024 global view (upstream's crop_mode=False, base_size=1024). The text was identical on both test fixtures, a single line and a full book page, including when upstream ran in float32 on CrispEmbed's own preprocessed pixels.

CrispEmbed encodes one global view. Upstream's default "Gundam" mode also adds up to nine 640 ร— 640 tiles, which helps on dense, small print; that mode is not implemented here.

Downloads last month
136
GGUF
Model size
3B params
Architecture
unlimited_ocr
Hardware compatibility
Log In to add your hardware

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cstr/jina-ocr-v1-GGUF

Quantized
(3)
this model