jina-ocr-v1 โ GGUF for CrispEmbed
GGUF conversion of jinaai/jina-ocr-v1 for CrispEmbed: a 3.4B mixture-of-experts OCR model (~570M active parameters). It uses DeepSeek-OCR's architecture: SAM ViT-B + CLIP-L/14 vision encoder and a DeepSeek-V2 MoE decoder.
Licence: CC BY-NC 4.0 โ non-commercial use only. ยฉ Jina AI. For commercial use, contact Jina AI. These files are a format conversion of the original weights and carry the same licence and attribution requirements.
| file | size |
|---|---|
jina-ocr-v1-f16.gguf |
6.7 GB |
jina-ocr-v1-q8_0.gguf |
3.6 GB |
jina-ocr-v1-q4_k.gguf |
2.3 GB |
crispembed -m jina-ocr-v1 --ocr page.png
The GGUF stores the model's own prompt ("Transcribe the provided document image into a clean
Markdown format, preserving the natural reading order.", in its <|User|> / <|Assistant|> chat
layout), its rope base (1e6), and plain greedy decoding, so no extra flags are needed.
Parity
Checked against transformers (trust_remote_code, the model's own processor) with
tools/ci-heavy/jina_ocr_v1.py,
on the same 1024 ร 1024 global view (upstream's crop_mode=False, base_size=1024). The text was
identical on both test fixtures, a single line and a full book page, including when upstream ran
in float32 on CrispEmbed's own preprocessed pixels.
CrispEmbed encodes one global view. Upstream's default "Gundam" mode also adds up to nine 640 ร 640 tiles, which helps on dense, small print; that mode is not implemented here.
- Downloads last month
- 136
8-bit
16-bit
Model tree for cstr/jina-ocr-v1-GGUF
Base model
jinaai/jina-ocr-v1