Instructions to use Hukyl/kgb-archive-doclayout-yolov10m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- YOLOv10
How to use Hukyl/kgb-archive-doclayout-yolov10m with YOLOv10:
from ultralytics import YOLOvv10 model = YOLOvv10.from_pretrained("Hukyl/kgb-archive-doclayout-yolov10m") source = 'http://images.cocodataset.org/val2017/000000039769.jpg' model.predict(source=source, save=True) - Notebooks
- Google Colab
- Kaggle
Ukrainian Historical Document Layout Detector — DocLayout-YOLOv10m (fine-tuned)
A fine-tuned DocLayout-YOLO (YOLOv10m-doclayout) detector for layout-region detection on scanned Ukrainian-language historical document pages. Trained as part of a course project on document digitisation; intended as the detection stage of an end-to-end OCR pipeline (detection → preprocessing → printed/handwritten OCR).
The training corpus is private and is not redistributed with this model.
Classes (7)
| ID | Label | Description |
|---|---|---|
| 0 | printed_text |
Printed Cyrillic body text |
| 1 | handwritten_text |
Handwritten Cyrillic text (notes, annotations) |
| 2 | signature |
Hand-signed names |
| 3 | seal |
Stamped seals / round office stamps |
| 4 | table |
Tabular regions |
| 5 | date_printed |
Printed dates (form fields) |
| 6 | date_handwritten |
Handwritten dates |
Intended Use
- Layout-region proposal for downstream OCR on Ukrainian-language historical document scans.
- Feeding crops into:
- PaddleOCR-VL / PP-OCRv5 eslav for
printed_text/date_printed - TrOCR (Cyrillic, fine-tuned) for
handwritten_text/date_handwritten
- PaddleOCR-VL / PP-OCRv5 eslav for
- Region-of-interest extraction for downstream tasks: NER, semantic search, signature clustering.
Out-of-Scope / Not Recommended
- Modern documents, non-Cyrillic scripts, non-archival photography.
- Privacy-sensitive deployments without legal review — historical record corpora can contain personal data; consult the appropriate ethics / legal authority before any downstream use that surfaces such data.
- Production OCR for live documents — this is a research artefact, not an audited product.
Training Data
| Source | Private Ukrainian-language historical-document corpus (mid-20th-century scans) |
Annotated images (post prepare-data) |
~330 |
| Annotated bounding boxes | ~13k |
| Annotators | Multi-annotator effort across several CVAT batches |
| Train / val split | ~280 train / ~50 val (offline-augmented training set: ~1.7k images / ~61k boxes) |
Approximate class frequency in the full corpus (before split)
| Class | Boxes | Share |
|---|---|---|
printed_text |
~10,000 | ~75 % |
handwritten_text |
~2,500 | ~19 % |
date_printed |
~275 | ~2 % |
signature |
~250 | ~2 % |
date_handwritten |
~210 | ~2 % |
seal |
~45 | <1 % |
table |
~10 | <1 % |
The corpus is dominated by printed text. seal and table are still small (low tens of instances corpus-wide).
Training Procedure
This release corresponds to a single-split final training run (not cross-validation).
Base model
- Architecture: YOLOv10m-doclayout (DocLayout-YOLO), 19,970,122 params, 68.0 GFLOPs.
- Initialisation: weights transplanted from
juliozhao/DocLayout-YOLO-DocStructBench(DocStructBench, 10 classes) onto a fresh 7-class head. Backbone (model.0–model.10) frozen during fine-tuning.
Hyperparameters
| Optimizer | AdamW (auto-selected by Ultralytics) |
| lr0 | 0.000909 |
| Momentum | 0.9 |
| Weight decay | 0.0005 |
| Warmup | 3 epochs |
| Schedule | cosine off, close_mosaic=10 (no effect; mosaic disabled) |
| Mixed precision | fp16 (AMP) |
| Image size | 1024 |
| Batch size | 4 |
| Max epochs | 100 |
| Early-stopping patience | 25 epochs |
| Frozen layers | model.0 – model.10 (freeze=11) |
| Hardware | NVIDIA RTX 4070, 12 GB VRAM |
Augmentation
- Online (Ultralytics defaults): HSV (0.015/0.7/0.4), translate 0.1, scale 0.5, RandAugment, erasing 0.4. Mosaic / MixUp / CopyPaste disabled.
- Offline (training set only): crops, small rotations, asymmetric margins, paper-tone shifts, Gaussian noise, JPEG compression artefacts. Validation set untouched.
Run
- Wall-clock ≈ 49 min on a single RTX 4070.
- Early stopping fired at epoch 38 / 100; best fitness (
0.1·mAP@50 + 0.9·mAP@50-95) reached at epoch 13. - The released
best.ptcorresponds to epoch 13.
Evaluation
Metrics are computed on the held-out validation set of the single split (~50 images, ~2.1k boxes). The validation set received no augmentation.
Headline (best.pt re-validated)
| Metric | Value |
|---|---|
| mAP@50 | 0.637 |
| mAP@50-95 | 0.402 |
| Precision | 0.681 |
| Recall | 0.610 |
Per-class
| Class | n (val, approx.) | Precision | Recall | mAP@50 | mAP@50-95 |
|---|---|---|---|---|---|
printed_text |
~1,700 | 0.721 | 0.822 | 0.816 | 0.527 |
handwritten_text |
~300 | 0.641 | 0.598 | 0.604 | 0.327 |
signature |
~35 | 0.607 | 0.838 | 0.831 | 0.492 |
seal |
~5 | 1.000 | 0.394 | 0.575 | 0.291 |
table |
~2 | 1.000 | 0.934 | 0.995 | 0.812 |
date_printed |
~40 | 0.346 | 0.395 | 0.324 | 0.196 |
date_handwritten |
~35 | 0.450 | 0.289 | 0.314 | 0.170 |
Operating-point recommendation
- Reliable for downstream OCR:
printed_text(0.82 mAP@50) andhandwritten_text(0.60 mAP@50). - Usable but high variance on this val draw:
signature— single-split estimate; treat as upper bound. - Data-starved:
seal,table. Predictions are precise (P = 1.0 in both) but recall is partial. - Weakest classes:
date_printedanddate_handwritten— both hover around 0.31 mAP@50 on this split. Treat date detection as best-effort; consider date extraction as a downstream NER task on OCR'd text.
How to Use
1. Install DocLayout-YOLO
pip install doclayout-yolo
# (or follow the upstream installation: https://github.com/opendatalab/DocLayout-YOLO)
2. Download the checkpoint
from huggingface_hub import hf_hub_download
ckpt_path = hf_hub_download(
repo_id="<USER_OR_ORG>/kgb-archive-doclayout-yolov10m",
filename="best.pt",
)
3. Inference
from doclayout_yolo import YOLOv10
model = YOLOv10(ckpt_path)
results = model.predict(
"path/to/page.png",
imgsz=1024,
conf=0.25,
iou=0.45,
device="cuda", # or "cpu"
)
for box, cls_id, conf in zip(
results[0].boxes.xyxy.tolist(),
results[0].boxes.cls.tolist(),
results[0].boxes.conf.tolist(),
):
print(int(cls_id), conf, box)
Class names are in data.yaml (also in this repo).
Limitations & Caveats
- Single-split metrics. All numbers above are computed on one held-out validation set of ~50 images. Per-class numbers for low-frequency classes (
signature,seal,table, dates) ride on tens of instances and are not statistically tight. tableis unmeasurable at this scale (~2 instances in val, ~10 in corpus). The 0.995 mAP@50 figure is essentially a 2/2 hit rate, not a robust estimate.- Dates are weak.
date_printedanddate_handwrittenboth around 0.31 mAP@50. Date extraction is better treated as a post-OCR NER task. - Domain shift. Trained exclusively on mid-20th-century Ukrainian-language scanned document pages. Performance on other historical corpora, modern documents, or non-Cyrillic scripts is unknown and likely poor.
- Privacy. Historical record corpora can contain personal data. Deployment in any context that surfaces or extracts such data must be cleared with the relevant ethics / legal authority.
- Annotation noise. Ground truth was produced by multiple annotators in parallel and has not been triple-adjudicated. Some residual error reflects annotation drift, not model error.
License
This model inherits the AGPL-3.0 license from its upstream lineage:
- DocLayout-YOLO — AGPL-3.0
- Ultralytics YOLOv10 — AGPL-3.0
If you build on these weights, your derivative work and any network-deployed service that uses them must be released under AGPL-3.0 as well. For commercial / proprietary use, consult the Ultralytics commercial licence.
Citation
If you use this model, please cite the upstream architecture:
@inproceedings{zhao2024doclayoutyolo,
title={DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception},
author={Zhao, Zhiyuan and others},
year={2024},
}
Model tree for Hukyl/kgb-archive-doclayout-yolov10m
Base model
juliozhao/DocLayout-YOLO-DocStructBenchEvaluation results
- mAP@0.5 (all classes, weighted) on Private Ukrainian historical-document layout corpusself-reported0.637
- mAP@0.5:0.95 (all classes, weighted) on Private Ukrainian historical-document layout corpusself-reported0.402
- Precision (all classes) on Private Ukrainian historical-document layout corpusself-reported0.681
- Recall (all classes) on Private Ukrainian historical-document layout corpusself-reported0.610