Ukrainian Historical Document Layout Detector — DocLayout-YOLOv10m (fine-tuned)

A fine-tuned DocLayout-YOLO (YOLOv10m-doclayout) detector for layout-region detection on scanned Ukrainian-language historical document pages. Trained as part of a course project on document digitisation; intended as the detection stage of an end-to-end OCR pipeline (detection → preprocessing → printed/handwritten OCR).

The training corpus is private and is not redistributed with this model.

Classes (7)

ID Label Description
0 printed_text Printed Cyrillic body text
1 handwritten_text Handwritten Cyrillic text (notes, annotations)
2 signature Hand-signed names
3 seal Stamped seals / round office stamps
4 table Tabular regions
5 date_printed Printed dates (form fields)
6 date_handwritten Handwritten dates

Intended Use

  • Layout-region proposal for downstream OCR on Ukrainian-language historical document scans.
  • Feeding crops into:
    • PaddleOCR-VL / PP-OCRv5 eslav for printed_text / date_printed
    • TrOCR (Cyrillic, fine-tuned) for handwritten_text / date_handwritten
  • Region-of-interest extraction for downstream tasks: NER, semantic search, signature clustering.

Out-of-Scope / Not Recommended

  • Modern documents, non-Cyrillic scripts, non-archival photography.
  • Privacy-sensitive deployments without legal review — historical record corpora can contain personal data; consult the appropriate ethics / legal authority before any downstream use that surfaces such data.
  • Production OCR for live documents — this is a research artefact, not an audited product.

Training Data

Source Private Ukrainian-language historical-document corpus (mid-20th-century scans)
Annotated images (post prepare-data) ~330
Annotated bounding boxes ~13k
Annotators Multi-annotator effort across several CVAT batches
Train / val split ~280 train / ~50 val (offline-augmented training set: ~1.7k images / ~61k boxes)

Approximate class frequency in the full corpus (before split)

Class Boxes Share
printed_text ~10,000 ~75 %
handwritten_text ~2,500 ~19 %
date_printed ~275 ~2 %
signature ~250 ~2 %
date_handwritten ~210 ~2 %
seal ~45 <1 %
table ~10 <1 %

The corpus is dominated by printed text. seal and table are still small (low tens of instances corpus-wide).


Training Procedure

This release corresponds to a single-split final training run (not cross-validation).

Base model

  • Architecture: YOLOv10m-doclayout (DocLayout-YOLO), 19,970,122 params, 68.0 GFLOPs.
  • Initialisation: weights transplanted from juliozhao/DocLayout-YOLO-DocStructBench (DocStructBench, 10 classes) onto a fresh 7-class head. Backbone (model.0–model.10) frozen during fine-tuning.

Hyperparameters

Optimizer AdamW (auto-selected by Ultralytics)
lr0 0.000909
Momentum 0.9
Weight decay 0.0005
Warmup 3 epochs
Schedule cosine off, close_mosaic=10 (no effect; mosaic disabled)
Mixed precision fp16 (AMP)
Image size 1024
Batch size 4
Max epochs 100
Early-stopping patience 25 epochs
Frozen layers model.0 – model.10 (freeze=11)
Hardware NVIDIA RTX 4070, 12 GB VRAM

Augmentation

  • Online (Ultralytics defaults): HSV (0.015/0.7/0.4), translate 0.1, scale 0.5, RandAugment, erasing 0.4. Mosaic / MixUp / CopyPaste disabled.
  • Offline (training set only): crops, small rotations, asymmetric margins, paper-tone shifts, Gaussian noise, JPEG compression artefacts. Validation set untouched.

Run

  • Wall-clock ≈ 49 min on a single RTX 4070.
  • Early stopping fired at epoch 38 / 100; best fitness (0.1·mAP@50 + 0.9·mAP@50-95) reached at epoch 13.
  • The released best.pt corresponds to epoch 13.

Evaluation

Metrics are computed on the held-out validation set of the single split (~50 images, ~2.1k boxes). The validation set received no augmentation.

Headline (best.pt re-validated)

Metric Value
mAP@50 0.637
mAP@50-95 0.402
Precision 0.681
Recall 0.610

Per-class

Class n (val, approx.) Precision Recall mAP@50 mAP@50-95
printed_text ~1,700 0.721 0.822 0.816 0.527
handwritten_text ~300 0.641 0.598 0.604 0.327
signature ~35 0.607 0.838 0.831 0.492
seal ~5 1.000 0.394 0.575 0.291
table ~2 1.000 0.934 0.995 0.812
date_printed ~40 0.346 0.395 0.324 0.196
date_handwritten ~35 0.450 0.289 0.314 0.170

Operating-point recommendation

  • Reliable for downstream OCR: printed_text (0.82 mAP@50) and handwritten_text (0.60 mAP@50).
  • Usable but high variance on this val draw: signature — single-split estimate; treat as upper bound.
  • Data-starved: seal, table. Predictions are precise (P = 1.0 in both) but recall is partial.
  • Weakest classes: date_printed and date_handwritten — both hover around 0.31 mAP@50 on this split. Treat date detection as best-effort; consider date extraction as a downstream NER task on OCR'd text.

How to Use

1. Install DocLayout-YOLO

pip install doclayout-yolo
# (or follow the upstream installation: https://github.com/opendatalab/DocLayout-YOLO)

2. Download the checkpoint

from huggingface_hub import hf_hub_download

ckpt_path = hf_hub_download(
    repo_id="<USER_OR_ORG>/kgb-archive-doclayout-yolov10m",
    filename="best.pt",
)

3. Inference

from doclayout_yolo import YOLOv10

model = YOLOv10(ckpt_path)

results = model.predict(
    "path/to/page.png",
    imgsz=1024,
    conf=0.25,
    iou=0.45,
    device="cuda",  # or "cpu"
)

for box, cls_id, conf in zip(
    results[0].boxes.xyxy.tolist(),
    results[0].boxes.cls.tolist(),
    results[0].boxes.conf.tolist(),
):
    print(int(cls_id), conf, box)

Class names are in data.yaml (also in this repo).


Limitations & Caveats

  • Single-split metrics. All numbers above are computed on one held-out validation set of ~50 images. Per-class numbers for low-frequency classes (signature, seal, table, dates) ride on tens of instances and are not statistically tight.
  • table is unmeasurable at this scale (~2 instances in val, ~10 in corpus). The 0.995 mAP@50 figure is essentially a 2/2 hit rate, not a robust estimate.
  • Dates are weak. date_printed and date_handwritten both around 0.31 mAP@50. Date extraction is better treated as a post-OCR NER task.
  • Domain shift. Trained exclusively on mid-20th-century Ukrainian-language scanned document pages. Performance on other historical corpora, modern documents, or non-Cyrillic scripts is unknown and likely poor.
  • Privacy. Historical record corpora can contain personal data. Deployment in any context that surfaces or extracts such data must be cleared with the relevant ethics / legal authority.
  • Annotation noise. Ground truth was produced by multiple annotators in parallel and has not been triple-adjudicated. Some residual error reflects annotation drift, not model error.

License

This model inherits the AGPL-3.0 license from its upstream lineage:

  • DocLayout-YOLO — AGPL-3.0
  • Ultralytics YOLOv10 — AGPL-3.0

If you build on these weights, your derivative work and any network-deployed service that uses them must be released under AGPL-3.0 as well. For commercial / proprietary use, consult the Ultralytics commercial licence.

Citation

If you use this model, please cite the upstream architecture:

@inproceedings{zhao2024doclayoutyolo,
  title={DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception},
  author={Zhao, Zhiyuan and others},
  year={2024},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Hukyl/kgb-archive-doclayout-yolov10m

Finetuned
(6)
this model

Evaluation results

  • mAP@0.5 (all classes, weighted) on Private Ukrainian historical-document layout corpus
    self-reported
    0.637
  • mAP@0.5:0.95 (all classes, weighted) on Private Ukrainian historical-document layout corpus
    self-reported
    0.402
  • Precision (all classes) on Private Ukrainian historical-document layout corpus
    self-reported
    0.681
  • Recall (all classes) on Private Ukrainian historical-document layout corpus
    self-reported
    0.610