Tashkeel-50M

A ~50M parameter Arabic diacritization (ุชุดูƒูŠู„) model fine-tuned from oddadmix/50M-2048-Emhotob on Misraj/Sadeed_Tashkeela.

It is a small causal LM intended for fast, on-device / low-cost Arabic tashkeel experiments.

Model details

Architecture LLaMA-style (LlamaForCausalLM)
Parameters ~50M
Context 2048 tokens
Hidden size 512
Layers 12
Vocab size 32,000
Precision bfloat16
Base model oddadmix/50M-2048-Emhotob
Training data Misraj/Sadeed_Tashkeela (train split)

Prompt format

Training and inference use this prompt:

ู‚ู… ุจุชุดูƒูŠู„ ู‡ุฐุฉ ุงู„ุฌู…ู„ู‡ : {undiacritized_text}

The model should continue with the diacritized Arabic text.

Quick start

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ahmedsamirtarjama/Tashkeel-50M"
device = "cuda" if torch.cuda.is_available() else "cpu"

tok = AutoTokenizer.from_pretrained(model_id)
tok.padding_side = "left"
if tok.pad_token is None:
    tok.pad_token = tok.eos_token

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16 if device == "cuda" else torch.float32,
).to(device)
model.eval()

text = "ุงู„ู„ุบุฉ ุงู„ุนุฑุจูŠุฉ ู„ุบุฉ ุฌู…ูŠู„ุฉ"
prompt = f"ู‚ู… ุจุชุดูƒูŠู„ ู‡ุฐุฉ ุงู„ุฌู…ู„ู‡ : {text}\n"
inputs = tok(prompt, return_tensors="pt").to(device)

with torch.inference_mode():
    out = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=False,
        pad_token_id=tok.pad_token_id,
    )

print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Training

Full fine-tune (not LoRA) with Hugging Face Trainer:

Hyperparameter Value
Epochs 1
Learning rate 3e-4
Scheduler cosine
Warmup steps 500
Batch size 32
Max sequence length 768 (longer examples dropped)
Loss next-token LM loss on the diacritized target only (prompt tokens masked with -100)
Precision bfloat16

Approximate training recipe:

<source prompt> + <diacritized target> + </s>

Evaluation

Evaluated on Misraj/SadeedDiac-25 with standard Morph/Total DER & WER (missing GT diacritics skipped).

Mapping used below: Total โ‰ˆ (CE) (with case endings), Morph โ‰ˆ (w/o CE) (without case endings). Hallucinations โ‰ˆ share of examples skipped due to word-count mismatch.

Model DER (CE) WER (CE) DER (w/o CE) WER (w/o CE) Hallucinations
Claude-3-7-Sonnet 1.39 4.67 0.77 2.31 0.82
Tashkeel-50M 3.08* 9.56* 2.26* 6.77* ~99
GPT-4 3.86 5.27 3.86 10.93 1.02
Gemini-Flash-2.0 3.19 7.99 2.38 5.50 1.17
Sadeed 7.29 13.74 5.26 9.92 7.19

*Tashkeel-50M DER/WER are computed only on examples where the prediction and reference have the same word count. Because most generations change length (truncation / repetition / insertions), they are skipped by the length-matching evaluator โ€” hence the high hallucination rate. Treat the starred numbers as optimistic and not a full apples-to-apples comparison with systems that preserve word identity on nearly all examples.

For production tashkeel, prefer stronger constrained models or add decoding constraints that keep the undiacritized skeleton fixed.

Intended use

  • Research and prototyping for Arabic diacritization
  • Baseline for small / efficient tashkeel models
  • Educational demos of causal-LM fine-tuning for sequence transduction

Limitations

  • Small capacity (~50M); quality lags dedicated / large instruction models on hard classical Arabic
  • Causal generation can truncate, repeat, or insert words; DER/WER only apply when word counts match
  • Prompt is Arabic-instruction style; changing the prompt may degrade quality
  • Not a general-purpose chat model

Citation

If you use this model, please also cite the base model and dataset:

@misc{tashkeel50m,
  title        = {Tashkeel-50M},
  author       = {Ahmed Samir},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/ahmedsamirtarjama/Tashkeel-50M}}
}
Downloads last month
146
Safetensors
Model size
51.8M params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ahmedsamirtarjama/Tashkeel-50M

Finetuned
(16)
this model

Dataset used to train ahmedsamirtarjama/Tashkeel-50M