Instructions to use ahmedsamirtarjama/Tashkeel-50M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ahmedsamirtarjama/Tashkeel-50M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ahmedsamirtarjama/Tashkeel-50M")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ahmedsamirtarjama/Tashkeel-50M") model = AutoModelForCausalLM.from_pretrained("ahmedsamirtarjama/Tashkeel-50M", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ahmedsamirtarjama/Tashkeel-50M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ahmedsamirtarjama/Tashkeel-50M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ahmedsamirtarjama/Tashkeel-50M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/ahmedsamirtarjama/Tashkeel-50M
- SGLang
How to use ahmedsamirtarjama/Tashkeel-50M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ahmedsamirtarjama/Tashkeel-50M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ahmedsamirtarjama/Tashkeel-50M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ahmedsamirtarjama/Tashkeel-50M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ahmedsamirtarjama/Tashkeel-50M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use ahmedsamirtarjama/Tashkeel-50M with Docker Model Runner:
docker model run hf.co/ahmedsamirtarjama/Tashkeel-50M
Tashkeel-50M
A ~50M parameter Arabic diacritization (ุชุดููู) model fine-tuned from oddadmix/50M-2048-Emhotob on Misraj/Sadeed_Tashkeela.
It is a small causal LM intended for fast, on-device / low-cost Arabic tashkeel experiments.
Model details
| Architecture | LLaMA-style (LlamaForCausalLM) |
| Parameters | ~50M |
| Context | 2048 tokens |
| Hidden size | 512 |
| Layers | 12 |
| Vocab size | 32,000 |
| Precision | bfloat16 |
| Base model | oddadmix/50M-2048-Emhotob |
| Training data | Misraj/Sadeed_Tashkeela (train split) |
Prompt format
Training and inference use this prompt:
ูู
ุจุชุดููู ูุฐุฉ ุงูุฌู
ูู : {undiacritized_text}
The model should continue with the diacritized Arabic text.
Quick start
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ahmedsamirtarjama/Tashkeel-50M"
device = "cuda" if torch.cuda.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(model_id)
tok.padding_side = "left"
if tok.pad_token is None:
tok.pad_token = tok.eos_token
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16 if device == "cuda" else torch.float32,
).to(device)
model.eval()
text = "ุงููุบุฉ ุงูุนุฑุจูุฉ ูุบุฉ ุฌู
ููุฉ"
prompt = f"ูู
ุจุชุดููู ูุฐุฉ ุงูุฌู
ูู : {text}\n"
inputs = tok(prompt, return_tensors="pt").to(device)
with torch.inference_mode():
out = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
pad_token_id=tok.pad_token_id,
)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Training
Full fine-tune (not LoRA) with Hugging Face Trainer:
| Hyperparameter | Value |
|---|---|
| Epochs | 1 |
| Learning rate | 3e-4 |
| Scheduler | cosine |
| Warmup steps | 500 |
| Batch size | 32 |
| Max sequence length | 768 (longer examples dropped) |
| Loss | next-token LM loss on the diacritized target only (prompt tokens masked with -100) |
| Precision | bfloat16 |
Approximate training recipe:
<source prompt> + <diacritized target> + </s>
Evaluation
Evaluated on Misraj/SadeedDiac-25 with standard Morph/Total DER & WER (missing GT diacritics skipped).
Mapping used below: Total โ (CE) (with case endings), Morph โ (w/o CE) (without case endings). Hallucinations โ share of examples skipped due to word-count mismatch.
| Model | DER (CE) | WER (CE) | DER (w/o CE) | WER (w/o CE) | Hallucinations |
|---|---|---|---|---|---|
| Claude-3-7-Sonnet | 1.39 | 4.67 | 0.77 | 2.31 | 0.82 |
| Tashkeel-50M | 3.08* | 9.56* | 2.26* | 6.77* | ~99 |
| GPT-4 | 3.86 | 5.27 | 3.86 | 10.93 | 1.02 |
| Gemini-Flash-2.0 | 3.19 | 7.99 | 2.38 | 5.50 | 1.17 |
| Sadeed | 7.29 | 13.74 | 5.26 | 9.92 | 7.19 |
*Tashkeel-50M DER/WER are computed only on examples where the prediction and reference have the same word count. Because most generations change length (truncation / repetition / insertions), they are skipped by the length-matching evaluator โ hence the high hallucination rate. Treat the starred numbers as optimistic and not a full apples-to-apples comparison with systems that preserve word identity on nearly all examples.
For production tashkeel, prefer stronger constrained models or add decoding constraints that keep the undiacritized skeleton fixed.
Intended use
- Research and prototyping for Arabic diacritization
- Baseline for small / efficient tashkeel models
- Educational demos of causal-LM fine-tuning for sequence transduction
Limitations
- Small capacity (~50M); quality lags dedicated / large instruction models on hard classical Arabic
- Causal generation can truncate, repeat, or insert words; DER/WER only apply when word counts match
- Prompt is Arabic-instruction style; changing the prompt may degrade quality
- Not a general-purpose chat model
Citation
If you use this model, please also cite the base model and dataset:
@misc{tashkeel50m,
title = {Tashkeel-50M},
author = {Ahmed Samir},
year = {2026},
howpublished = {\url{https://huggingface.co/ahmedsamirtarjama/Tashkeel-50M}}
}
- Base:
oddadmix/50M-2048-Emhotob - Data:
Misraj/Sadeed_Tashkeela
- Downloads last month
- 146
Model tree for ahmedsamirtarjama/Tashkeel-50M
Base model
oddadmix/50M-2048-Emhotob