Persian Named Entity Recognition (NER) using mDeBERTa-v3

This repository contains a fine-tuned version of microsoft/mdeberta-v3-base tailored for Persian Named Entity Recognition (NER).

Model Overview

  • Base Model: microsoft/mdeberta-v3-base
  • Dataset: mansoorhamidzadeh/Persian-NER-Dataset-500k
  • Task: Token Classification (39 IOB-format NER Classes)
  • Optimizer: Adafactor
  • Training Environment: PyTorch on WSL2 (Ubuntu 22.04)

Evaluation Results

The model was evaluated on the validation split using seqeval at the entity level:

Metric Score
Accuracy 93.67% (0.9367)
F1-Score 70.09% (0.7009)
Precision 68.72% (0.6872)
Recall 71.51% (0.7151)
Eval Loss 0.2700

Usage

You can easily use this model with Hugging Face transformers pipeline:

from transformers import pipeline

# Load pipeline directly from Hugging Face Hub
ner_pipeline = pipeline(
    "token-classification", 
    model="SalmaShirdel/persian-mdeberta-v3-ner", 
    aggregation_strategy="simple"
)

# Example Persian sentence
text = "دانشگاه تهران در خیابان انقلاب قرار دارد."
results = ner_pipeline(text)

for entity in results:
    print(f"Entity: {entity['word']} | Group: {entity['entity_group']} | Score: {entity['score']:.4f}")
Downloads last month
28
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train SalmaShirdel/persian-mdeberta-v3-ner

Space using SalmaShirdel/persian-mdeberta-v3-ner 1