Download README.md from doibungdidao/mt-eval-leaderboard: direct link, hf CLI and curl.
- Browser
- Download file 11.4 kB
-
https://huggingface.co/spaces/doibungdidao/mt-eval-leaderboard/resolve/main/README.md
- Command line
-
hf download hf://spaces/doibungdidao/mt-eval-leaderboard/README.md
-
curl -L -o README.md https://huggingface.co/spaces/doibungdidao/mt-eval-leaderboard/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.29.1
title: MT Eval Leaderboard
emoji: π
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: 5.34.0
app_file: app.py
pinned: false
license: apache-2.0
short_description: Translate public test sets with any HF model and rank them
π MT Eval Leaderboard
A self-contained Hugging Face Space that runs machine-translation evaluations and reports them like a leaderboard: pick any Hugging Face model, pick a test set and language pairs, hit Run β the Space translates the test set, scores it with BLEU and ChrF (sacrebleu), and adds the result to a leaderboard you can slice three ways.
Supported test sets
The dataset picker has two levels β collection (category) β dataset (release) β language pairs β so every result stays categorized and is scored per dataset per language pair.
WARMTH β consolidated catalog (recommended)
WARMTH gathers a large family of
MT-evaluation resources into one uniform parquet schema. A compact, pre-cleaned
copy is vendored into this repo at warmth_data/ (built by
scripts/build_warmth_bundle.py), so test sets load with no network access at
runtime. Each collection is a category; each release inside it is a
selectable dataset:
| Collection (category) | Example releases | Coverage |
|---|---|---|
| WMT24++ | wmt24pp |
enβ55 locales, post-edited refs |
| WMT General MT | WMT15 β¦ WMT25, WMT24-humeval, WMT25-humeval |
WMT general-MT test sets |
| NTREX-128 | NTREX-128, NTREX-additional |
news, enβ127 |
| FLORES+ | FLORES-200 |
200+ languages (dev / devtest) |
| WMT MQM Β· Metrics (hi-res) Β· Biomedical Β· Terminology Β· Bio-MQM | β¦ | domain / human-rated sets |
| IWSLT Β· mTEDx Β· MTNT Β· Multi30k Β· Tatoeba Β· DiaBLa | β¦ | speech, noisy, captions, dialogue |
Only (release, language-pair) combinations that carry both a source and a
reference are included β a source is required to translate, so the source-less
wmt-metrics QE releases are dropped when the bundle is built. WARMTH repeats
each segment once per MT system; the bundle is de-duplicated to one segment per
source, and WMT24++ "canary" watermark segments are removed. The full 544 MB of
system outputs and human scores is distilled to ~106 MB of sources+references.
Upstream sources consolidated by WARMTH include WMT general-MT, WMT24++, NTREX-128 and FLORES+, among others.
Leaderboard views
- Summary β one row per model Γ dataset. Scores are aggregated across language pairs; the Score button cycles average β median β min β max. Rows are ranked per dataset with the best value highlighted.
- By language pair β the full breakdown: model Γ dataset Γ language pair, best model per pair highlighted.
- Model comparison β executive pivot: rows are dataset Γ language pair, one column per model, a single metric at a time (the Metric button toggles BLEU β ChrF), and a Leader column showing the winner and its margin over the runner-up. Ends each dataset with an average row.
Model names link to their Hugging Face pages. Export CSV downloads the raw
results; Load demo data seeds clearly-flagged illustrative numbers (for
google/gemma-4-e4b-it and Qwen/Qwen3.6-35B-A3B) so you can explore the
views before running a real evaluation.
Deploying
π Full step-by-step walkthrough (token setup, secrets, running the workflow, deploying from
main, troubleshooting): DEPLOYMENT.md.
Recommended β GitHub Actions (no local setup; the runner reaches Hugging
Face for you). This repo ships .github/workflows/deploy-hf-space.yml:
- Add a repository secret
HF_TOKEN(a write token from huggingface.co/settings/tokens): Settings β Secrets and variables β Actions β New repository secret. - Go to the Actions tab β Deploy to Hugging Face Space β Run
workflow. Set the target Space id (defaults to
doibungdidao/mt-eval-leaderboard) and, optionally, hardware (l4x1,a100-large, β¦).
The workflow creates the Space, uploads the repo, and stores your HF_TOKEN
as the Space's runtime secret. It also re-deploys automatically on pushes to
the working branch (skips cleanly if the secret isn't set yet). The token
lives only in GitHub Secrets β it is never written into the repo.
Alternative β one command locally (from a machine with network access to Hugging Face):
pip install huggingface_hub
export HF_TOKEN=hf_xxx # a WRITE token from huggingface.co/settings/tokens
python deploy.py <username>/mt-eval-leaderboard --set-token-secret
deploy.py creates the Space, uploads the repo, and (with
--set-token-secret) stores your token as the Space's runtime HF_TOKEN.
The token is read from the environment only β never committed. Pass
--hardware l4x1 (or a100-large, etc.) to provision GPU hardware for the
local backends.
Or manually:
- Create a new Space (Gradio SDK) and upload this repository, or point the Space at this repo.
- Add an
HF_TOKENsecret (Settings β Variables and secrets) β required for the default Inference Providers API backend and for gated models. - Pick hardware β see Hardware sizing below. TL;DR: CPU basic (free) for the default API backend; a GPU sized to the model only if you use the local backend.
- Optional: enable persistent storage β results are saved to
/data/mteval_results.jsonand survive restarts. Without it they live in the container only.
Hardware sizing
Which instance you need depends entirely on the inference backend, because that decides where the model actually runs.
Backend: Inference Providers API (the default)
The model runs on Hugging Face's serverless providers, not on the Space. The Space itself only downloads test sets, sends API requests, and computes sacrebleu scores β all lightweight.
| Instance | Fits? | Notes |
|---|---|---|
| CPU basic (2 vCPU Β· 16 GB) β free | β Recommended | Plenty. The heaviest thing it ever does is parse the 128 MB WMT25 file once. |
| CPU upgrade (8 vCPU Β· 32 GB) | β | Only worth it if you raise the API concurrency in mteval/translate.py and want snappier metric computation on large runs. |
Any GPU instance is wasted money on this backend.
Backend: Local GPU (transformers)
The model is loaded in the Space with torch_dtype="auto" (bf16), so budget
roughly 2 GB of VRAM per billion parameters, plus ~20 % headroom for the
KV cache and activations. For MoE models, size for total parameters, not
active ones β all experts sit in memory.
| Model class | Example | bf16 weights | Recommended instance |
|---|---|---|---|
| β€ 4β8 B dense / small MoE | google/gemma-4-e4b-it |
~8β16 GB | Nvidia L4 (24 GB) β T4 small (16 GB) is marginal and lacks bf16 |
| 8β20 B | 9β12 B translators | ~18β40 GB | L40S (48 GB) or A10G large |
| 20β30 B dense | google/gemma-4-26b-it |
~52 GB | A100 large or H100 (80 GB) β L40S (48 GB) is too small once the KV cache is counted |
| 30 B+ MoE / dense | Qwen/Qwen3.6-35B-A3B (~35 B total) |
~70 GB | A100 large or H100 (80 GB) |
| 20 B+ quantized 4-bit (Unsloth bnb / NVFP4) | unsloth/gemma-4-26b-it-NVFP4, unsloth/Qwen3.6-35B-A3B-NVFP4 |
~14β22 GB | L4 / A10G (24 GB) β see Unsloth variants & NVFP4 |
Practical tips:
- A 35 B MoE with only ~3 B active params generates fast once loaded, but it still needs the full ~70 GB of weights resident β don't try to squeeze it onto a 24/48 GB card without quantization.
- To run the 35 B model on cheaper hardware, use its Unsloth 4-bit or NVFP4 variant β a first-class option in the UI, see the next section.
- Set the Space sleep timeout generously or use "sleep after 1 h" β GPU Spaces bill while awake, and evaluation runs are bursty.
- ZeroGPU Spaces are not supported by this code path as written (model loading
happens outside a
@spaces.GPUfunction); use a dedicated GPU instance for the local backend.
Disk is not a constraint on any tier: test-set caches total well under 1 GB; model weights for the local backend go to the ephemeral disk, which is sized to the hardware tier.
Unsloth variants & NVFP4
The Model variant control in the Run tab swaps the base model for its
pre-quantized Unsloth checkpoint. The repo
name is resolved automatically (google/gemma-4-e4b-it β
unsloth/gemma-4-e4b-it-unsloth-bnb-4bit / unsloth/gemma-4-e4b-it-NVFP4,
same for google/gemma-4-26b-it, Qwen/Qwen3.6-35B-A3B, and any other slug
you type); several naming
conventions are checked against the Hub at run time, and the leaderboard
records the exact resolved repo so quantized and full-precision runs rank
side by side β handy for showing how much quality 4-bit actually costs on
your language pairs.
| Variant | What it is | Backend to select | Hardware |
|---|---|---|---|
| Unsloth dynamic 4-bit (bnb) | bitsandbytes 4-bit with Unsloth's per-layer dynamic quantization (sensitive layers kept in higher precision) | Local GPU (transformers) β uncomment bitsandbytes in requirements.txt |
Works on any CUDA GPU, T4 and up; 35 B MoE fits in ~20 GB (L4/A10G) |
| Unsloth NVFP4 | NVIDIA's block-scaled FP4 format (per-16-value FP8 scales), near-bf16 quality at 4-bit size | Local GPU (vLLM) β uncomment vllm in requirements.txt. vLLM auto-detects the quantization from the checkpoint; transformers has no NVFP4 path, and the app blocks that combination with a clear error |
Full FP4 tensor-core speedups need a Blackwell GPU (B200-class). On Hopper/Ada instances vLLM falls back to weight-only 4-bit kernels: you keep the ~4Γ memory saving, not the compute speedup |
Extra speed on unsloth checkpoints with the transformers backend: uncomment
unsloth in requirements.txt and set a Space variable USE_UNSLOTH=1 β
the app then loads models through FastLanguageModel and silently falls back
to plain transformers if the package is missing.
Two caveats worth knowing:
- The Inference Providers API backend serves base models only β providers
generally don't host
unsloth/*quantized repos, and the app warns if you try that combination. - If a variant repo doesn't exist for your model, the run stops with the list
of names it tried; you can always paste an exact
unsloth/...repo id into the model field with variant set to Base model.
Notes
- The first WMT25 run downloads a ~128 MB JSONL once, reduces it to a compact cache, and deletes the original.
- BLEU tokenization adapts to the target language (
zhfor Chinese,charfor Japanese/Korean/Thai,13aotherwise). ChrF is language-agnostic. - "Max segments per language pair" caps how much of the test set is used β keep it low (~100) for smoke tests, raise it for reportable numbers.
- Re-running the same model Γ dataset Γ language pair overwrites the previous entry, so the board never shows stale duplicates.