--- title: MT Eval Leaderboard emoji: 🌍 colorFrom: indigo colorTo: blue sdk: gradio sdk_version: 5.34.0 app_file: app.py pinned: false license: apache-2.0 short_description: Translate public test sets with any HF model and rank them --- # 🌍 MT Eval Leaderboard A self-contained Hugging Face Space that **runs machine-translation evaluations and reports them like a leaderboard**: pick any Hugging Face model, pick a test set and language pairs, hit *Run* β€” the Space translates the test set, scores it with **BLEU** and **ChrF** (sacrebleu), and adds the result to a leaderboard you can slice three ways. ## Supported test sets The dataset picker has two levels β€” **collection (category) β†’ dataset (release) β†’ language pairs** β€” so every result stays categorized and is scored *per dataset per language pair*. ### WARMTH β€” consolidated catalog (recommended) [WARMTH](https://github.com/alvations/warmth) gathers a large family of MT-evaluation resources into one uniform parquet schema. A compact, pre-cleaned copy is **vendored into this repo** at `warmth_data/` (built by `scripts/build_warmth_bundle.py`), so test sets load with **no network access at runtime**. Each **collection** is a category; each **release** inside it is a selectable dataset: | Collection (category) | Example releases | Coverage | |---|---|---| | WMT24++ | `wmt24pp` | enβ†’55 locales, post-edited refs | | WMT General MT | `WMT15` … `WMT25`, `WMT24-humeval`, `WMT25-humeval` | WMT general-MT test sets | | NTREX-128 | `NTREX-128`, `NTREX-additional` | news, enβ†’127 | | FLORES+ | `FLORES-200` | 200+ languages (dev / devtest) | | WMT MQM Β· Metrics (hi-res) Β· Biomedical Β· Terminology Β· Bio-MQM | … | domain / human-rated sets | | IWSLT Β· mTEDx Β· MTNT Β· Multi30k Β· Tatoeba Β· DiaBLa | … | speech, noisy, captions, dialogue | Only (release, language-pair) combinations that carry **both** a source and a reference are included β€” a source is required to translate, so the source-less `wmt-metrics` QE releases are dropped when the bundle is built. WARMTH repeats each segment once per MT system; the bundle is de-duplicated to one segment per source, and WMT24++ "canary" watermark segments are removed. The full 544 MB of system outputs and human scores is distilled to ~106 MB of sources+references. Upstream sources consolidated by WARMTH include [WMT general-MT](https://github.com/wmt-conference), [WMT24++](https://huggingface.co/datasets/google/wmt24pp), [NTREX-128](https://github.com/MicrosoftTranslator/NTREX) and [FLORES+](https://github.com/openlanguagedata/flores), among others. ## Leaderboard views 1. **Summary** β€” one row per model Γ— dataset. Scores are aggregated across language pairs; the **Score** button cycles *average β†’ median β†’ min β†’ max*. Rows are ranked per dataset with the best value highlighted. 2. **By language pair** β€” the full breakdown: model Γ— dataset Γ— language pair, best model per pair highlighted. 3. **Model comparison** β€” executive pivot: rows are dataset Γ— language pair, one column per model, a single metric at a time (the **Metric** button toggles BLEU ⇄ ChrF), and a *Leader* column showing the winner and its margin over the runner-up. Ends each dataset with an average row. Model names link to their Hugging Face pages. **Export CSV** downloads the raw results; **Load demo data** seeds clearly-flagged illustrative numbers (for `google/gemma-4-e4b-it` and `Qwen/Qwen3.6-35B-A3B`) so you can explore the views before running a real evaluation. ## Deploying > πŸ“˜ Full step-by-step walkthrough (token setup, secrets, running the workflow, > deploying from `main`, troubleshooting): **[DEPLOYMENT.md](DEPLOYMENT.md)**. **Recommended β€” GitHub Actions** (no local setup; the runner reaches Hugging Face for you). This repo ships `.github/workflows/deploy-hf-space.yml`: 1. Add a repository secret **`HF_TOKEN`** (a *write* token from [huggingface.co/settings/tokens](https://huggingface.co/settings/tokens)): *Settings β†’ Secrets and variables β†’ Actions β†’ New repository secret*. 2. Go to the **Actions** tab β†’ *Deploy to Hugging Face Space* β†’ **Run workflow**. Set the target Space id (defaults to `doibungdidao/mt-eval-leaderboard`) and, optionally, hardware (`l4x1`, `a100-large`, …). The workflow creates the Space, uploads the repo, and stores your `HF_TOKEN` as the Space's runtime secret. It also re-deploys automatically on pushes to the working branch (skips cleanly if the secret isn't set yet). The token lives only in GitHub Secrets β€” it is never written into the repo. **Alternative β€” one command locally** (from a machine with network access to Hugging Face): ```bash pip install huggingface_hub export HF_TOKEN=hf_xxx # a WRITE token from huggingface.co/settings/tokens python deploy.py /mt-eval-leaderboard --set-token-secret ``` `deploy.py` creates the Space, uploads the repo, and (with `--set-token-secret`) stores your token as the Space's runtime `HF_TOKEN`. The token is read from the environment only β€” never committed. Pass `--hardware l4x1` (or `a100-large`, etc.) to provision GPU hardware for the local backends. **Or manually:** 1. Create a new Space (Gradio SDK) and upload this repository, or point the Space at this repo. 2. Add an **`HF_TOKEN`** secret (Settings β†’ Variables and secrets) β€” required for the default *Inference Providers API* backend and for gated models. 3. Pick hardware β€” see [Hardware sizing](#hardware-sizing) below. TL;DR: **CPU basic (free)** for the default API backend; a GPU sized to the model only if you use the local backend. 4. Optional: enable **persistent storage** β€” results are saved to `/data/mteval_results.json` and survive restarts. Without it they live in the container only. ## Hardware sizing Which instance you need depends entirely on the **inference backend**, because that decides where the model actually runs. ### Backend: Inference Providers API (the default) The model runs on Hugging Face's serverless providers, not on the Space. The Space itself only downloads test sets, sends API requests, and computes sacrebleu scores β€” all lightweight. | Instance | Fits? | Notes | |---|---|---| | **CPU basic** (2 vCPU Β· 16 GB) β€” free | βœ… **Recommended** | Plenty. The heaviest thing it ever does is parse the 128 MB WMT25 file once. | | CPU upgrade (8 vCPU Β· 32 GB) | βœ… | Only worth it if you raise the API concurrency in `mteval/translate.py` and want snappier metric computation on large runs. | Any GPU instance is wasted money on this backend. ### Backend: Local GPU (transformers) The model is loaded in the Space with `torch_dtype="auto"` (bf16), so budget roughly **2 GB of VRAM per billion parameters, plus ~20 % headroom** for the KV cache and activations. For MoE models, size for **total** parameters, not active ones β€” all experts sit in memory. | Model class | Example | bf16 weights | Recommended instance | |---|---|---|---| | ≀ 4–8 B dense / small MoE | `google/gemma-4-e4b-it` | ~8–16 GB | **Nvidia L4** (24 GB) β€” T4 small (16 GB) is marginal and lacks bf16 | | 8–20 B | 9–12 B translators | ~18–40 GB | **L40S** (48 GB) or A10G large | | 20–30 B dense | `google/gemma-4-26b-it` | ~52 GB | **A100 large or H100 (80 GB)** β€” L40S (48 GB) is too small once the KV cache is counted | | 30 B+ MoE / dense | `Qwen/Qwen3.6-35B-A3B` (~35 B total) | ~70 GB | **A100 large or H100 (80 GB)** | | 20 B+ **quantized 4-bit** (Unsloth bnb / NVFP4) | `unsloth/gemma-4-26b-it-NVFP4`, `unsloth/Qwen3.6-35B-A3B-NVFP4` | ~14–22 GB | **L4 / A10G (24 GB)** β€” see [Unsloth variants & NVFP4](#unsloth-variants--nvfp4) | Practical tips: - A 35 B MoE with only ~3 B *active* params generates fast once loaded, but it still needs the full ~70 GB of weights resident β€” don't try to squeeze it onto a 24/48 GB card without quantization. - To run the 35 B model on cheaper hardware, use its **Unsloth 4-bit or NVFP4 variant** β€” a first-class option in the UI, see the next section. - Set the Space **sleep timeout** generously or use "sleep after 1 h" β€” GPU Spaces bill while awake, and evaluation runs are bursty. - ZeroGPU Spaces are not supported by this code path as written (model loading happens outside a `@spaces.GPU` function); use a dedicated GPU instance for the local backend. Disk is not a constraint on any tier: test-set caches total well under 1 GB; model weights for the local backend go to the ephemeral disk, which is sized to the hardware tier. ## Unsloth variants & NVFP4 The **Model variant** control in the Run tab swaps the base model for its pre-quantized [Unsloth](https://huggingface.co/unsloth) checkpoint. The repo name is resolved automatically (`google/gemma-4-e4b-it` β†’ `unsloth/gemma-4-e4b-it-unsloth-bnb-4bit` / `unsloth/gemma-4-e4b-it-NVFP4`, same for `google/gemma-4-26b-it`, `Qwen/Qwen3.6-35B-A3B`, and any other slug you type); several naming conventions are checked against the Hub at run time, and the leaderboard records the exact resolved repo so quantized and full-precision runs rank side by side β€” handy for showing how much quality 4-bit actually costs on your language pairs. | Variant | What it is | Backend to select | Hardware | |---|---|---|---| | **Unsloth dynamic 4-bit (bnb)** | bitsandbytes 4-bit with Unsloth's per-layer dynamic quantization (sensitive layers kept in higher precision) | Local GPU (transformers) β€” uncomment `bitsandbytes` in `requirements.txt` | Works on any CUDA GPU, T4 and up; 35 B MoE fits in ~20 GB (L4/A10G) | | **Unsloth NVFP4** | NVIDIA's block-scaled FP4 format (per-16-value FP8 scales), near-bf16 quality at 4-bit size | **Local GPU (vLLM)** β€” uncomment `vllm` in `requirements.txt`. vLLM auto-detects the quantization from the checkpoint; transformers has no NVFP4 path, and the app blocks that combination with a clear error | Full FP4 tensor-core speedups need a **Blackwell** GPU (B200-class). On Hopper/Ada instances vLLM falls back to weight-only 4-bit kernels: you keep the ~4Γ— memory saving, not the compute speedup | Extra speed on unsloth checkpoints with the transformers backend: uncomment `unsloth` in `requirements.txt` and set a Space variable `USE_UNSLOTH=1` β€” the app then loads models through `FastLanguageModel` and silently falls back to plain transformers if the package is missing. Two caveats worth knowing: - The **Inference Providers API backend serves base models only** β€” providers generally don't host `unsloth/*` quantized repos, and the app warns if you try that combination. - If a variant repo doesn't exist for your model, the run stops with the list of names it tried; you can always paste an exact `unsloth/...` repo id into the model field with variant set to *Base model*. ### Notes - The first WMT25 run downloads a ~128 MB JSONL once, reduces it to a compact cache, and deletes the original. - BLEU tokenization adapts to the target language (`zh` for Chinese, `char` for Japanese/Korean/Thai, `13a` otherwise). ChrF is language-agnostic. - "Max segments per language pair" caps how much of the test set is used β€” keep it low (~100) for smoke tests, raise it for reportable numbers. - Re-running the same model Γ— dataset Γ— language pair overwrites the previous entry, so the board never shows stale duplicates.