mt-eval-leaderboard / README.md
doibungdidao's picture
Deploy MT Eval Leaderboard
9b7eb64 verified
|
Raw History Blame Contribute Delete
11.4 kB

A newer version of the Gradio SDK is available: 6.29.1

Upgrade
metadata
title: MT Eval Leaderboard
emoji: 🌍
colorFrom: indigo
colorTo: blue
sdk: gradio
sdk_version: 5.34.0
app_file: app.py
pinned: false
license: apache-2.0
short_description: Translate public test sets with any HF model and rank them

🌍 MT Eval Leaderboard

A self-contained Hugging Face Space that runs machine-translation evaluations and reports them like a leaderboard: pick any Hugging Face model, pick a test set and language pairs, hit Run β€” the Space translates the test set, scores it with BLEU and ChrF (sacrebleu), and adds the result to a leaderboard you can slice three ways.

Supported test sets

The dataset picker has two levels β€” collection (category) β†’ dataset (release) β†’ language pairs β€” so every result stays categorized and is scored per dataset per language pair.

WARMTH β€” consolidated catalog (recommended)

WARMTH gathers a large family of MT-evaluation resources into one uniform parquet schema. A compact, pre-cleaned copy is vendored into this repo at warmth_data/ (built by scripts/build_warmth_bundle.py), so test sets load with no network access at runtime. Each collection is a category; each release inside it is a selectable dataset:

Collection (category) Example releases Coverage
WMT24++ wmt24pp en→55 locales, post-edited refs
WMT General MT WMT15 … WMT25, WMT24-humeval, WMT25-humeval WMT general-MT test sets
NTREX-128 NTREX-128, NTREX-additional news, en→127
FLORES+ FLORES-200 200+ languages (dev / devtest)
WMT MQM Β· Metrics (hi-res) Β· Biomedical Β· Terminology Β· Bio-MQM … domain / human-rated sets
IWSLT Β· mTEDx Β· MTNT Β· Multi30k Β· Tatoeba Β· DiaBLa … speech, noisy, captions, dialogue

Only (release, language-pair) combinations that carry both a source and a reference are included β€” a source is required to translate, so the source-less wmt-metrics QE releases are dropped when the bundle is built. WARMTH repeats each segment once per MT system; the bundle is de-duplicated to one segment per source, and WMT24++ "canary" watermark segments are removed. The full 544 MB of system outputs and human scores is distilled to ~106 MB of sources+references.

Upstream sources consolidated by WARMTH include WMT general-MT, WMT24++, NTREX-128 and FLORES+, among others.

Leaderboard views

  1. Summary β€” one row per model Γ— dataset. Scores are aggregated across language pairs; the Score button cycles average β†’ median β†’ min β†’ max. Rows are ranked per dataset with the best value highlighted.
  2. By language pair β€” the full breakdown: model Γ— dataset Γ— language pair, best model per pair highlighted.
  3. Model comparison β€” executive pivot: rows are dataset Γ— language pair, one column per model, a single metric at a time (the Metric button toggles BLEU ⇄ ChrF), and a Leader column showing the winner and its margin over the runner-up. Ends each dataset with an average row.

Model names link to their Hugging Face pages. Export CSV downloads the raw results; Load demo data seeds clearly-flagged illustrative numbers (for google/gemma-4-e4b-it and Qwen/Qwen3.6-35B-A3B) so you can explore the views before running a real evaluation.

Deploying

πŸ“˜ Full step-by-step walkthrough (token setup, secrets, running the workflow, deploying from main, troubleshooting): DEPLOYMENT.md.

Recommended β€” GitHub Actions (no local setup; the runner reaches Hugging Face for you). This repo ships .github/workflows/deploy-hf-space.yml:

  1. Add a repository secret HF_TOKEN (a write token from huggingface.co/settings/tokens): Settings β†’ Secrets and variables β†’ Actions β†’ New repository secret.
  2. Go to the Actions tab β†’ Deploy to Hugging Face Space β†’ Run workflow. Set the target Space id (defaults to doibungdidao/mt-eval-leaderboard) and, optionally, hardware (l4x1, a100-large, …).

The workflow creates the Space, uploads the repo, and stores your HF_TOKEN as the Space's runtime secret. It also re-deploys automatically on pushes to the working branch (skips cleanly if the secret isn't set yet). The token lives only in GitHub Secrets β€” it is never written into the repo.

Alternative β€” one command locally (from a machine with network access to Hugging Face):

pip install huggingface_hub
export HF_TOKEN=hf_xxx          # a WRITE token from huggingface.co/settings/tokens
python deploy.py <username>/mt-eval-leaderboard --set-token-secret

deploy.py creates the Space, uploads the repo, and (with --set-token-secret) stores your token as the Space's runtime HF_TOKEN. The token is read from the environment only β€” never committed. Pass --hardware l4x1 (or a100-large, etc.) to provision GPU hardware for the local backends.

Or manually:

  1. Create a new Space (Gradio SDK) and upload this repository, or point the Space at this repo.
  2. Add an HF_TOKEN secret (Settings β†’ Variables and secrets) β€” required for the default Inference Providers API backend and for gated models.
  3. Pick hardware β€” see Hardware sizing below. TL;DR: CPU basic (free) for the default API backend; a GPU sized to the model only if you use the local backend.
  4. Optional: enable persistent storage β€” results are saved to /data/mteval_results.json and survive restarts. Without it they live in the container only.

Hardware sizing

Which instance you need depends entirely on the inference backend, because that decides where the model actually runs.

Backend: Inference Providers API (the default)

The model runs on Hugging Face's serverless providers, not on the Space. The Space itself only downloads test sets, sends API requests, and computes sacrebleu scores β€” all lightweight.

Instance Fits? Notes
CPU basic (2 vCPU Β· 16 GB) β€” free βœ… Recommended Plenty. The heaviest thing it ever does is parse the 128 MB WMT25 file once.
CPU upgrade (8 vCPU Β· 32 GB) βœ… Only worth it if you raise the API concurrency in mteval/translate.py and want snappier metric computation on large runs.

Any GPU instance is wasted money on this backend.

Backend: Local GPU (transformers)

The model is loaded in the Space with torch_dtype="auto" (bf16), so budget roughly 2 GB of VRAM per billion parameters, plus ~20 % headroom for the KV cache and activations. For MoE models, size for total parameters, not active ones β€” all experts sit in memory.

Model class Example bf16 weights Recommended instance
≀ 4–8 B dense / small MoE google/gemma-4-e4b-it ~8–16 GB Nvidia L4 (24 GB) β€” T4 small (16 GB) is marginal and lacks bf16
8–20 B 9–12 B translators ~18–40 GB L40S (48 GB) or A10G large
20–30 B dense google/gemma-4-26b-it ~52 GB A100 large or H100 (80 GB) β€” L40S (48 GB) is too small once the KV cache is counted
30 B+ MoE / dense Qwen/Qwen3.6-35B-A3B (~35 B total) ~70 GB A100 large or H100 (80 GB)
20 B+ quantized 4-bit (Unsloth bnb / NVFP4) unsloth/gemma-4-26b-it-NVFP4, unsloth/Qwen3.6-35B-A3B-NVFP4 ~14–22 GB L4 / A10G (24 GB) β€” see Unsloth variants & NVFP4

Practical tips:

  • A 35 B MoE with only ~3 B active params generates fast once loaded, but it still needs the full ~70 GB of weights resident β€” don't try to squeeze it onto a 24/48 GB card without quantization.
  • To run the 35 B model on cheaper hardware, use its Unsloth 4-bit or NVFP4 variant β€” a first-class option in the UI, see the next section.
  • Set the Space sleep timeout generously or use "sleep after 1 h" β€” GPU Spaces bill while awake, and evaluation runs are bursty.
  • ZeroGPU Spaces are not supported by this code path as written (model loading happens outside a @spaces.GPU function); use a dedicated GPU instance for the local backend.

Disk is not a constraint on any tier: test-set caches total well under 1 GB; model weights for the local backend go to the ephemeral disk, which is sized to the hardware tier.

Unsloth variants & NVFP4

The Model variant control in the Run tab swaps the base model for its pre-quantized Unsloth checkpoint. The repo name is resolved automatically (google/gemma-4-e4b-it β†’ unsloth/gemma-4-e4b-it-unsloth-bnb-4bit / unsloth/gemma-4-e4b-it-NVFP4, same for google/gemma-4-26b-it, Qwen/Qwen3.6-35B-A3B, and any other slug you type); several naming conventions are checked against the Hub at run time, and the leaderboard records the exact resolved repo so quantized and full-precision runs rank side by side β€” handy for showing how much quality 4-bit actually costs on your language pairs.

Variant What it is Backend to select Hardware
Unsloth dynamic 4-bit (bnb) bitsandbytes 4-bit with Unsloth's per-layer dynamic quantization (sensitive layers kept in higher precision) Local GPU (transformers) β€” uncomment bitsandbytes in requirements.txt Works on any CUDA GPU, T4 and up; 35 B MoE fits in ~20 GB (L4/A10G)
Unsloth NVFP4 NVIDIA's block-scaled FP4 format (per-16-value FP8 scales), near-bf16 quality at 4-bit size Local GPU (vLLM) β€” uncomment vllm in requirements.txt. vLLM auto-detects the quantization from the checkpoint; transformers has no NVFP4 path, and the app blocks that combination with a clear error Full FP4 tensor-core speedups need a Blackwell GPU (B200-class). On Hopper/Ada instances vLLM falls back to weight-only 4-bit kernels: you keep the ~4Γ— memory saving, not the compute speedup

Extra speed on unsloth checkpoints with the transformers backend: uncomment unsloth in requirements.txt and set a Space variable USE_UNSLOTH=1 β€” the app then loads models through FastLanguageModel and silently falls back to plain transformers if the package is missing.

Two caveats worth knowing:

  • The Inference Providers API backend serves base models only β€” providers generally don't host unsloth/* quantized repos, and the app warns if you try that combination.
  • If a variant repo doesn't exist for your model, the run stops with the list of names it tried; you can always paste an exact unsloth/... repo id into the model field with variant set to Base model.

Notes

  • The first WMT25 run downloads a ~128 MB JSONL once, reduces it to a compact cache, and deletes the original.
  • BLEU tokenization adapts to the target language (zh for Chinese, char for Japanese/Korean/Thai, 13a otherwise). ChrF is language-agnostic.
  • "Max segments per language pair" caps how much of the test set is used β€” keep it low (~100) for smoke tests, raise it for reportable numbers.
  • Re-running the same model Γ— dataset Γ— language pair overwrites the previous entry, so the board never shows stale duplicates.