Yunmo-Next Tokenizer

A 65,536-entry byte-level BPE tokenizer for Traditional Chinese, English, code and mathematics. It is the tokenizer of Yunmo-Next-1B-Base and ships as one standard Hugging Face tokenizer.json; no custom tokenization code is needed.

Specification

Vocabulary 65,536 IDs
Text IDs 0–65,279 (60,280 phase-1 merges/bytes + 5,000 phase-2 superwords)
Protocol IDs 65,280–65,535 (256); `<
Model byte-level BPE with byte fallback
Lossless any UTF-8 input round-trips; there is no UNK token
Normalization none: no Unicode normalization, script conversion, case folding or whitespace cleanup
Digits every decimal digit is its own token

Compression

Bytes of UTF-8 text per token on identical samples (higher means fewer tokens):

Tokenizer Vocabulary Trad. Chinese web Trad. Chinese Wikipedia Taiwan statutes Classical Chinese English (edu) Code Math
Yunmo-Next 65,536 4.35 4.18 4.98 4.21 5.04 3.32 3.85
TAIDE (Llama 3.1) 188,256 4.27 4.01 4.52 3.83 4.75 3.92 3.60
Qwen3.6 248,070 4.08 3.59 3.79 3.66 4.60 3.56 3.38
DeepSeek-V3 128,815 3.97 3.58 3.77 3.56 4.79 3.57 3.68
Breeze 61,875 4.01 3.62 3.94 3.75 4.19 2.93 3.09
Gemma 3 262,145 3.72 3.42 3.48 3.43 4.66 3.22 3.38
Llama 3.1 128,256 3.14 3.03 3.18 2.99 4.75 3.92 3.60

The measurement covered 18 tokenizers on 12 domains (the table shows a subset; several pairs share a vocabulary, e.g. Gemma 3/4, Qwen3/Qwen2.5 and Llama 3.1/Llama-3-Taiwan). Despite a vocabulary 2–4× smaller than most of them, Yunmo-Next ranks first on 9 domains: Traditional Chinese web, Wikipedia, classical texts, Taiwan statutes and court judgments, Zhuyin, Simplified Chinese web, English and mathematics. It is second on Taiwan patents (TAIDE: 4.01 vs 3.93) and third on general Simplified Chinese (Qwen3.6: 5.04, DeepSeek-V3: 4.96, Yunmo-Next: 4.68). Code is its weak domain: it ranks 9th, with 8 tokenizers including Llama 3.1, TAIDE, Qwen3.6 and DeepSeek-V3 encoding code more compactly.

Caveat. These domain samples come from the same source families used to train this tokenizer (Traditional Chinese web, Wikipedia, statutes, court judgments, patents, Zhuyin), and the comparison does not guarantee that the sampled documents were excluded from its training corpus. The domains are therefore in-distribution for Yunmo-Next and not for the other tokenizers, which favors Yunmo-Next. The only strictly held-out measurement is the comparison with the phase-1-only control below.

How it was built

Corpus allocated by target exposure

The tokenizer corpus is sized so that each domain's share of training tokens matches its planned share of pretraining tokens. With target token share q_d for domain d and a reference tokenizer producing τ_d tokens per byte in that domain, the byte budget is

b_d = (q_d / τ_d) / Σ_j (q_j / τ_j)

The final corpus is 101,241 documents / 200.2 MB drawn from 21 source families, plus a fixed held-out set of 10,166 documents / 20.1 MB that was selected by source and exact text hash before any candidate was scored and removed from all training sources. After training, budgets recomputed with the new tokenizer moved by at most 0.73 percentage points, under the 1-point threshold that would have triggered a rebuild.

Corpus share (UTF-8 bytes)
English 34.76%
Converted Traditional Chinese 16.70%
Native Traditional Chinese, general 12.66%
Code 12.36%
Mathematics 10.73%
Original Simplified Chinese 7.32%
Native Traditional Chinese, official/professional 2.75%
Traditional Chinese instructions 2.34%
Zhuyin (Bopomofo) 0.39%

Two-phase BPE

  1. Phase 1 learns 60,280 text entries inside standard byte-level pre-tokenization boundaries, with decimal digits isolated.
  2. Phase 2 encodes the raw documents with the phase-1 model and learns additional merges over the resulting token sequences, so new entries may cross pre-tokenization boundaries. Digits and newlines terminate a candidate, and each new entry spans at most 8 phase-1 tokens. The build processed 50,426,460 phase-1 tokens and kept 5,000 superwords.

The two merge lists are then serialized as a single standard BPE model, so the result loads with the tokenizers library alone. The approach follows the cross-boundary merging idea of SuperBPE.

Selection

The candidate (two-phase, 65,280 text entries) was compared with a control trained on the same documents, held-out set, vocabulary size and protocol. The control does not isolate digits, so it served only as a comparison and could not be selected as a fallback.

Check Result
Blind held-out tokens (10,166 documents) 4,555,786 vs 4,872,873 for the control (−6.5%); fewer tokens on all 8 strata, from −1.3% (Simplified Chinese) to −15.3% (Zhuyin) and −12.3% (English)
Round trip / unknown tokens 100% / 0
Encoding speed (median of 3 rounds) 5.37 MB/s vs 3.46 MB/s for the control
2-layer, 4.3M-parameter screens, 1,048,576 tokens, seeds 17 and 29 bits per byte −0.060 and −0.063 vs control
12-layer, 91.9M-parameter Yunmo-architecture model, 32,006,144 tokens, seed 17 bits per byte 1.987 vs 1.996 (−0.0095)

Limitations of this selection, stated as measured:

  • The final downstream comparison is a single seed. The two arms' capped validation sets ended up with different byte totals (2.19 MB vs 2.16 MB), so the −0.0095 is not a document-paired difference and no significance claim is made.
  • At an equal token budget the candidate saw about 7% more raw text, which is part of the measured effect; the experiment tests the tokenizer as a whole, not digit isolation, the source prior or superwords individually.
  • A preregistered gate on a public Traditional Chinese lexicon benchmark (Pangolin, lexicon cost below 2.48) was failed: the candidate scored 2.57 (control 2.62). The gate hierarchy was then reconciled to rank downstream model quality above that benchmark, and the failure is kept on record rather than relabeled as a pass.
  • No near-duplicate scan was run between the tokenizer training corpus and its held-out set; they are separated by exact content hash only.

Protocol IDs

The 256 protocol IDs encode sequence and message boundaries, roles, separate reasoning (analysis) and answer (final) channels, tool calls, document boundaries, fill-in-the-middle and reserved multimodal slots. The mapping is in protocol_abi_v2.json. Reserving them does not mean a model has learned the formats.

The protocol tokens are registered special tokens, so the literal strings match them: encoding the text <|eos|> returns ID 65,281 even with add_special_tokens=False. If you encode untrusted text, check for IDs ≥ 65,280. The Yunmo-Next runtime does this and rejects user text that would encode to a protocol ID; only its chat formatter inserts protocol tokens.

Usage

from tokenizers import Tokenizer

tok = Tokenizer.from_file("tokenizer.json")
ids = tok.encode("雲墨是一個以繁體中文為主的語言模型。", add_special_tokens=False).ids
print(len(ids))          # 8
print(tok.decode(ids))   # round-trips exactly

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support