parakeet-tdt-0.6b-v3 โ€” MLX 6-bit

NVIDIA Parakeet TDT 0.6B v3, in the MLX layout of mlx-community/parakeet-tdt-0.6b-v3, quantized to 6-bit for on-device speech-to-text on Apple silicon. Multilingual: 25 European languages.

What changed from the source

This is a redistribution of a modified model. Starting from the fp32 weights of mlx-community/parakeet-tdt-0.6b-v3:

  • every Linear and Embedding layer (attention and feed-forward projections, the pre-encode and joint projections, the prediction-network embedding โ€” 221 modules) is quantized with mlx.nn.quantize to 6-bit, group size 64 (affine);
  • every other tensor (convolutions, LSTM, normalization) is stored as bfloat16;
  • config.json gains "quantization": {"group_size": 64, "bits": 6}.

Tensor names are unchanged from the source checkpoint, so loaders that read the mlx-community layout and honour the quantization block (quantizing the modules whose weights arrive with .scales) load it directly.

Files

  • model.safetensors โ€” quantized weights (608 MB, vs 2.5 GB fp32)
  • config.json โ€” the NeMo config with the quantization block
  • vocab.txt, tokenizer.model, tokenizer.vocab โ€” unchanged from the source

Accuracy

Measured on 1,000 Russian clips (500 FLEURS, 500 GOLOS), greedy decode, on an M4 Mac:

build download FLEURS WER GOLOS WER speed
fp32 (mlx-community/parakeet-tdt-0.6b-v3) 2.5 GB 5.57 % 3.42 % 64ร—
8-bit (transcriber-app/parakeet-tdt-0.6b-v3-mlx-8bit) 744 MB 5.67 % 3.46 % 90ร—
6-bit (transcriber-app/parakeet-tdt-0.6b-v3-mlx-6bit) 608 MB 5.61 % 3.38 % 88ร—
5-bit (not published) 540 MB 5.93 % 3.50 % 87ร—
4-bit (transcriber-app/parakeet-tdt-0.6b-v3-mlx-4bit) 472 MB 5.95 % 3.67 % 90ร—

6-bit keeps fp32 accuracy โ€” level with 8-bit, within measurement noise โ€” at the smallest size that does. 5 and 4 bits lose about a third of a WER point.

Credits and licence

Other builds: 8-bit, 4-bit.

Downloads last month
17
Safetensors
Model size
0.6B params
Tensor type
BF16
ยท
U32
ยท
MLX
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for transcriber-app/parakeet-tdt-0.6b-v3-mlx-6bit

Finetuned
(98)
this model