Instructions to use transcriber-app/parakeet-tdt-0.6b-v3-mlx-6bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use transcriber-app/parakeet-tdt-0.6b-v3-mlx-6bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download transcriber-app/parakeet-tdt-0.6b-v3-mlx-6bit --local-dir parakeet-tdt-0.6b-v3-mlx-6bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
parakeet-tdt-0.6b-v3 โ MLX 6-bit
NVIDIA Parakeet TDT 0.6B v3, in the MLX layout of mlx-community/parakeet-tdt-0.6b-v3, quantized to 6-bit for on-device speech-to-text on Apple silicon. Multilingual: 25 European languages.
What changed from the source
This is a redistribution of a modified model. Starting from the fp32 weights of
mlx-community/parakeet-tdt-0.6b-v3:
- every
LinearandEmbeddinglayer (attention and feed-forward projections, the pre-encode and joint projections, the prediction-network embedding โ 221 modules) is quantized withmlx.nn.quantizeto 6-bit, group size 64 (affine); - every other tensor (convolutions, LSTM, normalization) is stored as bfloat16;
config.jsongains"quantization": {"group_size": 64, "bits": 6}.
Tensor names are unchanged from the source checkpoint, so loaders that read the
mlx-community layout and honour the quantization block (quantizing the modules whose
weights arrive with .scales) load it directly.
Files
model.safetensorsโ quantized weights (608 MB, vs 2.5 GB fp32)config.jsonโ the NeMo config with thequantizationblockvocab.txt,tokenizer.model,tokenizer.vocabโ unchanged from the source
Accuracy
Measured on 1,000 Russian clips (500 FLEURS, 500 GOLOS), greedy decode, on an M4 Mac:
| build | download | FLEURS WER | GOLOS WER | speed |
|---|---|---|---|---|
fp32 (mlx-community/parakeet-tdt-0.6b-v3) |
2.5 GB | 5.57 % | 3.42 % | 64ร |
8-bit (transcriber-app/parakeet-tdt-0.6b-v3-mlx-8bit) |
744 MB | 5.67 % | 3.46 % | 90ร |
6-bit (transcriber-app/parakeet-tdt-0.6b-v3-mlx-6bit) |
608 MB | 5.61 % | 3.38 % | 88ร |
| 5-bit (not published) | 540 MB | 5.93 % | 3.50 % | 87ร |
4-bit (transcriber-app/parakeet-tdt-0.6b-v3-mlx-4bit) |
472 MB | 5.95 % | 3.67 % | 90ร |
6-bit keeps fp32 accuracy โ level with 8-bit, within measurement noise โ at the smallest size that does. 5 and 4 bits lose about a third of a WER point.
Credits and licence
- Model: NVIDIA, parakeet-tdt-0.6b-v3 โ CC-BY-4.0.
- MLX conversion: mlx-community (via parakeet-mlx) โ CC-BY-4.0.
- Quantization: this repository, released under the same CC-BY-4.0 licence.
- Downloads last month
- 17
6-bit
Model tree for transcriber-app/parakeet-tdt-0.6b-v3-mlx-6bit
Base model
nvidia/parakeet-tdt-0.6b-v3