Petal
Collection
Voice transcription models for the Petal Mac app: Voxtral Mini and Cohere Transcribe conversions for Apple silicon. • 5 items • Updated
Hybrid CoreML encoder + ONNX decoder for Cohere Transcribe, optimized for Apple Silicon inference. The encoder runs on the Neural Engine via CoreML, the decoder runs with ONNX Runtime and KV cache on CPU.
Cohere Transcribe is a 2B-parameter encoder-decoder ASR model that holds #1 on the Open ASR Leaderboard with 5.42% average WER — beating Whisper Large v3, ElevenLabs Scribe v2, and Qwen3-ASR-1.7B.
| Property | Value |
|---|---|
| Total Parameters | 2.07B |
| Encoder | CoreML FP16 (3.5 GB) |
| Decoder | ONNX q4f16 (98 MB) |
| Projection | Float32 (5 MB) |
| Total Download | ~3.6 GB |
| License | Apache 2.0 |
coreml/
cohere_encoder.mlmodelc/ # CoreML encoder (ANE-optimized, FP16)
onnx/
decoder_model_merged_q4f16.onnx # ONNX decoder header
decoder_model_merged_q4f16.onnx_data # ONNX decoder weights (q4f16)
config.json
generation_config.json
preprocessor_config.json
tokenizer.json
tokenizer_config.json
decoder_proj_weight.bin # Encoder→decoder projection (1280→1024)
decoder_proj_bias.bin
This model is designed for use with Petal, a macOS menu bar app for local-first audio transcription.
Architecture:
Performance on Apple Silicon:
Apache 2.0 — original model by Cohere.
Base model
CohereLabs/cohere-transcribe-03-2026