Speech-to-text for Apple silicon (Core ML)
Collection
Audio to text models Whisper, Parakeet, Nemotron and Cohere Transcribe as compiled Core ML models for the Neural Engine. • 14 items • Updated
openai/whisper-small for Apple silicon: a transcribe model package of compiled Core ML models, mostly on the Neural Engine. The transcribe-model-darwin Rust crate loads it (AnyModel::load).
| Package | whisper-small: format 3, architecture whisper |
| Source weights | https://openaipublic.azureedge.net/main/whisper/models/9ecf779972d90ba49c06d968637d720dd632c55bbf19d441fb42bf17a411e794/small.pt |
| License | MIT (LICENSE, NOTICE) |
| Languages | 99: af, am, ar, as, az, ba, be, bg, bn, bo, br, bs, ca, cs, cy, da, de, el, en, es, et, eu, fa, fi, fo, fr, gl, gu, ha, haw, he, hi, hr, ht, hu, hy, id, is, it, ja, jw, ka, kk, km, kn, ko, la, lb, ln, lo, lt, lv, mg, mi, mk, ml, mn, mr, ms, mt, my, ne, nl, nn, no, oc, pa, pl, ps, pt, ro, ru, sa, sd, si, sk, sl, sn, so, sq, sr, su, sv, sw, ta, te, tg, th, tk, tl, tr, tt, uk, ur, uz, vi, yi, yo, zh |
| Size | 175 MiB in 17 files |
| Requires | macOS 15.0 on Apple silicon |
| Built with | transcribe-models 0.0.1 at commit f2d171f91efaba91053e99586c694d82c4238479, coremltools 9.0 |
| Model | Compute units | Functions | Weights |
|---|---|---|---|
frontend |
CPU | audio_480000 |
not compressed |
encoder |
CPU and Neural Engine | audio_480000 |
6-bit k-means palettes per tensor |
decoder |
CPU and Neural Engine | one | 6-bit k-means palettes per tensor |
Download manifest.json and every file it lists from https://huggingface.co/yorganci/whisper-small-6bit-coreml/resolve/<revision>/<path>, checking each file's size and SHA-256 against the manifest, then load the directory:
use transcribe_core::{Audio, TranscribeOptions, Transcriber};
use transcribe_model_darwin::{AnyModel, LoadOptions};
let mut model = AnyModel::load("whisper-small", LoadOptions::default())?;
let transcript = model.transcribe(Audio::new(&samples, 16_000), &TranscribeOptions::default())?;
The first load on a machine compiles the models for the Neural Engine (seconds to a couple of minutes); later loads are fast.
Error rates of this build with transcribe-eval (Apple M4; characters for Japanese, words otherwise):
| Set | Error rate | Substitutions | Deletions | Insertions | Reference units |
|---|---|---|---|---|---|
| LibriSpeech test-clean, 20 utterances | 2.20 % | 10 | 0 | 1 | 500 |
| LibriSpeech test-clean, 62 s clip, long-form with timestamps | 2.82 % | 2 | 1 | 1 | 142 |
| FLEURS de, 10 utterances | 12.32 % | 22 | 1 | 3 | 211 |
| FLEURS fr, 10 utterances | 23.02 % | 30 | 3 | 25 | 252 |
| FLEURS es, 10 utterances | 7.42 % | 11 | 3 | 5 | 256 |
| FLEURS ja, 10 utterances | 9.96 % | 35 | 6 | 5 | 462 |
whisper-small
OpenAI Whisper small (https://github.com/openai/whisper), Copyright (c) 2022 OpenAI, licensed under the MIT License.
Source: https://openaipublic.azureedge.net/main/whisper/models/9ecf779972d90ba49c06d968637d720dd632c55bbf19d441fb42bf17a411e794/small.pt
License: MIT (LICENSE)
Changes: this package is a modified form of the source weights. transcribe-models
(https://github.com/atahanyorganci/transcribe) ported the model to PyTorch and converted
it to Core ML with coremltools 9.0, in float16 (the front end's constants in float32),
with these weight compressions:
- frontend: not compressed
- encoder: 6-bit k-means palettes per tensor
- decoder: 6-bit k-means palettes per tensor
Base model
openai/whisper-small