Higgs Audio STT
Collection
Audio Transcription Model โข 4 items โข Updated โข 5
How to use bosonai/higgs-audio-v3-stt-v2 with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("automatic-speech-recognition", model="bosonai/higgs-audio-v3-stt-v2", trust_remote_code=True) # Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("bosonai/higgs-audio-v3-stt-v2", trust_remote_code=True, device_map="auto")This is bosonai/higgs-audio-understanding-v3-1.7b (checkpoint-65000) with a
LoRA (rank=64, alpha=128) merged into the base weights. The LoRA was trained
on a curated ASR mix (AMI IHM train, Earnings22 skipped, GigaSpeech XS train,
LibriSpeech train.100 + train.500, SpgiSpeech S train, TEDLium train, VoxPopuli-en
train). No ESB test data was used in training.
| Dataset | WER (%) |
|---|---|
| AMI | 10.03 |
| Earnings-22 | 8.95 |
| GigaSpeech | 8.16 |
| LibriSpeech clean | 1.39 |
| LibriSpeech other | 2.80 |
| SPGISpeech | 3.76 |
| TED-LIUM | 2.76 |
| VoxPopuli | 6.07 |
| Macro avg | 5.49 |
Evaluated with the official
run_eval_higgs_audio.py
script, max_new_tokens=1024, greedy decoding, Whisper English normalizer.
from transformers import AutoModel, AutoTokenizer
import torch, numpy as np, soundfile as sf
model = AutoModel.from_pretrained(
"bosonai/higgs-audio-v3-stt-v2",
torch_dtype=torch.bfloat16,
trust_remote_code=True,
attn_implementation="eager",
device_map="cuda:0",
)
tok = AutoTokenizer.from_pretrained("bosonai/higgs-audio-v3-stt-v2")
model.audio_out_bos_token_id = tok.convert_tokens_to_ids("<|audio_out_bos|>")
model.audio_eos_token_id = tok.convert_tokens_to_ids("<|audio_eos|>")
# Load bundled transcribe.py
from transformers.utils import cached_file
import runpy, os, sys
path = cached_file("bosonai/higgs-audio-v3-stt-v2", "transcribe.py")
for f in ["higgs_audio_collator.py","modeling_higgs_audio_xcodec.py","utils.py","common.py","configuration_higgs_audio.py"]:
cached_file("bosonai/higgs-audio-v3-stt-v2", f)
sys.path.insert(0, os.path.dirname(path))
transcribe_batch = runpy.run_path(path)["transcribe_batch"]
audio, sr = sf.read("example.wav")
print(transcribe_batch(model, tok, [audio.astype(np.float32)], sample_rates=sr))
git clone https://github.com/huggingface/open_asr_leaderboard
cd open_asr_leaderboard
python run_eval_higgs_audio.py \
--model_id bosonai/higgs-audio-v3-stt-v2 \
--dataset_path hf-audio/open-asr-leaderboard-sorted \
--dataset ami --split test --device 0 --batch_size 4
Base model
bosonai/higgs-audio-v3-stt