higgs-audio-v3-1.7b-stt-v2

This is bosonai/higgs-audio-understanding-v3-1.7b (checkpoint-65000) with a LoRA (rank=64, alpha=128) merged into the base weights. The LoRA was trained on a curated ASR mix (AMI IHM train, Earnings22 skipped, GigaSpeech XS train, LibriSpeech train.100 + train.500, SpgiSpeech S train, TEDLium train, VoxPopuli-en train). No ESB test data was used in training.

Reported result (Open ASR Leaderboard methodology)

Dataset WER (%)
AMI 10.03
Earnings-22 8.95
GigaSpeech 8.16
LibriSpeech clean 1.39
LibriSpeech other 2.80
SPGISpeech 3.76
TED-LIUM 2.76
VoxPopuli 6.07
Macro avg 5.49

Evaluated with the official run_eval_higgs_audio.py script, max_new_tokens=1024, greedy decoding, Whisper English normalizer.

Usage

from transformers import AutoModel, AutoTokenizer
import torch, numpy as np, soundfile as sf

model = AutoModel.from_pretrained(
    "bosonai/higgs-audio-v3-stt-v2",
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
    attn_implementation="eager",
    device_map="cuda:0",
)
tok = AutoTokenizer.from_pretrained("bosonai/higgs-audio-v3-stt-v2")
model.audio_out_bos_token_id = tok.convert_tokens_to_ids("<|audio_out_bos|>")
model.audio_eos_token_id     = tok.convert_tokens_to_ids("<|audio_eos|>")

# Load bundled transcribe.py
from transformers.utils import cached_file
import runpy, os, sys
path = cached_file("bosonai/higgs-audio-v3-stt-v2", "transcribe.py")
for f in ["higgs_audio_collator.py","modeling_higgs_audio_xcodec.py","utils.py","common.py","configuration_higgs_audio.py"]:
    cached_file("bosonai/higgs-audio-v3-stt-v2", f)
sys.path.insert(0, os.path.dirname(path))
transcribe_batch = runpy.run_path(path)["transcribe_batch"]

audio, sr = sf.read("example.wav")
print(transcribe_batch(model, tok, [audio.astype(np.float32)], sample_rates=sr))

Reproducing the benchmark

git clone https://github.com/huggingface/open_asr_leaderboard
cd open_asr_leaderboard
python run_eval_higgs_audio.py \
    --model_id bosonai/higgs-audio-v3-stt-v2 \
    --dataset_path hf-audio/open-asr-leaderboard-sorted \
    --dataset ami --split test --device 0 --batch_size 4
Downloads last month
186
Safetensors
Model size
3B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for bosonai/higgs-audio-v3-stt-v2

Finetuned
(1)
this model

Space using bosonai/higgs-audio-v3-stt-v2 1

Collection including bosonai/higgs-audio-v3-stt-v2