Canary-Qwen-2.5B Fine-Tuned for ATC ASR (LoRA + Regularization, v3)

Fine-tuned nvidia/canary-qwen-2.5b on the UWB-ATCC corpus for Air Traffic Control speech recognition using LoRA adaptation plus SpecAugment/dropout/weight-decay regularization.

Correction (2026-09-08): this repo's README previously mislabeled the hosted weights as the plain (un-regularized) LoRA baseline, citing 23.32% WER. That was incorrect โ€” the actual uploaded consolidated_model.pt is the v3 (regularized) model. This has been verified directly: real inference against these exact weights reproduces 20.70% WER, bit-for-bit identical to v3_results.json in this repo and to an independent evaluation of the separately-preserved local v3 checkpoint. See v3_results.json, learning_curve_v3.json, and training_config.yaml in this repo (all of which were already correct) for the full result and config record.

Results

Model Params Trained WER
Canary-Qwen (zero-shot) 0 81.49%
Canary-Qwen (LoRA, no regularization) 27.8M (0.97%) 23.32%
Canary-Qwen (LoRA + regularization, v3 โ€” this checkpoint) 27.8M (0.97%) 20.70%
W2V2 Large (no LM) 317M (100%) 14.54%
W2V2 Large (with KenLM) 317M (100%) 12.69%

Training

  • Dataset: UWB-ATCC (Prague Airport ATC, 11,543 train / 2,886 test utterances)
  • Steps: 10,000 | LR: 5e-4 | Warmup: 1,000
  • LoRA: r=128, alpha=256, dropout=0.1, targets=[q_proj, v_proj]
  • Regularization: SpecAugment (freq_masks=2, time_masks=10), weight_decay=1e-2
  • Strategy: FSDP across 4x RTX 2080 Ti | Precision: fp16-true (eps=1e-4)
  • Framework: NVIDIA NeMo 2.8.0rc0
  • Verified training time: ~21 hours (a previously-documented ~5.3h figure for this configuration was found to be incorrect and corrected 2026-09-08)

Learning Curve (500-utterance test subset per checkpoint)

Step WER
0 81.49%
500 39.51%
1,000 32.68%
2,000 27.28%
3,000 25.08%
5,000 23.00%
7,500 23.81%
10,000 22.30% (full 2,886-utterance test set at step 10,000: 20.70%)

Unlike the un-regularized baseline (which plateaus around 24.5% WER after step 5,000), this model keeps improving throughout training โ€” the regularization prevents overfitting on the small (10.5-hour) fine-tuning set.

Usage

from nemo.collections.speechlm2.models import SALM
import torch

model = SALM.from_pretrained('nvidia/canary-qwen-2.5b')
state = torch.load('consolidated_model.pt', map_location='cpu')
model.load_state_dict(state, strict=False)
model.cuda().eval()

answer_ids = model.generate(
    prompts=[[{
        'role': 'user',
        'content': f'Transcribe the following: {model.audio_locator_tag}',
        'audio': ['atc_audio.wav']
    }]],
    max_new_tokens=128,
)
print(model.tokenizer.ids_to_text(answer_ids[0].cpu()))

Part of

Pilot-to-ATC Research โ€” Comparative evaluation of W2V2 vs Canary-Qwen for ATC domain ASR.

Downloads last month
71
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for suideepmax/canary-qwen-2.5b-atc-lora

Finetuned
Qwen/Qwen3-1.7B
Adapter
(2)
this model