Instructions to use suideepmax/canary-qwen-2.5b-atc-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use suideepmax/canary-qwen-2.5b-atc-lora with NeMo:
# tag did not correspond to a valid NeMo domain.
- Notebooks
- Google Colab
- Kaggle
Canary-Qwen-2.5B Fine-Tuned for ATC ASR (LoRA + Regularization, v3)
Fine-tuned nvidia/canary-qwen-2.5b on the UWB-ATCC corpus for Air Traffic Control speech recognition using LoRA adaptation plus SpecAugment/dropout/weight-decay regularization.
Correction (2026-09-08): this repo's README previously mislabeled the hosted weights as the plain (un-regularized) LoRA baseline, citing 23.32% WER. That was incorrect โ the actual uploaded consolidated_model.pt is the v3 (regularized) model. This has been verified directly: real inference against these exact weights reproduces 20.70% WER, bit-for-bit identical to v3_results.json in this repo and to an independent evaluation of the separately-preserved local v3 checkpoint. See v3_results.json, learning_curve_v3.json, and training_config.yaml in this repo (all of which were already correct) for the full result and config record.
Results
| Model | Params Trained | WER |
|---|---|---|
| Canary-Qwen (zero-shot) | 0 | 81.49% |
| Canary-Qwen (LoRA, no regularization) | 27.8M (0.97%) | 23.32% |
| Canary-Qwen (LoRA + regularization, v3 โ this checkpoint) | 27.8M (0.97%) | 20.70% |
| W2V2 Large (no LM) | 317M (100%) | 14.54% |
| W2V2 Large (with KenLM) | 317M (100%) | 12.69% |
Training
- Dataset: UWB-ATCC (Prague Airport ATC, 11,543 train / 2,886 test utterances)
- Steps: 10,000 | LR: 5e-4 | Warmup: 1,000
- LoRA: r=128, alpha=256, dropout=0.1, targets=[q_proj, v_proj]
- Regularization: SpecAugment (freq_masks=2, time_masks=10), weight_decay=1e-2
- Strategy: FSDP across 4x RTX 2080 Ti | Precision: fp16-true (eps=1e-4)
- Framework: NVIDIA NeMo 2.8.0rc0
- Verified training time: ~21 hours (a previously-documented ~5.3h figure for this configuration was found to be incorrect and corrected 2026-09-08)
Learning Curve (500-utterance test subset per checkpoint)
| Step | WER |
|---|---|
| 0 | 81.49% |
| 500 | 39.51% |
| 1,000 | 32.68% |
| 2,000 | 27.28% |
| 3,000 | 25.08% |
| 5,000 | 23.00% |
| 7,500 | 23.81% |
| 10,000 | 22.30% (full 2,886-utterance test set at step 10,000: 20.70%) |
Unlike the un-regularized baseline (which plateaus around 24.5% WER after step 5,000), this model keeps improving throughout training โ the regularization prevents overfitting on the small (10.5-hour) fine-tuning set.
Usage
from nemo.collections.speechlm2.models import SALM
import torch
model = SALM.from_pretrained('nvidia/canary-qwen-2.5b')
state = torch.load('consolidated_model.pt', map_location='cpu')
model.load_state_dict(state, strict=False)
model.cuda().eval()
answer_ids = model.generate(
prompts=[[{
'role': 'user',
'content': f'Transcribe the following: {model.audio_locator_tag}',
'audio': ['atc_audio.wav']
}]],
max_new_tokens=128,
)
print(model.tokenizer.ids_to_text(answer_ids[0].cpu()))
Part of
Pilot-to-ATC Research โ Comparative evaluation of W2V2 vs Canary-Qwen for ATC domain ASR.
- Downloads last month
- 71