Bangla Parler-TTS (fine-tuned from Indic Parler-TTS)

Bangla (Bengali) text-to-speech model, fine-tuned from ai4bharat/indic-parler-tts on a curated, quality-filtered subset of the SUBAK.KO Bangla speech corpus. Like the base model, it generates speech conditioned on a text prompt (what to say, in Bangla) and a natural-language description (how to say it: gender, pace, clarity).

Training data

The training set was built by curating SUBAK.KO (an ASR corpus) into a TTS-suitable dataset:

  • per-clip quality scoring (SNR proxy, clipping, duration 1-15 s)
  • kept only the cleanest clips (SNR floor + per-speaker balancing), ~10.5k train clips / 800 validation / 800 test
  • style descriptions rebuilt to be name-free and attribute-based, with speaker gender taken from corpus speaker codes (ground truth), e.g. "A Bengali female speaker speaks at a moderate pace, in a very clear sound."

Training setup

  • Base: ai4bharat/indic-parler-tts (text encoder frozen; decoder + LM head trained)
  • 8 epochs, ~2.6k optimizer steps, effective batch 32 (4 x 8 grad-accum)
  • LR 1.5e-5, cosine schedule, 200 warmup steps, weight decay 0.01
  • bfloat16, gradient checkpointing, SDPA attention, single A100-40GB

Installation

Using Parler-TTS is as simple as "Bengali". Simply install the library once:

pip install git+https://github.com/huggingface/parler-tts.git

Usage

import torch
from parler_tts import ParlerTTSForConditionalGeneration
from transformers import AutoTokenizer
import soundfile as sf

device = "cuda:0" if torch.cuda.is_available() else "cpu"

model = ParlerTTSForConditionalGeneration.from_pretrained("Tanmoyshome/bn-parler-tts").to(device)
tokenizer = AutoTokenizer.from_pretrained("Tanmoyshome/bn-parler-tts")
description_tokenizer = AutoTokenizer.from_pretrained(model.config.text_encoder._name_or_path)

prompt = "রোগ যদি ধরতে চাও রোগীর ইতিহাস জেনে নাও তাতে কিছু তথ্য নাও সেই মত পথ্য দাও"
description = ("A Bengali female speaker speaks at a moderate pace "
               "in a very clear environment.")

description_input_ids = description_tokenizer(description, return_tensors="pt").to(device)
prompt_input_ids = tokenizer(prompt, return_tensors="pt").to(device)

generation = model.generate(
    input_ids=description_input_ids.input_ids,
    attention_mask=description_input_ids.attention_mask,
    prompt_input_ids=prompt_input_ids.input_ids,
    prompt_attention_mask=prompt_input_ids.attention_mask,
)
audio_arr = generation.cpu().numpy().squeeze()
sf.write("bn_tts_out.wav", audio_arr, model.config.sampling_rate)

Requires pip install git+https://github.com/huggingface/parler-tts.git soundfile.

Prompting tips

Descriptions in training follow a fixed attribute style; you will get the most reliable control by mirroring it:

  • gender: "A Bengali male speaker" / "A Bengali female speaker"
  • pace: "at a slow / moderate / fast pace", "speaks quickly / slowly"
  • clarity: "in a very clear recording", "in a clear environment"

Limitations

  • Trained on multi-speaker read speech; voice identity may vary between generations (no named-voice control).
  • Descriptions control gender / pace / clarity; fine-grained emotion or named speakers are not supported.
  • Bangla only; other Indic languages should use the base model.

Base model / attribution

Fine-tuned from ai4bharat/indic-parler-tts (AI4Bharat), built on the Parler-TTS architecture. Please also refer to the base model card for architecture details and broader usage guidance.

Downloads last month
57
Safetensors
Model size
0.9B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Tanmoyshome/bn-parler-tts

Finetuned
(5)
this model

Dataset used to train Tanmoyshome/bn-parler-tts