Zonos v0.1

State of the art text-to-speech model [model]. [blog], [Zyphra Audio (hosted service)]

Unleashed

Use this space to generate long-form speech up to around ~2 minutes in length. To generate an unlimited length, clone this space and run it locally.

Tips

  • When providing prefix audio, include the text of the prefix audio in your speech text to ensure a smooth transition.
  • The appropriate range of Speaking Rate and Pitch STD are highly dependent on the speaker audio. Start with the defaults and adjust as needed.
  • Emotion sliders do not completely function intuitively, and require some experimentation to get the desired effect.
Zonos Model Variant
Language

Long-Form Parameters

1 300
0 1
0 1
0 1

Generation Parameters

1 5
0 1

Conditioning Parameters

All of these types of conditioning are optional and can be disabled.

'Speaker Noised' is a conditioning value that the model understands, not a processing step. Check this box if your input audio is noisy.

-1200 1200
0 1
0 1
0 1
0 1
0 1
0 1
0 1
0 1
1 5
0.5 0.8
0 22050
0 300
5 30