Overview
About this model
ViiTorVoice-NAR is a non-autoregressive speech generation model for voice cloning, local speech editing, and emotion / paralinguistic speech control.
View model card on HuggingFaceSpecifications
Model details
- Modality
- speech_clone
Modalities: text, audio → audio