The AI voice synthesis model Zonos has been released, supporting voice cloning.

Zyphra has released the first beta version of its speech synthesis AI model, Zonos, under the Apache 2.0 license. The toolkit offered with the model supports a voice cloning feature, allowing speech to be synthesized in the desired voice, for which the model only needs a 30-second reference recording of the speaker's voice. Synthesis is supported in English, Japanese, Chinese, French, and German.

The model encompasses 1.6 billion parameters and has been trained on 200,000 hours of audio recordings. It supports the synthesis of monotone speech (as in audiobooks) and emotional speech (as in live conversation), as well as prefix-based synthesis (where an audio recording with the beginning of the speech is provided, based on which the model synthesizes the continuation according to the specified text, replicating the original characteristics of the speech, such as continuing to speak in whispers).

The output generates sound with a sampling rate of 44kHz. It supports the substitution of synthesizable inserts to simulate performances with multiple speakers or to build interactive dialogues, as well as the addition of tags to control speech speed, pitch, and emotional expressions such as joy, fear, sadness, and anger.

According to the developers, the quality of the generated speech does not lag behind or exceeds all publicly available open-source and commercial synthesis systems (tests provide comparisons with ElevenLabs, Cartesia, and FishSpeech). A drawback noted is a higher concentration of audio artifacts, such as coughing, breathing sounds, or creaking, at the beginning or end of the generated audio material.

  • Zonos:
  • ElevenLabs:
  • Cartesia:
  • Fish Speech v1.5:

To use the model on your system, a ready-to-work image for Docker has been prepared, which includes a web interface for managing synthesis based on the Gradio platform. To get started, simply install the image with the command 'git clone https://github.com/Zyphra/Zonos.git; cd Zonos; docker compose up' and open the page 'http://localhost:7860' in your browser. Having an NVIDIA GPU of at least the 3000 series with 6GB of video memory is recommended. The performance on a system with an RTX 4090 GPU exceeds the requirements needed for real-time synthesis by two times.

The AI voice synthesis model Zonos has been released, supporting voice cloning.


Source: opennet.ru
Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster