New Release of the Silero Speech Synthesis System

A new public release of the neural network-based Silero Text-to-Speech synthesis system is now available. The project primarily aims to create a modern high-quality speech synthesis system that rivals commercial solutions from corporations and is accessible to all without the need for expensive server hardware.

The models are distributed under the GNU AGPL license, but the developing company does not disclose the training mechanism for the models. To run it, PyTorch and frameworks supporting the ONNX format can be used. The speech synthesis in Silero is based on the use of deep modified modern neural network algorithms and methods of digital signal processing.

It is noted that the main issue with modern neural network solutions for speech synthesis is that they are often only available as paid cloud solutions, while public products have high hardware requirements, lower quality, or are not finished and ready-to-use products. For example, to run one of the new popular end-to-end synthesis architectures, VITS, seamlessly in synthesis mode (i.e., not for training models), graphics cards with more than 16 gigabytes of VRAM are required.

Contrary to the prevailing trend, Silero solutions successfully run even on a single x86 Intel processor thread with AVX2 instructions. On 4 processor threads, it allows for synthesizing between 30 and 60 seconds in 8 kHz mode, 15-20 seconds in 24 kHz mode, and about 10 seconds in 48 kHz mode.

Key features of the new Silero release:

  • The model size has been reduced by half to 50 megabytes;
  • Models can make pauses;
  • Four high-quality voices are available in Russian (and an infinite number of random voices). Pronunciation examples;
  • Models have become 10 times faster and can synthesize up to 20 seconds of audio per second on 4 processor threads in 24 kHz mode;
  • All voice options for one language are packed into a single model;
  • Models can accept entire paragraphs of text as input, with support for SSML tags;
  • Synthesis works immediately at three selectable sampling rates — 8, 24, and 48 kilohertz;
  • The 'child problems' have been solved: instability and word skipping.
  • Flags have been added to control automatic accentuation and the placement of the letter 'ё'.

Currently, there are 4 voices in Russian publicly available for the latest synthesis version, but the next version with the following changes will be released soon:

  • The synthesis speed will increase by another 2-4 times;
  • Synthesis models for the CIS languages will be updated: Kalmyk, Tatar, Uzbek, and Ukrainian;
  • Models for European languages will be added;
  • Models for Indian languages will be added;
  • Models for the English language will be added.

Some of the systemic issues inherent in Silero synthesis include:

  • Unlike more traditional synthesis solutions such as RHVoice, Silero synthesis does not have integration with SAPI, easy-to-install clients, and integrations for Windows and Android;
  • The speed, while unprecedentedly high for such a solution, may be insufficient for real-time synthesis on weak processors at high quality;
  • The automatic accent placement solution does not handle homographs (words like замок and замОк) and still makes errors, but this shortcoming will be addressed in future releases;
  • The current synthesis version does not work on processors without AVX2 instructions (or requires special modification of PyTorch settings), as one of the modules inside the model is quantized;
  • The current synthesis version essentially has PyTorch as its only dependency; everything is 'embedded' within the model and JIT packages. The model source codes are not published, nor is the code for running models from PyTorch clients for other languages;
  • Libtorch, available for mobile platforms, is much bulkier than the ONNX runtime, but the ONNX version of the model is not yet provided.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster