The Group of Companies CRT, part of the Sberbank ecosystem, announced the development of an advanced speech synthesis platform that is said to ensure the smoothness and expressiveness of any text reading.
The presented solution is the third generation of speech synthesis system. High-quality audio signals are generated by complex neural network models. The developers claim that the result of these algorithms is the most realistic synthesis of Russian speech.

The platform includes a module for predicting stress in words that are not yet in the basic dictionary. Additionally, automatic correction of common spelling errors is provided. Thanks to deep linguistic analysis of text, pronunciation will match language norms even in complex cases.
Another advantage of the platform is that it does not require expensive servershardware equipped with GPU accelerators. The technology can be used in two ways - through a cloud service or embedded into your own solution.

Among the possible applications of the development are chatbots and voice assistants, information and notification services, voice services with instant synthesis of any text during a call, and more.
"In automated communication scenarios with clients, the technology allows for individual interaction with each subscriber, as there are no fixed messages, and any text can be synthesized during the call," say the developers.
The technology can be tested .
Source: 3dnews.ru
