Release of Speech Synthesizer RHVoice 1.8.0

The open speech synthesis system RHVoice 1.8.0 has been released, initially developed to provide quality support for the Russian language but later adapted for other languages, including English, Portuguese, Ukrainian, Kyrgyz, Tatar, and Georgian. The code is written in C++ and is distributed under the LGPL 2.1 license. It supports operation on GNU/Linux, Windows, and Android. The program is compatible with standard TTS interfaces for converting text to speech: SAPI5 (Windows), Speech Dispatcher (GNU/Linux), and Android Text-To-Speech API, and can also be used in the NVDA screen reader. The creator and main developer of RHVoice is Olga Yakovleva, who continues to develop the project despite being completely blind.

Version 1.8 for the Android platform introduces a new system for managing voice and language data, allowing updates to voice data to be downloaded without updating the mobile application. Checking for the availability of updates for added voices and languages occurs automatically. Additionally, the new release features support for the Polish language and a new voice for the Macedonian language. Compatibility with the latest alpha and beta releases of the NVDA screen reader has been ensured. Issues with building on the Linux platform, which arose in the absence of Speech Dispatcher, have been resolved.

It should be noted that RHVoice utilizes developments from the HTS project (HMM/DNN-based Speech Synthesis System) and employs a parametric synthesis method with statistical models (Statistical Parametric Synthesis based on HMM — Hidden Markov Model). The advantage of the statistical model is its low overhead and low CPU power demands. All operations are performed locally on the user's system. Three levels of speech quality are supported (the lower the quality, the higher the performance and the lower the response time).

A drawback of the statistical model is its relatively low quality of pronunciation, which does not reach the level of synthesizers that generate speech based on combinations of fragments of natural speech. Nonetheless, the result is quite intelligible and resembles the playback of a recording over a loudspeaker. For comparison, the Silero project, which provides an open engine for speech synthesis based on machine learning technologies and a set of models for the Russian language, surpasses RHVoice in quality.

There are 14 voice options available for the Russian language and 6 for English. The voices are formed based on recordings of natural speech. In the settings, you can adjust the speed, pitch, and volume. The Sonic library can be used to change the tempo. Automatic language detection and switching based on input text analysis is possible (for example, native synthesis models may be used for words and quotes in another language). Voice profiles are supported, defining combinations of voices for different languages.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster