Update of Mozilla Common Voice voice data 7.0

NVIDIA and Mozilla have announced an update to the voice data sets collected as part of the Common Voice initiative, which now includes pronunciation examples from 182,000 people, a 25% increase compared to six months ago. The data is published as public domain (CC0). The proposed sets can be used in machine learning systems to build speech recognition and synthesis models.

Compared to the last update, the volume of speech material in the collection has increased from 9 to 13.9 thousand hours of speech. The number of supported languages has grown from 60 to 76, with support for Belarusian, Kazakh, Uzbek, Bulgarian, Armenian, Azerbaijani, and Bashkir languages added for the first time. The Russian language set covers 2,136 participants and 173 hours of speech material (up from 1,412 participants and 111 hours), while the Ukrainian language set includes 615 participants and 66 hours (up from 459 participants and 30 hours).

Over 75,000 people participated in preparing materials in English, recording 2,637 hours of verified speech (up from 66,000 participants and 1,686 hours). Interestingly, the language with the second-largest amount of collected data is Kinyarwanda, with 2,260 hours gathered. This is followed by German (1,040), Catalan (920), and Esperanto (840). The languages showing the most dynamic growth in voice data include Thai (growing twentyfold, from 12 to 250 hours), Luganda (from 8 to 80 hours), Esperanto (from 100 to 840 hours), and Tamil (from 24 to 220 hours).

As part of its participation in the Common Voice project, NVIDIA has prepared pre-trained models for machine learning systems based on the collected data (supported by PyTorch). The models are distributed as part of the free and open NVIDIA NeMo toolkit, which is already being used, for example, in the automated voice services of MTS and Sberbank. These models are designed for use in speech recognition, speech synthesis, and natural language processing systems, and may be useful for researchers working on voice dialogue systems, transcription platforms, and automated call centers. Unlike earlier available projects, the published models are not limited to recognizing English and cover various languages, accents, and forms of speech.

It is worth noting that the Common Voice project aims to facilitate collaborative work for accumulating a database of voice samples that considers the full diversity of voices and speaking styles. Users are encouraged to voice the phrases displayed on the screen or to evaluate the quality of data added by other users. The accumulated database with recordings of various pronunciations of standard phrases of human speech can be used without restrictions in machine learning systems and research projects.

According to the author of the Vosk speech recognition library, the shortcomings of the Common Voice dataset include the one-sidedness of the voice material (predominance of male voices aged 20-30 years, with a lack of material featuring voices of women, children, and the elderly), the absence of vocabulary variability (repetition of the same phrases), and the distribution of recordings in a distorted MP3 format.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster