Mozilla has released an update to its Common Voice voice data sets, including pronunciation samples from around 200,000 individuals. The data is published as public domain (CC0). The proposed sets can be utilized in machine learning systems to build speech recognition and synthesis models. Compared to the previous update, the volume of speech material in the collection has increased by 30% — from 13.9 to 18.2 thousand hours of speech. The number of supported languages has grown from 67 to 87.
The Russian language set includes 2,452 participants and 193 hours of speech material (previously 2,136 participants and 173 hours), the Belarusian language set comprises 6,160 participants and 987 hours (previously 3,831 participants and 356 hours), and the Ukrainian language set consists of 684 participants and 76 hours (previously 615 participants and 66 hours). More than 79,000 individuals contributed to the English language materials, recording 2,886 hours of validated speech (up from 75,000 participants and 2,637 hours).
It is worth noting that the Common Voice project aims to facilitate collaborative efforts to build a database of voice templates that reflects the diverse range of voices and speech styles. Users are encouraged to read aloud phrases displayed on the screen or evaluate the quality of data submitted by other users. The accumulated database of recordings featuring various pronunciations of standard human speech phrases can be used without restrictions in machine learning systems and research projects. According to the author of the Vosk speech recognition library, the shortcomings of the Common Voice set include a lack of diversity in voice material (predominance of male voices aged 20-30, and a shortage of recordings from women, children, and elderly individuals), a lack of vocabulary variability (repetition of the same phrases), and the prevalence of recordings in the distorted MP3 format.
Additionally, the release of NVIDIA NeMo 1.6 can be noted, which provides machine learning methods for creating speech recognition, speech synthesis, and natural language processing systems. NeMo includes ready-trained models for machine learning systems based on the PyTorch framework, prepared by NVIDIA using Common Voice speech data, covering various languages, accents, and speech varieties. These models may be useful for researchers working on voice dialogue systems, transcription platforms, and automated call centers. For example, NVIDIA NeMo is used in the automated voice services of MTS and Sberbank. The NeMo code is written in Python using PyTorch and is distributed under the Apache 2.0 license.
Source: opennet.ru
