Mozilla has updated its Common Voice voice data sets, featuring pronunciation samples from over 200,000 individuals. The data is published as public domain (CC0). The proposed sets can be used in machine learning systems to build speech recognition and synthesis models.
Compared to the previous update, the amount of speech material in the collection has increased from 23.8 to 25.8 thousand hours of speech. More than 88,000 individuals participated in preparing materials in English, dictating 3,161 hours of speech (there were 84,000 participants and 3,098 hours previously). The Belarusian language set includes 7,903 participants and 1,419 hours of speech material (up from 6,965 participants and 1,217 hours), the Russian language set includes 2,815 participants and 229 hours (previously 2,731 participants and 215 hours), the Uzbek language set has 2,092 participants and 262 hours (up from 2,025 participants and 258 hours), and the Ukrainian language set includes 780 participants and 87 hours (up from 759 participants and 87 hours).
The Common Voice project aims to organize collaborative efforts to accumulate a database of voice templates that reflects the full diversity of voices and speech styles. Users are invited to voice the phrases displayed on the screen or assess the quality of data contributed by other users. The accumulated database, containing recordings of various pronunciations of standard phrases in human speech, can be used without restrictions in machine learning systems and research projects.
Source: opennet.ru
