Mozilla has updated the Common Voice voice data sets, which include pronunciation samples from over 200,000 individuals. The data is published as public domain (CC0). The offered sets can be used in machine learning systems to build speech recognition and synthesis models. Compared to the last update, the volume of speech material in the collection has increased from 31.8 to 32.6 thousand hours of speech, of which over 20 thousand hours have passed validation. The number of supported languages has increased from 129 to 131.
In preparing the materials in English, 93.9 thousand people participated, dictating 3,587 hours of speech (there were 93.3 thousand participants and 3,554 hours previously). The Belarusian language set covers 8,444 participants and 1,846 hours of speech material (previously there were 8,400 participants and 1,815 hours), the Russian language set consists of 3,296 participants and 278 hours (previously 3,241 participants and 277 hours), the Uzbek language set includes 2,200 participants and 265 hours (previously 2,189 participants and 265 hours), and the Ukrainian language set has 1,104 participants and 114 hours (previously there were 1,091 participants and 113 hours).
The Common Voice project aims to organize collaborative efforts to accumulate a database of voice templates that reflects the full diversity of voices and speech styles. Users are invited to voice the phrases displayed on the screen or assess the quality of data contributed by other users. The accumulated database, containing recordings of various pronunciations of standard phrases in human speech, can be used without restrictions in machine learning systems and research projects.
Source: opennet.ru
