Mozilla has updated the Common Voice voice data sets, which include pronunciation samples from over 200,000 people. The data is published as public domain (CC0). The proposed sets can be used in machine learning systems to build speech recognition and synthesis models. Compared to the previous update, the amount of speech material in the collection has increased from 31.1 to 31.8 thousand hours of speech, of which 20.8 thousand hours have undergone verification. The number of supported languages has increased from 124 to 129 (the languages added are Xhosa, Kalenjin, Kivanda, Dholuo, and Tswana).
A total of 93,300 people participated in preparing materials in English, dictating 3,554 hours of speech (up from 92,300 participants and 3,508 hours). The Belarusian language set includes 8,400 participants and 1,815 hours of speech material (up from 8,291 participants and 1,766 hours), the Russian language set has 3,241 participants and 277 hours (up from 3,206 participants and 274 hours), the Uzbek language set consists of 2,189 participants and 265 hours (up from 2,170 participants and 264 hours), and the Ukrainian language set has 1,091 participants and 113 hours (up from 1,075 participants and 112 hours).
The Common Voice project aims to organize collaborative efforts to accumulate a database of voice templates that reflects the full diversity of voices and speech styles. Users are invited to voice the phrases displayed on the screen or assess the quality of data contributed by other users. The accumulated database, containing recordings of various pronunciations of standard phrases in human speech, can be used without restrictions in machine learning systems and research projects.
Source: opennet.ru
