Mozilla has updated its Common Voice voice data sets, which include pronunciation samples from over 200,000 individuals. The data is published as public domain (CC0). The provided sets can be used in machine learning systems to build speech recognition and synthesis models. Compared to the last update, the volume of speech material in the collection has increased from 28.7 to 30.3 thousand hours of speech, of which 19.7 thousand hours have been validated. The number of supported languages has risen from 114 to 120 (Yiddish, Latgalian, Ligurian, Ossetian, Telugu, and Western Sierra Puebla Nahuatl have been added).
A total of 90,670 individuals contributed to the preparation of English language materials, dictating 3,438 hours of speech (up from 88,900 participants and 3,347 hours). The Belarusian language set includes 8,249 participants and 1,641 hours of speech material (up from 8,205 participants and 1,632 hours), the Russian language set encompasses 3,133 participants and 265 hours (up from 3,053 participants and 260 hours), the Uzbek language set consists of 2,151 participants and 264 hours (up from 2,141 participants and 263 hours), and the Ukrainian language set comprises 1,058 participants and 108 hours (up from 1,024 participants and 105 hours).
The Common Voice project aims to organize collaborative efforts to accumulate a database of voice templates that reflects the full diversity of voices and speech styles. Users are invited to voice the phrases displayed on the screen or assess the quality of data contributed by other users. The accumulated database, containing recordings of various pronunciations of standard phrases in human speech, can be used without restrictions in machine learning systems and research projects.
Source: opennet.ru
