The code for the Whisper speech recognition and translation system has been released

The OpenAI project, focused on developing open-access projects in the field of artificial intelligence, has published findings related to the Whisper speech recognition system. It claims that the system provides reliability and accuracy levels for English speech recognition that are comparable to human recognition. The code for the reference implementation based on the PyTorch framework and a set of pre-trained models ready for use has been made open source. The code is released under the MIT license.

The model was trained using 680,000 hours of speech data collected from various datasets, covering different languages and thematic areas. Approximately one-third of the speech data used for training comes from languages other than English. The proposed system accurately handles situations such as accented pronunciations, background noise, and technical jargon. In addition to transcribing speech to text, the system can also translate speech from any language into English and identify the presence of speech in an audio stream.

The models are available in two forms: an English language model and a multilingual model that supports languages including Russian, Ukrainian, and Belarusian. Each form is divided into five variants, differing in size and the number of parameters covered by the model. The larger the model size, the higher the accuracy and quality of recognition, but also greater requirements for GPU video memory and lower performance. For example, the smallest variant includes 39 million parameters and requires 1 GB of video memory, whereas the largest variant includes 1,550 million parameters and requires 10 GB of video memory. The minimum variant is 32 times faster than the maximum.

The code for the Whisper speech recognition and translation system has been released

The system utilizes a "Transformer" neural network architecture, which includes an encoder and decoder that interact with each other. Sound is divided into 30-second segments, transformed into a log-Mel spectrogram, and sent to the encoder. The output of the encoder is directed to the decoder, which predicts a text representation mixed with special tokens, allowing a single model to tackle tasks such as language identification, accounting for the chronology of phrase pronunciation, speech transcription in different languages, and translation to English.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster