Mozilla is developing the Whisperfile speech recognition toolkit

Mozilla is developing the Whisperfile speech recognition toolset, which includes an independent, high-performance implementation of the Whisper machine learning model, developed and released by OpenAI. The toolset is built on top of whisper.cpp, an implementation of the Whisper model in C/C++, created by Georgi Gerganov (the author of llama.cpp). The code is written in C++ and is distributed under the MIT license.

Whisperfile is being developed by the Mozilla Ocho team and complements the llamafile project, which is designed to create universal executables for running large language machine learning models (LLM). Similar to llamafile, the whisperfile project allows you to generate an executable file based on a machine learning model parameter file in GGUF format, which can run on various operating systems on AMD64 and ARM64 processor hardware. The compiled code can link with the standard C library Cosmopolitan, enabling the creation of application builds that run on Linux, FreeBSD, macOS, OpenBSD, NetBSD, and Windows.

When running the executable, a speech audio file in wav, mp3, ogg, or flac format is passed as an input parameter, and the recognized text is saved as output. In practice, the project can be used for tasks such as generating text subtitles for videos, creating logs of voice and video calls, converting recorded voice materials into text, and organizing voice input. With Whisperfile, such tasks can be solved on a local system without relying on external services.

Additionally, it supports operating as an HTTP server that processes speech recognition requests through a Web API. To accelerate the operation of the model, GPUs and AVX instructions may be utilized. The toolset can also output confidence scores, allowing recognized words to be highlighted based on the accuracy of their identification.

Mozilla is developing the Whisperfile speech recognition toolkit

The Whisper model has been trained on 680,000 hours of speech data covering various topics and languages (2/3 of the data is in English). The model performs well in recognizing accented speech, understands technical jargon, supports automatic language detection, and can operate in the presence of background noise. For English speech, the system demonstrates a level of reliability and accuracy in automatic recognition that is close to human recognition. Besides transcribing speech to text, the model can also be used for translating speech into another language.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster