Google has released the open audio codec Lyra V2.

Google has introduced the Lyra V2 audio codec, which utilizes machine learning methods to achieve maximum speech transmission quality over very slow communication channels. The new version features a transition to a new neural network architecture, support for additional platforms, enhanced bitrate management capabilities, improved performance, and higher sound quality. The reference implementation of the code is written in C++ and is distributed under the Apache 2.0 license.

In terms of voice data transmission quality at low speeds, Lyra significantly outperforms traditional codecs that use digital signal processing methods. To achieve high-quality voice transmission under limited data transfer requirements, Lyra employs a speech model based on a machine learning system that reconstructs missing information based on typical speech characteristics, in addition to conventional sound compression and signal transformation methods.

The codec includes an encoder and a decoder. The encoder works by extracting voice data parameters every 20 milliseconds, compressing them, and transmitting them to the recipient over the network with a bitrate ranging from 3.2 kbps to 9.2 kbps. On the recipient's side, the decoder uses a generative model to reconstruct the original speech signal based on transmitted audio parameters, which include logarithmic mel-spectrograms reflecting the characteristics of speech energy across various frequency ranges and prepared in consideration of human auditory perception models.

Lyra V2 features a new generative model based on the SoundStream convolutional neural network, which has low computational resource requirements, allowing real-time decoding even on lower-powered systems. The sound generation model has been trained using several thousand hours of voice recordings in over 90 languages. TensorFlow Lite is used to execute the model. The performance of the proposed implementation is sufficient for encoding and decoding speech on budget smartphones.

In addition to using another generative model, the new version is also noteworthy for including a codec architecture with an RVQ (Residual Vector Quantizer) quantizer, which is executed on the sender side before data transmission and on the receiver side after data reception. The quantizer transforms the parameters output by the codec into sets of packets, encoding information relative to the chosen bitrate. To ensure varying levels of quality, quantizers are provided for three bitrates (3.2 kbps, 6 kbps, and 9.2 kbps); the higher the bitrate, the better the quality, but with higher bandwidth requirements.

Google has released the open audio codec Lyra V2.

The new architecture has reduced signal transmission delays from 100 to 20 milliseconds. In comparison, the Opus codec for WebRTC demonstrated tested delays at bitrates of 26.5 ms, 46.5 ms, and 66.5 ms. Additionally, the performance of the encoder and decoder has significantly improved — up to 5 times faster compared to the previous version. For example, on the Pixel 6 Pro smartphone, the new codec encodes and decodes a 20-millisecond sample in 0.57 ms, which is 35 times faster than needed for real-time transmission.

In addition to performance, there's also an improvement in sound recovery quality — on the MUSHRA scale, the speech quality at bitrates of 3.2 kbps, 6 kbps, and 9.2 kbps using the Lyra V2 codec corresponds to bitrates of 10 kbps, 13 kbps, and 14 kbps using the Opus codec.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster