Standard Intelligence has announced the release of hertz-dev, the first open AI model for full-duplex speech synthesis, which can be used as a foundation for creating real-time voice communication systems or generating conversational speech. The model allows for the generation of speech closely resembling the voice data it was trained on, ensuring interaction in a style akin to live human communication without delays reminiscent of a choppy phone call. The project's developments are distributed under the Apache 2.0 license.
On a system with an NVIDIA GeForce RTX 4090 GPU, the average latency before generation is 120 ms (theoretically up to 65 ms), which is roughly twice as fast as existing publicly available models. The released version is built using a transformer architecture, encompasses 8.5 billion parameters, and is trained using 500 billion tokens. The context size considered by the model (the number of tokens it can process and remember during speech generation) is 2048 tokens or approximately 4 minutes of speech.
Source: opennet.ru
