Google Translatotron – a technology for simultaneous speech translation that mimics the user's voice

Developers from Google have introduced a new project that created a technology capable of translating spoken sentences from one language to another. The main distinction of the new translator called Translatotron from its counterparts is that it works exclusively with sound, without using any intermediate text. This approach has significantly sped up the performance of the translator. Another notable point is that the system accurately mimics the pitch and tone of the speaker.

Creating Translatotron was possible thanks to continuous work that lasted several years. Researchers from Google have long considered the possibility of direct speech conversion, but until recently, it had not been possible to realize this idea.

Google Translatotron – a technology for simultaneous speech translation that mimics the user's voice

The systems currently used for simultaneous translation typically operate based on one algorithm. At the initial stage, the original speech is transformed into text. Then, the text in one language is converted into text in another language. Finally, the resulting text is transformed back into speech in the target language. This method works reasonably well, but it is not without its drawbacks. Errors can occur at each stage, compounding one another and leading to a decrease in translation quality.

To achieve the necessary results, the researchers studied sound spectrograms. They aimed to ensure that the spectrogram in one language would convert into a spectrogram in another language, bypassing the text conversion stages.


Google Translatotron – a technology for simultaneous speech translation that mimics the user's voice

It is worth noting that despite the complexity of such a transformation, speech processing occurs in a single step, rather than in three steps as before. With sufficient computing power, Translatotron will perform simultaneous translation significantly faster. Another important aspect is that this approach preserves the characteristics and intonation of the original voice.

At this stage, Translatotron cannot boast the same high translation accuracy as standard systems. However, researchers say that most translations performed are of sufficient quality. Further work on Translatotron will continue, as researchers aim to improve the quality of simultaneous speech translation.



Source: 3dnews.ru
Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster