
The architecture of the new language model is similar to Llama or Qwen, but it was developed entirely from scratch. This similarity allows the use of the same tools. The pretrain version of the large language model YandexGPT 5 Lite has 8 billion parameters with a context length of 32k tokens. Special attention was paid to the Russian language during the model's training, with materials in Russian making up over 70% of the dataset.
The senior model YandexGPT 5 is available in Alice and on the Yandex website, but it will not be made available to the public.
In its category, the model achieves parity with global SOTA on several key benchmarks for pretrain models, and in many others, it surpasses them. For example, according to the results of internal blind pairwise comparisons (side-by-side) for a wide range of queries, YandexGPT 5 Pro surpasses YandexGPT 4 Pro in 67% of cases and does not fall short of GPT-4o.
Source: linux.org.ru
