The code for Jina Embedding, a model for vector representation of text meaning, has been opened.

Jina has launched a machine learning model for vector representation of text under the Apache 2.0 license — jina-embeddings-v2. The model converts arbitrary text, containing up to 8192 characters, into a small sequence of real numbers that form a vector associated with the original text, effectively reproducing its semantics. Jina Embedding is the first open machine learning model that matches the characteristics of proprietary text vectorization models from the OpenAI project (text-embedding-ada-002), also capable of processing texts containing up to 8192 tokens.

The distance between two generated vectors can be used to determine the semantic relationship between the original texts. In practice, the generated vectors can be used for text similarity analysis, organizing searches for related materials (ranking results by semantic closeness), grouping texts by meaning, generating recommendations (suggesting a list of similar text strings), detecting anomalies, identifying plagiarism, and classifying tests. Examples of applications include using the model for analyzing legal documents, business analytics, medical research for processing scientific articles, literary criticism, analyzing financial reports, and improving how chatbots handle complex queries.

Two variants of the jina-embeddings model are available for download (basic — 0.27 GB and lightweight — 0.07 GB), trained on 400 million pairs of text sequences in English, covering various domains of knowledge. During training, sequences of 512 tokens were extrapolated to a size of 8192 using the ALiBi (Attention with Linear Biases) method.

The Basic model includes 137 million parameters and is designed for use on stationary systems with GPUs. The reduced model includes 33 million parameters, offers lower accuracy, and is aimed at application on mobile devices and systems with limited memory. A larger model covering 435 million parameters is also planned for release soon. In addition, a multilingual variant of the model is in development, currently focused on supporting German and Spanish. A plugin for using the jina-embeddings model through the LLM toolkit has been prepared separately.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster