Google has opened the source code of the AI system Magika for identifying file content types.

Google has announced the open sourcing of the Magika project, designed to determine content types based on analysis of the data present in a file. Magika can accurately identify the programming languages, compression methods, installation packages, executable code, markup types, and formats of audio, video, documents, and images used in the content. The associated toolkit and pre-trained machine learning model have been published under the Apache 2.0 license.

Unlike similar projects that determine MIME types based on content, Magika stands out by employing machine learning methods, achieving high performance, and exceptional accuracy in identification. The model has been trained using the Keras framework on 25 million file examples and supports recognition of 116 data types with at least 99% accuracy. The model is packaged in the ONNX format and is only 1 MB in size. The use of deep learning methods has allowed for a 50% increase in accuracy compared to the previously used Google system based on manually defined rules.

Google has opened the source code of the AI system Magika for identifying file content types.

In Google, the system is used for file classification in services such as Gmail, Drive, Code Insight, and Safe Browsing for security checks and compliance with service rules. Work is underway to integrate Magika into the VirusTotal platform as a step for initial file filtering before specific analyzers are executed. Deployed in Google's infrastructure, Magika's configuration enables scanning of several million files per second and hundreds of billions of files per week. After model loading, the output generation time is 5-6 ms when tested on a single CPU core. The identification time is almost unaffected by the file size.

To utilize Magika in their projects, a command-line utility, a Python package, and a JavaScript library capable of running in browsers or Node.js-based projects have been prepared. The command-line interface and API support batch operations, allowing multiple files to be checked in a single request. There is a recursive scanning mode for the entire directory content and three prediction modes to adjust error tolerance (high confidence, medium confidence, and best guess).

Google has opened the source code of the AI system Magika for identifying file content types.


Source: opennet.ru
Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster