Google has released Magika 1.0, a toolkit for determining the content type of files

Google has introduced the release of the Magika 1.0 toolkit, designed to identify content types based on the analysis of data present in the file. Magika can accurately determine the programming languages, compression methods, installation packages, executable code, markup types, and formats of audio, video, documents, and images contained within. The associated toolkit and ready model of machine learning are distributed under the Apache 2.0 license. Wrappers have been prepared for the languages Rust, Python, JavaScript/TypeScript, and Go.

Unlike similar projects that determine MIME types based on content, Magika stands out by employing machine learning methods, high performance and accuracy in identification. The model was trained using the Keras framework on 100 million file examples (with a dataset size of over 3 TB) and supports recognition of 200 data types with an accuracy of at least 99%. The model is packaged in ONNX format and is only a few megabytes in size. Utilizing deep learning methods has allowed for a 50% increase in accuracy compared to the previously used Google system based on manually specified rules.

In Google, the system is used for classifying files in services like Gmail, Drive, Code Insight, and Safe Browsing during security checks and compliance with service regulations. Integration of Magika into VirusTotal and abuse.ch platforms has been ensured as a link for primary filtering of files before executing specific analyzers. The deployment of Magika in Google's infrastructure provides scanning of several million files per second and several hundred billion files per week. After loading the model, the output generation time is 5 ms when tested on a single CPU core. The identification time is almost independent of the file size.

To utilize Magika in your projects, a command-line utility, packages for Python, Rust, and Go, as well as a JavaScript library capable of operating in a browser or within Node.js-based projects have been prepared. The command-line interface and API support batch processing, allowing multiple files to be checked in a single request. There is a recursive scanning mode for the entire contents of a directory and three prediction modes to adjust for error resilience (high confidence, medium confidence, and best guess).

Initially, the project was developed in Python, but with the preparation of release 1.0, the content type detection engine was rewritten in Rust, achieving higher performance while maintaining an adequate level of code security. The ONNX Runtime framework is utilized for executing the machine learning model, and the Tokio library is used for parallel asynchronous request handling. On a MacBook Pro (M4), the engine's performance allows for processing around 1000 files per second.

In addition to the new engine, the changes in release 1.0 include an expansion in the number of supported types from approximately 100 to 200; the addition of a new command-line client written in Rust; improved accuracy in detecting text formats such as configuration files and code; and a redesign of modules for Python and TypeScript to facilitate their integration with other projects. Among the newly supported content types are formats used in machine learning and AI; programming languages such as Swift, Kotlin, TypeScript, Dart, Solidity, Web Assembly, and Zig; DevOps components (Dockerfiles, TOML, HashiCorp, Bazel build files, and YARA rules); SQLite databases; AutoCAD files (dwg, dxf), Adobe Photoshop (psd), and font formats (woff, woff2). There has been an improvement in the separation of code for C++ and C, JavaScript, and TypeScript.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster