OpenXLA has been introduced, a toolkit for optimizing and compiling machine learning models

The largest companies involved in machine learning development have unveiled the OpenXLA project, aimed at collaboratively advancing tools for compiling and optimizing models for machine learning systems. The project has taken over the development of tools that unify the compilation of models prepared in TensorFlow, PyTorch, and JAX frameworks for efficient training and execution on various GPUs and specialized accelerators. Companies such as Google, NVIDIA, AMD, Intel, Meta, Apple, Arm, Alibaba, and Amazon have joined the collaborative effort on this project.

It is expected that, through the combined efforts of leading research teams and community representatives, the development of machine learning systems will be stimulated and the issue of infrastructure fragmentation for various frameworks and hardware will be addressed. OpenXLA allows for effective support of diverse hardware, regardless of which framework a machine learning model is based on. It is anticipated that OpenXLA will reduce model training time, increase throughput, decrease latencies, lower computing resource costs, and shorten time-to-market.

OpenXLA has been introduced, a toolkit for optimizing and compiling machine learning models

OpenXLA consists of three main components, the code for which is distributed under the Apache 2.0 license:

  • XLA (Accelerated Linear Algebra) — a compiler that optimizes machine learning models for high-performance execution on different hardware platforms, including GPUs, CPUs, and specialized accelerators from various manufacturers.
  • StableHLO — a specification and basic implementation of a set of high-level operations (HLO, High-Level Operations) for use in machine learning system models. It acts as a layer between machine learning frameworks and compilers that transform the model for execution on specific hardware. Layers for generating models in StableHLO format have been prepared for PyTorch, TensorFlow, and JAX frameworks. The MHLO set, which underlies StableHLO, has been extended with support for serialization and versioning.
  • IREE (Intermediate Representation Execution Environment) is a compiler and runtime that transforms machine learning models into a universal intermediate representation based on the MLIR (Multi-Level Intermediate Representation) format from the LLVM project. Key features include ahead-of-time compilation, support for flow control, the ability to utilize dynamic elements in models, optimization for various CPUs and GPUs, and low overhead.

Key benefits of the OpenXLA toolkit:

  • Achieving optimal performance without the need for writing device-specific code. Providing ready-made optimizations, including simplification of algebraic expressions, efficient memory placement, and execution scheduling to reduce peak memory consumption and overhead.
  • Simplifying scaling and parallelizing computations. The developer only needs to add annotations for a subset of critical tensors, upon which the compiler can automatically generate code for parallel computations.
  • Ensuring portability by supporting various hardware platforms, such as AMD and NVIDIA GPUs, x86 and ARM-based CPUs, Google's TPU ML accelerators, AWS's Trainium and Inferentia, Graphcore, and Cerebras Wafer-Scale Engine.
  • Support for connecting extensions that implement additional capabilities, such as the ability to write deep machine learning primitives using CUDA, HIP, SYCL, Triton, and other languages for parallel computing. The possibility for manual tuning of bottlenecks in models.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster