NVIDIA has announced the development of the CUDA Rust toolkit, which allows the use of the Rust language to develop kernels that run on the GPU side. The toolkit ensures memory safety at compile time and prevents race conditions. Next year, CUDA Rust is expected to reach a level suitable for developing production projects, comparable to the CUDA C++ and CUDA Python toolkits.
The CUDA Rust toolkit supports two models for developing concurrently executing kernels — SIMT (Single Instruction, Multiple Threads) and Tile. The SIMT model allows for low-level definition of a single thread's logic, launching thousands of such threads and managing them. The Tile model offers a higher level of abstraction, where instead of explicitly managing individual threads, algorithms are defined to operate on blocks of data (tiles), with all thread manipulations, access synchronization, memory management, and data distribution to tensor cores handled by the Tile IR compiler.
For developing in Rust using the SIMT model, the cuda-oxide compiler is being developed, which allows code written in Rust, using the standard type system and ownership model of Rust, to be compiled directly into instructions for execution in virtual machine CUDA PTX (Parallel Thread Execution). Kernels for the GPU are created in conventional Rust but run in a no_std environment and can only utilize functions from the libcore library and specialized Rust abstractions, without access to the standard Rust library (libstd).
In CUDA kernels written in Rust, protection through the type system (safe) is allowed, with the use of unsafe blocks and access to low-level hardware instructions. To ensure safety, a type called DisjointSlice is proposed, guaranteeing that each thread has exclusive access only to its own data. The cuda-oxide code is distributed under the Apache 2.0 license.
To apply the Tail model, the cutile-rs library is recommended, allowing the use of idiomatic Rust to create code that compiles directly into CUDA cores. Cutile-rs applies Rust's strict ownership and borrowing rules to code executed on the GPU. The core is designed as a single-threaded program that works with a data block, while the compiler divides the computations into threads and ensures their synchronization. Threads are given shared access to immutable tensors, and safe access to mutable tensors is achieved by partitioning them into non-overlapping blocks. The cutile-rs code is distributed under the Apache 2.0 license.
Source: opennet.ru
