Nvidia is closing a long-standing gap for Rust developers: the GPU kernel itself. Two new projects, cuda-oxide and cutile-rs, let developers write kernels directly in Rust that compile natively to PTX, rather than wrapping code written in another language. cuda-oxide provides a custom rustc codegen backend that compiles SIMT-style kernels through the Pliron IR framework and LLVM, and it requires a pinned nightly toolchain. cutile-rs takes the Tile-based route, running on stable Rust 1.89 and above with CUDA 13.3 and no custom LLVM, with the compiler handling thread mapping and memory layout through CUDA Tile IR JIT compilation.

The Rust push mirrors the rest of Nvidia's systems layer. The Nova Linux driver is written in Rust, Nvidia Dynamo is built on a Rust core, and NVTX already has Rust bindings. Both new projects enforce memory safety at compile time: cuda-oxide uses DisjointSlice and launch contracts to prevent aliasing, while cutile-rs relies on tensor partitioning and ownership to guarantee exclusive access. Nvidia says it will keep growing and maturing CUDA Rust into 2027 and beyond, alongside the already mature CUDA C++ and CUDA Python toolchains.

Adoption is already happening on the Tile track. cutile-rs is published on crates.io and is used in HuggingFace's Grout inference engine and mistral.rs, while cuda-oxide remains in early alpha. Nvidia plans to support inter-language interoperability between CUDA Rust, CUDA C++ and CUDA Python, so choosing a frontend does not lock developers out of the others. For founders building inference and serving stacks, that means kernel-level performance work no longer requires leaving the Rust ecosystem behind.