Skip to content

How to Write Your First CUDA Kernel in Rust

Learn how to write and run your first NVIDIA GPU kernel in Rust with cuda-oxide, from installing the toolchain to compiling PTX and checking memory safety.

How to Write Your First CUDA Kernel in Rust

On this page

NVIDIA has opened a new path for Rust developers who want to write GPU kernels without switching to CUDA C++. Its CUDA Rust effort now includes cuda-oxide, an experimental Rust-to-PTX compiler for traditional SIMT kernels. This tutorial walks through the practical setup on Linux, creates a small vector-add kernel, explains the ownership rules that make the example interesting, and shows how to verify the generated GPU code.

What CUDA Rust actually lets you build

A GPU kernel is a function that runs across many GPU threads at once. In the traditional CUDA workflow, developers normally write those kernels in CUDA C++, while applications written in Rust call into the resulting GPU code through another layer. cuda-oxide changes that boundary by providing a custom Rust compiler backend that takes Rust kernel functions through Rust's intermediate representation and eventually produces NVIDIA PTX, the intermediate instruction format used to load GPU programs.

NVIDIA announced the project on September 8, 2026, describing two different approaches for Rust GPU programming: cuda-oxide for SIMT-style kernels and cuTile Rust for tile-based kernels. The SIMT model is the closer match to conventional CUDA programming because you explicitly reason about individual threads, blocks, and indexes. NVIDIA describes cuda-oxide as an early-alpha project, so it is better treated as an experimental programming environment than as a drop-in replacement for an established CUDA toolchain. 

Check the hardware and software requirements first

The first obstacle is not Rust. It is the GPU toolchain. The current cuda-oxide documentation targets Linux, has been tested on Ubuntu 24.04, requires an NVIDIA GPU from the Ampere generation or newer, and currently expects CUDA Toolkit 13.0 or newer, LLVM 21 or newer, and Clang 21. The project also pins a Rust nightly toolchain because its compiler backend depends on Rust compiler internals. 

That means this tutorial is not a good fit for a Windows-only machine or a computer without an NVIDIA GPU. The project documentation does provide a development-container route, which can reduce the amount of software you install directly on the host. The host still needs a compatible NVIDIA driver and GPU access, however, because the kernel eventually has to execute on NVIDIA hardware. 

Install the pinned Rust toolchain and CUDA dependencies

On a supported Linux machine, install the CUDA Toolkit first and make sure the CUDA headers and compiler are visible to your shell. You will also need LLVM with its NVIDIA PTX backend and Clang with the development headers required by the host-side bindings. The project provides a diagnostic command specifically because a missing compiler component can otherwise produce an error several layers away from the real problem.

Once Rust is installed, use the pinned nightly toolchain rather than choosing an arbitrary nightly release. The current project documentation pins nightly-2026-08-28 and lists rust-src, rustc-dev, rust-analyzer, clippy, rustfmt, and llvm-tools among the configured components. 

rustup toolchain install nightly-2026-08-28

rustup component add 
rust-src 
rustc-dev 
rust-analyzer 
clippy 
rustfmt 
llvm-tools 
--toolchain nightly-2026-08-28

Install the CUDA, LLVM, and Clang packages appropriate for your Linux distribution before continuing. If you are unsure whether everything is visible, do not start writing kernel code yet. Fix the environment first.

Install cargo-oxide and run its diagnostic check

cargo-oxide is the Cargo subcommand that drives the CUDA Rust compilation pipeline. Instead of manually invoking several compiler stages, you use commands such as cargo oxide run, cargo oxide inspect, and cargo oxide sanitize. The current documentation recommends installing it with the project's pinned nightly toolchain. ([NV Labs][1])

cargo +nightly-2026-08-28 install 
--git [https://github.com/NVlabs/cuda-oxide.git](https://github.com/NVlabs/cuda-oxide.git) 
cargo-oxide

Now run the diagnostic command:

cargo oxide doctor

This is worth doing before creating a project because the command checks the Rust toolchain, CUDA installation, LLVM, CUDA libraries, compiler backend, NVIDIA driver, and GPU. If it reports that cuda.h or stddef.h is missing, the problem is your development environment rather than the Rust kernel you have not written yet. ([GitHub][2])

Create a Rust GPU project

Once cargo oxide doctor passes, create a project with the CUDA-specific Cargo command:

cargo oxide new my_first_kernel
cd my_first_kernel

The generated project contains the host and device pieces needed by the compiler. The useful part of this design is that you are not maintaining a separate CUDA source tree for the kernel and a Rust application for everything else. The project is designed around single-source Rust code, with the GPU module marked so the CUDA Rust compiler knows which functions should become device code. ([GitHub][3])

Write a vector-add kernel that runs on the GPU

Vector addition is deliberately simple: every output element is the sum of the elements at the same position in two input arrays. It is a good first kernel because there is almost no algorithmic complexity, leaving you with the parts that actually matter when learning GPU programming: indexing, memory ownership, and launching thousands of parallel operations.

Inside the generated project, the device module can define a kernel using the #[cuda_module] and #[kernel] attributes. The important detail is the output buffer. A normal Rust mutable slice represents one exclusive mutable borrow, but a GPU kernel needs many threads to write different parts of the same output at the same time. DisjointSlice expresses that those writes are separated into non-overlapping pieces.

use cuda_device::{cuda_module, kernel, thread, DisjointSlice};

#[cuda_module]
mod kernels {
use super::*;

```
#[kernel]
pub fn vec_add(
    a: &[f32],
    b: &[f32],
    mut output: DisjointSlice<f32>,
) {
    let index = thread::index_1d();

    if let Some(value) = output.get_mut(index) {
        let i = index.get();
        *value = a[i] + b[i];
    }
}
```

}

The kernel gets one logical index for each GPU thread. The two input slices can be read by every thread, while DisjointSlice restricts each thread to the output element associated with its index. That is more than a naming convention: the type is part of the project's attempt to carry Rust's ownership reasoning into GPU execution. NVIDIA's published research on the related cuTile Rust system describes the same broader goal of extending ownership guarantees to GPU kernels. ([NVIDIA][4])

Understand why the output is not a normal mutable slice

This is the part that makes CUDA Rust different from simply putting Rust syntax around CUDA calls. Imagine 1,024 GPU threads all receiving the same &mut [f32]. Ordinary Rust would reject that because multiple mutable references cannot safely coexist. On a GPU, however, the operation is perfectly reasonable when thread zero writes element zero, thread one writes element one, and so on.

DisjointSlice provides a type for that pattern. The kernel obtains an index representing the current thread and uses it to request its own writable element. If the index is outside the buffer, get_mut returns no element rather than allowing an unchecked memory access. The result is a small example of how Rust's type system can describe parallel memory access instead of simply asking the programmer to promise that the access is safe.

Build and run the kernel

Before running the program, make sure an NVIDIA GPU is visible to the system. Then use the project's standard run command:

cargo oxide run vecadd

The exact generated example name depends on the project template you use, so if you created your own project and kernel name, run the corresponding example reported by cargo oxide list. The official first-kernel workflow uses cargo oxide run to compile the Rust kernel, produce the GPU representation, launch it, and verify the vector-add result. ([GitHub][5])

A successful run should report that all 1,024 vector elements are correct in the standard example. The number itself is not a performance benchmark; it simply gives you enough parallel work to prove that multiple GPU threads executed the kernel and produced the expected output. ([GitHub][3])

Inspect the generated GPU code instead of guessing

One advantage of the toolchain is that you can inspect what the compiler produced. Run the inspection command against your example:

cargo oxide inspect vecadd

This lets you examine generated PTX without having to reconstruct the entire compiler pipeline manually. PTX is useful here because it gives you a view between the Rust source and the final GPU machine instructions. If you are learning GPU optimization, that intermediate view becomes valuable later when you start changing indexing, memory access, or launch configuration. ([NV Labs][1])

Use Compute Sanitizer before trusting a kernel

A kernel that produces the right answer once is not necessarily correct. GPU memory errors can depend on timing and may not appear in a small test. NVIDIA's CUDA tooling includes Compute Sanitizer, and cuda-oxide exposes a sanitizer command for running examples through it.

cargo oxide sanitize vecadd --tool memcheck

Memory checking is especially useful after you change indexes or introduce shared memory. Shared memory is a fast memory region used by threads in the same GPU block, but incorrect synchronization or indexing can turn a working kernel into a race or invalid access. The project documentation specifically lists sanitizer support as part of the development workflow. ([NV Labs][1])

Know what is experimental before building on it

The first successful kernel can make CUDA Rust look production-ready, but the project is not there yet. NVIDIA describes cuda-oxide as experimental and in alpha, with bugs, incomplete features, and API changes expected. The repository also lists Linux as the supported environment, so a developer looking for a cross-platform Rust GPU solution should not mistake this project for a general-purpose replacement for every existing GPU framework. ([GitHub][3])

There is also an important distinction between the two Rust approaches NVIDIA announced. cuda-oxide follows the SIMT programming model, where you explicitly work with threads and indexing. cuTile Rust moves up a level and lets the compiler reason about tiles of data instead. NVIDIA's research reports that cuTile Rust reached 96% of cuBLAS performance for its evaluated GEMM workload on a B200, but that result belongs to the research evaluation and should not be interpreted as a general performance guarantee for every Rust kernel. ([NVIDIA][4])

Where to go after the first kernel

Vector addition teaches the mechanics, but the interesting work starts when the kernel has enough computation to benefit from the GPU. The next useful exercises are matrix operations, reductions, shared-memory algorithms, and kernels that process larger tensors. At that point, measure the result rather than assuming that moving code from the CPU to the GPU automatically makes it faster.

For debugging, keep cargo oxide inspect and Compute Sanitizer in the workflow. For compiler problems, run cargo oxide doctor before changing application code. And because the project is evolving quickly, check its current documentation before copying an older command or dependency version into a new project. CUDA Rust is still early, but the first kernel already demonstrates the interesting part: Rust's ownership model can become part of how developers reason about parallel GPU memory rather than something they abandon at the GPU boundary.

M

Written by

M. Rizwan Mirza

I’m M. Rizwan Mirza, a Full Stack Developer with over 12 years of experience in web development and software solutions. I work with modern web technologies and enjoy building practical, reliable, and user-friendly digital solutions. I’m also part of TechWare House, where I work on web development projects and technology solutions. One of my favorite websites is TheQuranic.com. Through WizTechnoz, I share my knowledge, experience, tutorials, and useful insights about technology.

51 posts published

All posts by this author

0 Comments

No comments yet. Be the first to share your thoughts.

Join the conversation

Log in or create a free account to leave a comment. You can edit or delete your own comments any time.