🛠️ project Enki – Write GPU compute kernels in pure stable Rust
Hey guys!
I've been building a GPU compute platform on pure stable-rust to let you write compute kernels directly in standard rust, Built entirely on Vulkan 1.3 (Compute) for the host api runtime & MLIR/LLVM for the backend JIT compiler.
It can make you write iters-enums-generics-etc.. in your function, and run it across GPU silicon, because the syntax is literally rust, It can execute on native CPU.
Here is a simple example:
use enki::*;
use glam::Vec2;
const COUNT: usize = 5;
// Declare the compute kernel with #[nam]
#[nam]
fn scale_vectors(_space: &Space, input: &Vec2, output: &mut Vec2, factor: f32) {
*output = *input * factor;
}
fn main() {
// Initialize the headless GPU runtime
let enki = Enki::init();
// Allocate physical data directly in GPU VRAM
let in_gpu = gpu_vec![
Vec2::new(1.0, 2.0),
Vec2::new(3.0, 4.0),
Vec2::new(5.0, 6.0),
Vec2::new(7.0, 8.0),
Vec2::new(9.0, 10.0),
];
let mut out_gpu = gpu_vec![Vec2::ZERO; COUNT];
let factor = 2.5f32;
// Record and dispatch directly to GPU silicon
enki.flow(|_| {
scale_vectors.run(
&Space::gpu_x(COUNT),
&in_gpu,
&mut out_gpu,
GpuParam::new(factor),
);
});
// Dual Execution: Run the exact same function on CPU native rust
let in_cpu = vec![
Vec2::new(1.0, 2.0),
Vec2::new(3.0, 4.0),
Vec2::new(5.0, 6.0),
Vec2::new(7.0, 8.0),
Vec2::new(9.0, 10.0),
];
let mut out_cpu = vec![Vec2::ZERO; COUNT];
for i in 0..COUNT {
scale_vectors(&Space::cpu_x(i, COUNT), &in_cpu[i], &mut out_cpu[i], factor);
}
// Verify bit-for-bit equivalence
assert_eq!(&out_cpu[..], &out_gpu.to_vec()[..]);
println!("\nGPU Results: {:?}", out_gpu);
println!("\nCPU Results: {:?}", out_cpu);
println!("\nCPU and GPU outputs match.");
}
Result:
enki_matmul on master [✘!+] is 📦 v0.1.0 via 🦀 v1.98.1
❯ cargo run
Compiling enki_matmul v0.1.0 (/home/mohiman/Projects/enki_demo_pack/enki_matmul)
Finished `dev` profile [optimized + debuginfo] target(s) in 1.86s
Running `target/debug/enki_matmul`
GPU Results: [Vec2(2.5, 5.0), Vec2(7.5, 10.0), Vec2(12.5, 15.0), Vec2(17.5, 20.0), Vec2(22.5, 25.0)]
CPU Results: [Vec2(2.5, 5.0), Vec2(7.5, 10.0), Vec2(12.5, 15.0), Vec2(17.5, 20.0), Vec2(22.5, 25.0)]
CPU and GPU outputs match.
enki_matmul on master [✘!+] is 📦 v0.1.0 via 🦀 v1.98.1 took 2s
❯
I've verified it on Linux (Docker Container Testing) & Windows (Wine) & colab T4, and my intel UHD laptop, every thing was great, i'd really love your help testing across more diverse hardware (AMD, Nvidia, Intel)
You can test SDF showcase in 3 commands:
git clone https://github.com/enkiruntime/enki_sdf.git
cd enki_sdf
cargo run --release
(pressing space will toggle execution from GPU (Enki) to CPU (rust-rayon))
- GitHub Repo: https://github.com/enkiruntime/enki
Enki is currently in Public Alpha (v0.2.x) release, I would really appreciate any feedback, architectural thoughts, or bug reports.
4
u/R4ND0M1Z3R_reddit 6h ago
How is it different compared to cubecl?
5
u/Xe8zz 6h ago
that's a good question, CubeCl (from the burn team) is great but enki works quite different under the hood
CubeCl is a proc macro transpiler, which means it parses the AST logic and transpile it into a shading language (like wgsl or cuda) enki doesn't do AST transpiling, it lets rustc compile your standard code into real bitcode (LLVM IR or .bc) and then the Parsu jit compiler lowers that LLVM IR directly into SPIR-V,
That means you can use deeper rust language features (enums with payload, iters, generic and external crates like "glam") naturally, also CubeCL uses descriptor sets and bindings tables, enki built on Vulkan 1.3 Buffer Device Addresses (BDA) memory pointer.
2
u/STSchif 5h ago
This sounds incredibly cool. Unfortunately I don't really have any use case for it currently, but there just certainly be things it can speed up transparently.
I wonder if it could be used for an embedded search engine like tantivy. If the index fits into vram it should theoretically be a lot faster than cpu, especially with compute intensive ranking functions. I'd wager you could run ranking functions of thousands of documents simultaneously.
1
u/R4ND0M1Z3R_reddit 5h ago
How is it able to operate on LLVM IR? Does it require 2 step compilation to IR then using llc or do you abuse the fact that process macros are loaded by compiler somehow?
1
u/Xe8zz 4h ago
haha, no the proc macro doesn't use the AST or abuse the macros, in the first run using 'cargo run' enki will ask you:
warning: missing required compilation profile and LLVM bitcode for GPU JIT synthesis
= note: GPU synthesis requires optimization (`opt-level = 2`) to lower host abstractions into silicon primitives
= note: execution requires `target.'cfg(all())'.rustflags = ["--emit=llvm-bc"]`
= help: configuration can be appended to `.cargo/config.toml` safely without modifying existing blocks
--> append configuration to `.cargo/config.toml`? [Y/n]
if 'Y' enki will append `target.'cfg(all())'.rustflags = ["--emit=llvm-bc"]` flag to let rust produce the .bc bitcode, this happens only on the first time run, from there parsu will take the 'nam' LLVM IR code from the 'target' deps folder and compiles it to spirv, that's why i used the 'nam' macro is simply just a marker/anchor so the runtime knows exactly which function and call-graph to locate and extract from the .bc file
if you don't want enki to configur the `.cargo/config.toml` file you can just run it with 'enki run' using cargo-enki CLI tool.
1
u/________-__-_______ 4h ago
That means you need to compile your code twice for any changes to take effect, right?
2
u/Xe8zz 4h ago
Nope, it's just standard single compilation pass, because of the '--emit=llvm-bc' flag, 'cargo run' compiles your binary code and updates the '.bc' bitcode simultaneously in the same pass, whenever you change your kernel and run 'cargo run' the updated bitcode is emitted immediatly and enki JIT compiler (Parsu) will JIT the new logic right away.
1
u/teerre 5h ago
Going to be deleted, but I don't know how much the claim actually follows. Yeah, the code is the same (supposing it works, I haven't tried), but it's the same because it's written in this shader-ish style. Nobody would write
scale_vectors(&Space::cpu_x(i, COUNT), &in_cpu[i], &mut out_cpu[i], factor); in a pure idiomatic Rust program
3
u/________-__-_______ 4h ago
To me the interesting is that you can run arbitrary Rust code inside the shader. Invoking it initially may be a bit of extra work but the computation can be as idiomatic as you want.
2
u/Xe8zz 4h ago
yeah that's fair, but keep in mind this is designed for parallel spatial workloads, (mapping grid coordinates), not sequential CPU code, knowing your spatial index is necessary for actually running any thing, also this is the only way to make the cpu execute the same nam function and get the same result (because the indexing is the same) but sequential for depug/testing, the point is the code inside the nam function is pure rust, you can call any math library from Crates.io, call functions with impl&struct, use en
ums/iters/etc..
note that if you tried to call a function that uses an internal heap alloc this what will you get this type of error:error[E0002]: dynamic heap allocation is not supported inside #[nam]
--> src/main.rs:11:16
|
11 | let data = json!({
| ^^^^^^^ heap allocation occurs here
|
= note: GPU nams execute in parallel across silicon cells without a
dynamic memory heap manager.
= help: use fixed-size stack arrays `[T; N]` or pre-allocate a `GpuVec`
on the CPU before dispatch.
For more information about this nam diagnostic, try `enki --explain E0002`.
error: could not compile `enki_matmul` (nam `scale_vectors`) due to 1 previous error
from this function:
use enki::*;
use glam::Vec2;
use serde_json::json;
const COUNT: usize = 5;
#[nam]
fn scale_vectors(_space: &Space, input: &Vec2, output: &mut Vec2, factor: f32) {
*output = *input * factor;
let data = json!({
"runtime": "enki",
"threads": [1024, 2048],
"nested": {
"status": "testing"
}
});
let formatted = data.to_string();
}
6
u/Living-Mall-7108 7h ago
the #[nam] proc macro is a wild naming choice but i respect the commitment to the bit