Rust + CUDA Inference Benchmark
42%Benchmarking custom CUDA kernels against the current baseline.
- Current task
- Profiling kernel memory access and occupancy
- Next task
- Implement shared-memory optimization
- baseline Latency Ms
- 8.7
- current Latency Ms
- 7.1
gpucudarustbenchmark
