Skip to main content
A

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Anyone Ai
1 day ago
Contract
Remote
Worldwide
Remote Engineering
Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads. We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking. WHAT YOU’LL WORK ON You’ll work with GPU and accelerator kernel tasks involving: - Kernel implementation and debugging - CUDA and Triton optimization - Translation between kernel frameworks - Hardware migration - Operator fusion - Performance profiling and benchmarking - Numerical correctness verification - Compilation and runtime debugging - Memory hierarchy optimization - Kernel-level AI workload performance You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware. WHAT WE’RE LOOKING FOR - 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels - Strong experience with at least two of the following: - CUDA - Triton - NKI... Click Apply to read the full job description.