platform, deep learning inference, TensorRT and related compiler/runtime technologies, CUDA/GPU performance, model compression..., TensorRT/compiler flows, CUDA kernels, memory behavior, scheduling, runtime services, and automotive platform constraints...
architecture, memory hierarchy, and kernel optimization (HIP/CUDA). Hands-on delivery on large-scale C++/HIP/CUDA codebases..., such as ROCm (rocBLAS, hipDNN, Composable Kernel, AITemplate), the CUDA ecosystem (cuBLAS, cuDNN, CUTLASS, Thrust, CUB, NCCL...
Advanced Micro Devices ⚡⚡ Sat, 27 Jun 2026 02:07:48 GMT
and maintain high-performance GPU kernels using Triton or CUDA to accelerate custom model layers and critical workloads. Improve... building high-performance or distributed systems. Familiarity with CUDA or Triton for performance-critical workloads is highly...
Eng II - SW
with Python General Purpose Graphical Processing Unit (GPU) OPENCL or CUDA development experience is highly desired Experience...
BAE Systems ⚡ ⚡ Sun, 28 Jun 2026 02:15:19 GMT