Designworks Talent LLC
Designworks Talent LLC
Designworks Talent LLC is a recruitment and talent advisory firm serving high-growth startups and global enterprises. Founded in 2009, the company combines experienced recruiters with AI-enabled tools to support faster, more targeted hiring. Its work spans full-service searches, project-based recruiting, and flexible on-demand models, alongside scalable hiring strategy, recruiting operations, enterprise talent acquisition leadership, data-informed sourcing, and candidate experience. The team’s approach is designed to give organizations adaptable support as their hiring needs evolve.

Staff GPU Performance Kernel Engineer (Hybrid)

Lead GPU kernel optimization for an AI infrastructure cloud platform, improving utilization, latency, throughput, and scalability. Work across training and inference workloads to strengthen performance across the GPU fleet.

Description

  • Analyze and optimize GPU kernels to increase utilization, throughput, and latency performance.
  • Diagnose and remove data-plane bottlenecks in large-scale AI workloads.
  • Tune performance-critical workloads across training and inference environments.
  • Partner with AI infrastructure, machine learning, and platform engineering teams to optimize system behavior based on workload characteristics.
  • Create benchmarking approaches and performance measurement practices for GPU infrastructure.
  • Assess emerging GPU hardware, profiling tools, and optimization methods as platforms evolve.
  • Advance engineering practices that improve GPU efficiency, scalability, and fleet reliability.
  • Improve the performance layer supporting next-generation AI infrastructure.

Requirements

  • Substantial experience developing and optimizing GPU kernels with CUDA, ROCm, or comparable GPU programming frameworks.
  • A track record of improving GPU utilization, lowering latency, or raising throughput for production AI workloads.
  • Deep knowledge of GPU architecture, memory hierarchies, parallel computing, and the path from application code to hardware execution.
  • Experience profiling and debugging performance problems in complex AI or distributed computing systems.
  • Ability to independently lead technically complex initiatives and deliver solutions in a fast-paced engineering environment.
  • A strong systems programming foundation and performance engineering approach.
  • Preferred: experience optimizing workloads for both NVIDIA and AMD GPU architectures.
  • Preferred: experience with GPU compilers, runtime optimization, or low-level systems performance.
  • Preferred: contributions to open-source GPU performance projects, compiler tools, or AI systems optimization.
  • Preferred: experience with large-scale AI training, inference platforms, HPC, or cloud GPU infrastructure.
  • Preferred: familiarity with Nsight Systems, Nsight Compute, ROCm profiling tools, or comparable technologies.
  • U.S. work authorization is required.
  • Visa sponsorship is not currently available; employees are expected to work in the office at least three days per week once the permanent office is established.
  • benefits?

Benefits

  • Eligible roles may offer merit increases, annual bonuses, and long-term incentives tied to individual performance.
  • Medical, dental, and vision coverage for U.S.-based employees.
  • 401(k) plan with company matching.
  • Paid holidays provided according to the company calendar.
  • Hybrid work arrangement.

Related Jobs