Vultr
Vultr
Vultr is a global cloud infrastructure company that provides developers and businesses with on-demand computing, storage, networking, and database services. Its platform includes virtual machines, bare-metal servers, GPU-accelerated infrastructure, managed databases, object and block storage, Kubernetes, and deployable marketplace applications. With support for AMD and NVIDIA GPUs, fast networking, and infrastructure across more than 32 data center regions, Vultr serves software, artificial intelligence, and high-performance computing workloads through developer-focused APIs and scalable cloud tools.

Senior GPU Infrastructure Engineer - United States Remote

Lead the validation, performance optimization, and reliability of Vultr’s GPU infrastructure for AI training and inference. You’ll qualify new hardware, troubleshoot distributed systems, and improve cluster automation at scale.

Description

  • Design and scale GPU infrastructure supporting AI training and inference workloads
  • Manage end-to-end validation for GPU clusters and incoming hardware platforms
  • Lead qualification and initial deployment activities for new GPU platforms
  • Identify and address performance constraints across GPU, CPU, PCIe, and networking layers
  • Create and maintain validation frameworks, test suites, and performance reference points
  • Build and improve automation for cluster provisioning and validation workflows
  • Set performance targets and define repeatable validation methods
  • Diagnose complex distributed-system problems, including issues involving libraries such as NCCL
  • Strengthen system reliability through preventive testing and performance tuning
  • Coach engineers and expand the team’s technical capabilities
  • Coordinate cross-functional efforts to increase GPU cluster reliability and efficiency

Requirements

  • At least five years of experience in GPU infrastructure, high-performance computing, or distributed systems
  • Advanced knowledge of Linux systems and server hardware
  • Demonstrated experience operating or supporting large-scale GPU clusters
  • Strong Python programming ability beyond basic scripting
  • Hands-on experience with Ansible or comparable automation frameworks
  • Experience establishing validation standards and performance baselines for GPU infrastructure
  • Strong debugging ability spanning hardware, operating system, and network layers
  • Working knowledge of high-speed networking concepts
  • Clear communication skills and a collaborative approach to cross-team work
  • Based in the United States
  • Legally authorized to work in the United States
  • Must indicate whether employment visa sponsorship is required

Benefits

  • The company covers 100% of employee medical, dental, and vision insurance premiums
  • 401(k) matching at 100% of contributions up to 4%, with immediate vesting
  • Up to $2,500 annually for professional development reimbursement
  • Eleven paid holidays, accrued paid time off, and PTO rollover
  • Additional PTO after three-year and ten-year anniversaries
  • One month of paid sabbatical every five years
  • Annual anniversary bonus
  • $500 remote-office setup stipend in the first year and $400 in each subsequent year
  • Internet reimbursement of up to $75 per month
  • Gym membership reimbursement of up to $50 per month
  • Company-paid Wellable subscription

Related Jobs

Basis Technologies

Office and Administrative Coordinator — Basis Technologies, Chicago (Hybrid)

Basis Technologies

Manage daily operations at Basis Technologies’ Chicago headquarters while supporting executives with scheduling, travel, and administrative projects. Coordinate facilities, vendors, events, and workplace logistics in a hybrid role.

Open