Vultr
Vultr
Vultr — dasturchilar va bizneslarga talab asosida hisoblash, saqlash, tarmoq va ma’lumotlar bazasi xizmatlarini taqdim etuvchi global bulut infratuzilmasi kompaniyasi. Uning platformasi virtual mashinalar, bare-metal serverlar, GPU tezlashtirilgan infratuzilma, boshqariladigan ma’lumotlar bazalari, obyekt va blokli saqlash, Kubernetes hamda joylashtirishga tayyor marketplace ilovalarini o‘z ichiga oladi. AMD va NVIDIA GPU qo‘llab-quvvatlovi, tezkor tarmoq hamda 32 dan ortiq ma’lumotlar markazi hududlaridagi infratuzilma bilan Vultr dasturchilarga yo‘naltirilgan API’lar va kengaytiriladigan bulut vositalari orqali dasturiy ta’minot, sun’iy intellekt va yuqori unumli hisoblash ish yuklamalariga xizmat ko‘rsatadi.

Senior GPU Infrastructure Engineer - United States Remote

Lead the validation, performance optimization, and reliability of Vultr’s GPU infrastructure for AI training and inference. You’ll qualify new hardware, troubleshoot distributed systems, and improve cluster automation at scale.

Tavsif

  • Design and scale GPU infrastructure supporting AI training and inference workloads
  • Manage end-to-end validation for GPU clusters and incoming hardware platforms
  • Lead qualification and initial deployment activities for new GPU platforms
  • Identify and address performance constraints across GPU, CPU, PCIe, and networking layers
  • Create and maintain validation frameworks, test suites, and performance reference points
  • Build and improve automation for cluster provisioning and validation workflows
  • Set performance targets and define repeatable validation methods
  • Diagnose complex distributed-system problems, including issues involving libraries such as NCCL
  • Strengthen system reliability through preventive testing and performance tuning
  • Coach engineers and expand the team’s technical capabilities
  • Coordinate cross-functional efforts to increase GPU cluster reliability and efficiency

Talablar

  • At least five years of experience in GPU infrastructure, high-performance computing, or distributed systems
  • Advanced knowledge of Linux systems and server hardware
  • Demonstrated experience operating or supporting large-scale GPU clusters
  • Strong Python programming ability beyond basic scripting
  • Hands-on experience with Ansible or comparable automation frameworks
  • Experience establishing validation standards and performance baselines for GPU infrastructure
  • Strong debugging ability spanning hardware, operating system, and network layers
  • Working knowledge of high-speed networking concepts
  • Clear communication skills and a collaborative approach to cross-team work
  • Based in the United States
  • Legally authorized to work in the United States
  • Must indicate whether employment visa sponsorship is required

Imtiyozlar

  • The company covers 100% of employee medical, dental, and vision insurance premiums
  • 401(k) matching at 100% of contributions up to 4%, with immediate vesting
  • Up to $2,500 annually for professional development reimbursement
  • Eleven paid holidays, accrued paid time off, and PTO rollover
  • Additional PTO after three-year and ten-year anniversaries
  • One month of paid sabbatical every five years
  • Annual anniversary bonus
  • $500 remote-office setup stipend in the first year and $400 in each subsequent year
  • Internet reimbursement of up to $75 per month
  • Gym membership reimbursement of up to $50 per month
  • Company-paid Wellable subscription

O‘xshash ish o‘rinlari