NVIDIA
NVIDIA
NVIDIA dezvoltă tehnologii de calcul accelerat și inteligență artificială utilizate în domenii precum jocurile video, centrele de date, cloud computing, asistența medicală, producția, industria auto și robotica. Activitatea companiei acoperă unități de procesare grafică, platforme de inteligență artificială, simulare și soluții dedicate anumitor industrii, inclusiv NVIDIA Omniverse pentru fluxuri de lucru colaborative 3D, NVIDIA DRIVE pentru dezvoltarea vehiculelor autonome și NVIDIA Clara pentru aplicații medicale. Compania reunește echipe tehnice numeroase care lucrează la infrastructura și software-ul ce susțin calculul avansat, simularea și analiza bazată pe inteligență artificială.

Senior Distributed Systems Software Engineer – DGX Cloud | NVIDIA

Build and scale Kubernetes-based GPU infrastructure for NVIDIA DGX Cloud AI workloads. Develop scheduling, monitoring, and health-management systems for reliable, high-performance clusters.

Descriere

  • Help build the production systems that operate large-scale GPU clusters for AI workloads on the DGX Cloud team.
  • Create custom Kubernetes software for scheduling GPU resources.
  • Develop monitoring and health-management capabilities that improve GPU asset reliability, availability, and scalability.
  • Use diagnostics from GPU hardware alongside cluster and network telemetry to inform system operations.
  • Partner with teams across NVIDIA to keep production AI clusters reliable, consistent, and highly performant.
  • Analyze system failures and strengthen services through a structured incident-management process.

Cerințe

  • Substantial Kubernetes software-engineering experience spanning cluster operations, operator development, node-health monitoring, and GPU resource scheduling.
  • Professional software-engineering experience in a highly technical organization with demonstrable results.
  • Hands-on development with Kubernetes APIs and frameworks, beyond administering clusters.
  • Clear communication skills and the ability to collaborate with multifunctional teams, principals, and architects across organizations and regions.
  • At least 8 years of experience in a comparable role, including work on large-scale production systems.
  • Working knowledge of established software-engineering principles, tools, and practices.
  • Bachelor’s degree in computer science, engineering, physics, mathematics, or a related field, or equivalent experience.
  • Proficiency in systems programming languages such as Go or Python.
  • Strong command of data structures and algorithms.
  • Technical ability to manage and automate large-scale distributed systems across cloud providers.
  • Advanced practical experience with cluster-management platforms including Kubernetes, Slurm, and Bright Cluster Manager.
  • A demonstrated record of operational excellence in maintaining reliable, performant AI infrastructure.

Beneficii

  • Equity compensation.
  • Company benefits.

Locuri de muncă similare

Kreato Global | BPO and Language Solutions

Remote English-Spanish OPI/VRI Interpreter

Kreato Global | BPO and Language Solutions

Interpret remotely between English- and Spanish-speaking people in medical, financial, social service, and customer care settings. Provide language support for Kreato Global across Latin America.

Deschide
RecruitGo

Remote Executive Assistant, Outreach and Marketing — Philippines

RecruitGo

Handle administrative support, CRM, outreach, and marketing for RecruitGo, an Employer of Record serving global clients with talent from emerging markets. Support two UK-facing clients with day-to-day administration and creative campaigns.

Deschide