NVIDIA
NVIDIA
NVIDIA develops accelerated computing and artificial intelligence technologies used across gaming, data centers, cloud computing, healthcare, manufacturing, automotive, and robotics. Its work spans graphics processing units, AI platforms, simulation, and industry-focused solutions, including NVIDIA Omniverse for collaborative 3D workflows, NVIDIA DRIVE for autonomous vehicle development, and NVIDIA Clara for healthcare applications. The company brings together large technical teams working on the infrastructure and software that support advanced computing, simulation, and AI-driven analytics.

Senior Distributed Systems Software Engineer – DGX Cloud | NVIDIA

Build and scale Kubernetes-based GPU infrastructure for NVIDIA DGX Cloud AI workloads. Develop scheduling, monitoring, and health-management systems for reliable, high-performance clusters.

Description

  • Help build the production systems that operate large-scale GPU clusters for AI workloads on the DGX Cloud team.
  • Create custom Kubernetes software for scheduling GPU resources.
  • Develop monitoring and health-management capabilities that improve GPU asset reliability, availability, and scalability.
  • Use diagnostics from GPU hardware alongside cluster and network telemetry to inform system operations.
  • Partner with teams across NVIDIA to keep production AI clusters reliable, consistent, and highly performant.
  • Analyze system failures and strengthen services through a structured incident-management process.

Requirements

  • Substantial Kubernetes software-engineering experience spanning cluster operations, operator development, node-health monitoring, and GPU resource scheduling.
  • Professional software-engineering experience in a highly technical organization with demonstrable results.
  • Hands-on development with Kubernetes APIs and frameworks, beyond administering clusters.
  • Clear communication skills and the ability to collaborate with multifunctional teams, principals, and architects across organizations and regions.
  • At least 8 years of experience in a comparable role, including work on large-scale production systems.
  • Working knowledge of established software-engineering principles, tools, and practices.
  • Bachelor’s degree in computer science, engineering, physics, mathematics, or a related field, or equivalent experience.
  • Proficiency in systems programming languages such as Go or Python.
  • Strong command of data structures and algorithms.
  • Technical ability to manage and automate large-scale distributed systems across cloud providers.
  • Advanced practical experience with cluster-management platforms including Kubernetes, Slurm, and Bright Cluster Manager.
  • A demonstrated record of operational excellence in maintaining reliable, performant AI infrastructure.

Benefits

  • Equity compensation.
  • Company benefits.

Related Jobs

Kreato Global | BPO and Language Solutions

Remote English-Spanish OPI/VRI Interpreter

Kreato Global | BPO and Language Solutions
201 – 500 Employees
HealthcareHospitalityLogistics

Interpret remotely between English- and Spanish-speaking people in medical, financial, social service, and customer care settings. Provide language support for Kreato Global across Latin America.

Open