NVIDIA
NVIDIA
NVIDIA o‘yinlar, ma’lumotlar markazlari, bulutli hisoblash, sog‘liqni saqlash, ishlab chiqarish, avtomobilsozlik va robototexnika sohalarida qo‘llaniladigan tezlashtirilgan hisoblash hamda sun’iy intellekt texnologiyalarini ishlab chiqadi. Kompaniyaning faoliyati grafik protsessorlar, sun’iy intellekt platformalari, simulyatsiya va muayyan sohalarga mo‘ljallangan yechimlarni qamrab oladi. Bular qatoriga hamkorlikdagi 3D ish jarayonlari uchun NVIDIA Omniverse, avtonom transport vositalarini ishlab chiqish uchun NVIDIA DRIVE va sog‘liqni saqlash ilovalari uchun NVIDIA Clara kiradi. Kompaniya ilg‘or hisoblash, simulyatsiya va sun’iy intellekt asosidagi tahlilni qo‘llab-quvvatlaydigan infratuzilma hamda dasturiy ta’minot ustida ishlaydigan yirik texnik jamoalarni birlashtiradi.

Senior Distributed Systems Software Engineer – DGX Cloud | NVIDIA

Build and scale Kubernetes-based GPU infrastructure for NVIDIA DGX Cloud AI workloads. Develop scheduling, monitoring, and health-management systems for reliable, high-performance clusters.

Tavsif

  • Help build the production systems that operate large-scale GPU clusters for AI workloads on the DGX Cloud team.
  • Create custom Kubernetes software for scheduling GPU resources.
  • Develop monitoring and health-management capabilities that improve GPU asset reliability, availability, and scalability.
  • Use diagnostics from GPU hardware alongside cluster and network telemetry to inform system operations.
  • Partner with teams across NVIDIA to keep production AI clusters reliable, consistent, and highly performant.
  • Analyze system failures and strengthen services through a structured incident-management process.

Talablar

  • Substantial Kubernetes software-engineering experience spanning cluster operations, operator development, node-health monitoring, and GPU resource scheduling.
  • Professional software-engineering experience in a highly technical organization with demonstrable results.
  • Hands-on development with Kubernetes APIs and frameworks, beyond administering clusters.
  • Clear communication skills and the ability to collaborate with multifunctional teams, principals, and architects across organizations and regions.
  • At least 8 years of experience in a comparable role, including work on large-scale production systems.
  • Working knowledge of established software-engineering principles, tools, and practices.
  • Bachelor’s degree in computer science, engineering, physics, mathematics, or a related field, or equivalent experience.
  • Proficiency in systems programming languages such as Go or Python.
  • Strong command of data structures and algorithms.
  • Technical ability to manage and automate large-scale distributed systems across cloud providers.
  • Advanced practical experience with cluster-management platforms including Kubernetes, Slurm, and Bright Cluster Manager.
  • A demonstrated record of operational excellence in maintaining reliable, performant AI infrastructure.

Imtiyozlar

  • Equity compensation.
  • Company benefits.

O‘xshash ish o‘rinlari

Peraton

Senior Hybrid Connectivity Engineer

Peraton

Designs and troubleshoots connectivity across AWS classified cloud, Kubernetes, and on-premises networks. Supports Peraton’s national security and government technology missions.

Ochish