NVIDIA
NVIDIA
NVIDIA develops accelerated computing and artificial intelligence technologies used across gaming, data centers, cloud computing, healthcare, manufacturing, automotive, and robotics. Its work spans graphics processing units, AI platforms, simulation, and industry-focused solutions, including NVIDIA Omniverse for collaborative 3D workflows, NVIDIA DRIVE for autonomous vehicle development, and NVIDIA Clara for healthcare applications. The company brings together large technical teams working on the infrastructure and software that support advanced computing, simulation, and AI-driven analytics.

Senior Site Reliability Engineer, Production Engineering – India Remote

Support NVIDIA Cloud products as a senior site reliability engineer focused on Kubernetes administration, automation, observability, and incident response. Help maintain highly available production services across global operations.

Description

  • Lead a global Service Reliability Operations center
  • Support NVIDIA Cloud products and services
  • Collaborate with Site Reliability Engineering, Security Operations, DevOps, and other teams
  • Support production Kubernetes services by increasing automation and reducing manual work
  • Administer large-scale Kubernetes and systems environments while monitoring security to protect SLAs, integrity, and reliability
  • Use alerts, alarms, and observability tools to identify, prevent, and respond to incidents
  • Investigate logs, metrics, and system behavior to diagnose technical issues
  • Direct root cause analysis and deliver effective resolutions
  • Initiate and lead incident management calls
  • Coordinate subject matter experts and service owners through timely escalation and resolution
  • Create monitors, alarms, and alerts that strengthen service reliability and customer experience

Requirements

  • At least 7 years of demonstrated experience administering large-scale production Kubernetes systems in highly available Internet, cloud, or data center environments
  • Strong preference for on-premises infrastructure experience
  • Bachelor’s degree in Computer Science, Engineering, Mathematics, or equivalent experience
  • Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management
  • Familiarity with GPU or DPU hardware and high-performance computing cluster environments
  • Strong Linux administration experience, including DNS, DHCP, IP tables, routing, and firewalls
  • Ability to troubleshoot and maintain services on large-scale bare-metal infrastructure
  • Experience with CI/CD tools such as Jenkins and ArgoCD
  • Experience writing scripts
  • Python, Golang, or Rust programming experience is preferred but not required
  • Strong communication and interpersonal skills, with the ability to present persuasively to cross-functional groups

Benefits

  • Access to 24/7 production engineering team support
  • Flexibility to work split-weekend shifts

Related Jobs