NVIDIA
NVIDIA
NVIDIA develops accelerated computing and artificial intelligence technologies used across gaming, data centers, cloud computing, healthcare, manufacturing, automotive, and robotics. Its work spans graphics processing units, AI platforms, simulation, and industry-focused solutions, including NVIDIA Omniverse for collaborative 3D workflows, NVIDIA DRIVE for autonomous vehicle development, and NVIDIA Clara for healthcare applications. The company brings together large technical teams working on the infrastructure and software that support advanced computing, simulation, and AI-driven analytics.

Senior Storage Software Engineer – DGX Cloud

Lead storage software for NVIDIA’s AI infrastructure, with a focus on open-source distributed file systems and object storage. Diagnose large-scale GPU cluster issues and establish standards for performance, durability, and tuning.

Description

  • Develop open-source parallel and distributed file systems and distributed object storage.
  • Contribute fixes and features upstream while working with project communities and maintainers.
  • Write and review production code as a hands-on storage software lead.
  • Inspect kernel, NFS, NVMe-oF, and SPDK code to diagnose software defects.
  • Make final technical decisions on storage deliveries using measurable targets.
  • Triage, troubleshoot, and identify root causes for complex storage issues across very large GPU clusters.
  • Analyze I/O and metadata performance, data corruption, and recovery behavior.
  • Assess storage architecture, capabilities, performance, and durability.
  • Conduct scale testing, benchmarking, and recovery exercises.
  • Qualify new builds against defined performance and durability targets.
  • Establish and recommend configuration, tuning, and operational practices for high-performance file systems on GPU infrastructure.
  • Help operators and internal customers implement storage guidance.
  • Partner with training, inference, accelerated computing, SRE, operations, networking, and security teams.
  • Work with cloud providers, neocloud operators, and storage vendors on shared architecture.
  • Apply modern AI coding and agentic tools to accelerate development, debugging, validation, and operations.

Requirements

  • Bachelor’s, master’s, or doctoral degree in Computer Science, Electrical Engineering, or a related discipline, or equivalent experience.
  • More than 12 years of direct storage software engineering experience.
  • Extensive experience with a high-performance parallel or distributed file system operating at multi-petabyte scale.
  • A track record of contributing to open-source distributed or parallel file system projects.
  • Hands-on experience developing and reviewing production code, inspecting file system, kernel, NVMe-oF, or SPDK source, and running scale tests or recovery exercises.
  • Experience diagnosing and resolving storage issues in large GPU or HPC clusters, including I/O and metadata performance analysis.
  • Strong command of at least one systems language: C, C++, Rust, or Go.
  • Proficiency in Python.
  • Working knowledge of Linux kernel storage and networking stacks, including the block layer, RDMA, RoCE, InfiniBand, NVMe, page cache, VFS, and multipath.
  • Strong understanding of object storage technologies such as S3 and Swift.
  • Strong understanding of block storage technologies such as NVMe-oF and iSCSI.
  • Excellent written and verbal communication skills.
  • Ability to operate in a 24/7 production environment.
  • A security-first approach to engineering and operations.
  • Experience maintaining or making sustained contributions to widely used public projects.
  • Experience designing or operating storage for AI training or inference at very large GPU scale.
  • Experience with kernel or file system development, metadata scalability, data placement, failure recovery, or HSM or equivalent systems.
  • Experience with Kubernetes and CSI driver development for storage.
  • Hands-on experience optimizing performance with SPDK, libfabric, or FUSE.

Benefits

  • Equity compensation
  • Employee benefits

Related Jobs

Pure Storage

Senior Systems Engineer – Queensland Remote

Pure Storage

Senior Systems Engineer supporting enterprise storage solution sales for Everpure customers in Queensland, Australia. The role advises enterprise accounts, leads technical evaluations, and collaborates with sales and alliance teams.

Open
Ventra Health

Senior Accounting Manager - Chennai Hybrid

Ventra Health
1,001 – 5,000 Employees
B2BHealthcareSaaS

Lead accounting operations for Ventra Health’s physician revenue-cycle business, covering close, reconciliations, reporting, billing, and compliance. Manage and develop Finance staff in a hybrid Chennai role.

Open
Tapcheck

Senior Performance Marketing Manager - Tapcheck Hybrid

Tapcheck
51 – 200 Employees
FinanceFintechHR Tech

Lead Tapcheck’s paid acquisition strategy across Google Ads, Microsoft Ads, and LinkedIn Ads for its on-demand pay platform. Own pipeline growth, conversion optimization, attribution, and revenue efficiency.

Open
IGS Energy

Residential Solar Services Partner | IGS Energy | Ohio Remote

IGS Energy
1,001 – 5,000 Employees
eCommerceEnergy

IGS Energy is hiring a remote Residential Solar Services Partner in Ohio to coordinate maintenance services, contractors, customers, and invoicing. The role manages service cases, documentation, payments, and customer experience from initiation through completion.

Open