Cisco
Cisco
Cisco is an enterprise technology company focused on networking, security, and digital infrastructure. Its portfolio includes routers, switches, optical transceivers, programmable silicon, edge computing platforms, Webex collaboration tools, observability products, AI-enabled software, and cloud-managed services. Cisco serves enterprises, service providers, and governments, supporting the design, operation, and security of large-scale networks and data centers through technology, professional services, training, and ongoing support.

Software Engineer, AI Compute Performance (Mid-Senior, Hybrid) at Cisco

Optimize GPU-accelerated AI and HPC workloads for Cisco Validated Infrastructure Services. Automate performance testing and diagnose bottlenecks across large-scale compute systems.

Description

  • Evaluate and optimize AI compute performance across hardware, system software, and applications
  • Execute and assess AI training, inference, and HPC workloads on GPU-based platforms
  • Diagnose compute, memory, PCIe, storage, and workload-scaling problems
  • Apply telemetry, logging, tracing, and profiling data to locate performance constraints
  • Build automation for performance testing, regression identification, and failure triage
  • Partner with hardware and software teams to address system and workload performance issues
  • Record technical findings, test methods, and measured performance outcomes
  • Help Cisco Validated Infrastructure Services verify large AI cluster performance before customer handoff

Requirements

  • Bachelor’s degree plus at least 5 years of related experience, a master’s degree plus at least 3 years, a PhD with relevant research experience, or equivalent related experience
  • Hands-on experience debugging, testing, or tuning Linux-based systems
  • Experience executing or assessing AI, machine learning, or HPC workloads
  • At least 3 years using a programming or scripting language such as Python, C, C++, Go, or Bash
  • Working familiarity with GPU computing technologies including CUDA, ROCm, NCCL, or RCCL
  • Familiarity with performance analysis tools such as Nsight, DCGM, rocprofiler, or comparable tools
  • Understanding of GPU system elements including memory, NUMA, PCIe, and scale-up interconnects
  • Exposure to NVIDIA NVLink and NVSwitch or AMD Infinity Fabric/xGMI
  • Experience with Kubernetes, Slurm, or another distributed workload environment
  • Strong ownership, technical documentation, and cross-functional collaboration abilities

Benefits

  • Medical, dental, and vision insurance
  • 401(k) plan with Cisco matching contributions
  • Paid parental leave
  • Short- and long-term disability insurance
  • Basic life insurance
  • Potential eligibility for Cisco restricted stock unit grants
  • Ten paid holidays per full calendar year
  • One floating holiday for non-exempt employees
  • A paid day off for the employee’s birthday
  • Paid year-end holiday shutdown
  • Four paid personal wellness days
  • Sixteen paid vacation days per full calendar year for non-exempt employees
  • Flexible vacation program without a defined limit for eligible exempt employees
  • Eighty hours of sick time provided upon hire and every January 1 thereafter
  • Annual carryover of up to 80 unused sick hours
  • Additional paid leave for critical or emergency family matters
  • Optional ten paid volunteer days per full calendar year
  • Annual bonuses for non-sales roles, subject to Cisco policies

Related Jobs

Knowtion Health

Remote Talent Acquisition Manager

Knowtion Health

Oversee recruiting systems, requisitions, analytics, and contingent workforce operations at Knowtion Health. Help support scalable hiring processes for a growing healthcare company.

Open