Mirantis
Mirantis
Mirantis develops cloud-native infrastructure and container management software for organizations building, securing, and operating Kubernetes environments at scale. Its portfolio includes Mirantis Kubernetes Engine, Mirantis OpenStack for Kubernetes, Mirantis Container Cloud, Mirantis Container Runtime, and Mirantis Secure Registry, supporting cluster operations and software supply-chain security. The company also contributes to the open-source ecosystem through projects and tools such as Lens Desktop, a Kubernetes IDE, alongside enterprise technical support. Mirantis works with organizations in consulting, healthcare, logistics, public services, financial services, and technology-focused industries adopting modern cloud infrastructure.

Senior Site Reliability Engineer (SRE), Remote

Join Mirantis as a senior SRE working on NVIDIA-backed Kubernetes and cloud AI infrastructure. Help improve reliability, security, automation, and scalability for enterprise environments.

Description

  • Design, develop, and operate cloud AI solutions within the CNCF ecosystem, including Kubernetes
  • Deploy AI infrastructure on NVIDIA-certified hardware in line with approved architecture and implementation plans
  • Maintain the reliability, security, and performance of container platforms
  • Coach team members and Mirantis customers
  • Partner with distributed international teams on technical issues and process improvements
  • Build, maintain, implement, and troubleshoot cloud and AI infrastructure based on open-source software
  • Work with stakeholders to define and refine technical requirements
  • Improve platform performance, reliability, and scalability
  • Investigate, debug, and resolve complex technical problems
  • Contribute to peer code reviews
  • Keep up with evolving cloud operations, development practices, and industry standards
  • Create and deploy AI-powered automation throughout the DevOps lifecycle
  • Transfer technical knowledge to customers during delivery engagements
  • Help stakeholders establish technical strategies and address complex challenges
  • Coordinate the reliable integration of cloud and software services

Requirements

  • At least five years of professional DevOps experience focused on cloud and infrastructure technologies such as Kubernetes and/or OpenStack
  • Background in high-performance data center computing, networking, and storage
  • Exposure to Go plus working knowledge of Python and JavaScript
  • Strong understanding of distributed systems, microservices, and CI/CD pipelines
  • Advanced troubleshooting and debugging ability across networking, storage, Linux, and Kubernetes, with a focus on performance and security
  • Proven ability to lead technical work and collaborate with varied teams
  • Confidence making independent decisions while supporting customers with limited daily supervision
  • Fluent written and spoken English
  • Strong communication skills in customer-facing settings
  • Commitment to innovation, ongoing learning, and high-quality delivery
  • Willingness to travel internationally up to 25% when required
  • Bachelor’s degree in Computer Science or a related discipline, or equivalent experience
  • Five or more years of experience in DevOps, software development, or a comparable position
  • Preferred: substantial experience designing network and/or storage architectures
  • Preferred: experience with HPC or GPU platforms, including GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand, NVLink, DCGM health checks, GPU driver and firmware lifecycles, or NVIDIA AI Enterprise
  • Preferred: participation in open-source communities through upstream contributions or conference presentations
  • Preferred: experience with Rancher, OpenShift, and VMware

Benefits

  • Professional development and training opportunities
  • Participation in conferences and working groups
  • Company outings, happy hours, hackathons, and technical talks
  • Competitive compensation package with a comprehensive benefits plan
  • Remote work option

Related Jobs