JumpCloud
JumpCloud
JumpCloud develops a unified identity, device, and access management platform for organizations with distributed IT environments. Its SaaS products bring together cloud directory services, cross-platform device management, Active Directory modernization, identity lifecycle controls, conditional access, passwordless authentication, multifactor authentication, and zero trust security. The platform also supports automated onboarding and offboarding, compliance workflows, SaaS management, and integrations with HR systems, helping enterprise teams manage users and devices across operating systems while reducing operational complexity.

Site Reliability Engineer – India Remote

Improve the reliability of JumpCloud’s AI-powered IT management platform across AWS and GCP. Work across observability, Kubernetes, disaster recovery, automation, and incident response.

Description

  • Design, deploy, and maintain the reliability, availability, and performance of critical JumpCloud systems and APIs across AWS and GCP
  • Partner with application teams to operationalize SLIs, SLOs, and error budgets
  • Build and improve end-to-end observability for microservices and cloud infrastructure with tools such as Datadog
  • Create monitoring around the Golden Signals: latency, traffic, errors, and saturation
  • Join on-call rotations, respond to incidents, and lead blameless post-incident reviews
  • Operate production Kubernetes EKS clusters through GitOps workflows including Argo CD and Kargo
  • Provision and secure multi-cloud infrastructure with modular Terraform configurations
  • Maintain disaster recovery dashboards and runbooks, plus multi-region failover automation and validation tests aligned with RTO and RPO targets
  • Develop production-ready Python or Go scripts and automation that reduce operational toil
  • Apply AI-assisted tools such as Cursor, Claude Code, and GitHub Copilot to scripting, runbook creation, and incident triage

Requirements

  • At least five years of professional software engineering experience in SRE, DevOps, or Platform Engineering for 24/7 mission-critical systems
  • Proficiency in Python or Go for SRE tooling, custom automation, and cloud integrations
  • Production experience with Kubernetes, container orchestration, and GitOps pipelines such as Argo CD
  • Experience creating, maintaining, and modularizing Terraform configurations
  • Experience operating AWS workloads, including EKS, IAM, VPC networking, Route 53, and ALB/NLB, or comparable GCP workloads
  • Practical knowledge of FinOps, cost-allocation tagging, resource right-sizing, and cloud-spend dashboards
  • Experience creating disaster recovery dashboards, conducting failover drills, and monitoring system health and recovery metrics
  • Hands-on experience with Datadog or a comparable platform, PagerDuty, alerting practices, and SLI/SLO frameworks
  • Operational experience configuring and troubleshooting production service meshes such as Istio and highly available proxies such as HAProxy or NGINX
  • Strong troubleshooting ability and a record of improving operational efficiency through code
  • Collaborative team orientation and alignment with company core values
  • Availability and willingness to participate in on-call shifts
  • Fluent spoken and written English
  • Preferred: experience with GitHub Actions or GitLab Pipelines
  • Preferred: basic knowledge of chaos engineering or resilience testing
  • Preferred: familiarity with HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator
  • Preferred: basic knowledge of DevSecOps tools and infrastructure-as-code vulnerability remediation

Benefits

  • Remote-first work based in India
  • Work in a fast-moving SaaS environment
  • Opportunities for professional growth and knowledge sharing
  • Collaborate with talented teams across the globe
  • Contribute employee perspectives to product and feature development
  • Work with a supportive executive team and board
  • Equal opportunity employment
  • Applications from third-party resume providers are not accepted

Related Jobs

The Clorox Company

Shopper Insights & Analytics Manager - Amazon, Hybrid

The Clorox Company

Lead shopper and category analytics that support The Clorox Company’s Amazon growth. Turn behavior, research, and retail data into decisions on assortment, pricing, content, conversion, and repeat purchasing.

Open