JumpCloud
JumpCloud
JumpCloud develops a unified identity, device, and access management platform for organizations with distributed IT environments. Its SaaS products bring together cloud directory services, cross-platform device management, Active Directory modernization, identity lifecycle controls, conditional access, passwordless authentication, multifactor authentication, and zero trust security. The platform also supports automated onboarding and offboarding, compliance workflows, SaaS management, and integrations with HR systems, helping enterprise teams manage users and devices across operating systems while reducing operational complexity.

Senior Site Reliability Engineer, India Remote

Senior SRE responsible for multi-region AWS and GCP infrastructure, Kubernetes, disaster recovery, observability, FinOps, automation, and incident management at JumpCloud.

Description

  • Own the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication systems across AWS and GCP
  • Build and continuously improve disaster recovery processes, automated multi-region failover, and business continuity plans
  • Define and enforce service-level indicators, objectives, and error budget practices
  • Lead the observability strategy with Datadog and Golden Signals monitoring
  • Manage on-call escalations and major incidents while supporting 99.99% availability service-level agreements
  • Run blameless post-incident reviews and deliver systemic root-cause fixes
  • Design, operate, and scale production EKS clusters with Argo CD and Kargo GitOps workflows
  • Develop and maintain Terraform infrastructure across multi-account, multi-region cloud environments
  • Create FinOps and cost-optimization dashboards covering multi-cloud spend, unit economics, and resource utilization
  • Develop production-quality Python or Go tools, platform automation, and custom integrations
  • Promote AI-assisted development workflows with Cursor, Claude Code, and GitHub Copilot
  • Maintain operational runbooks and architecture decision records
  • Mentor junior and mid-level engineers

Requirements

  • At least 8 years of professional experience in SRE, DevOps, or platform engineering for highly available, mission-critical distributed systems operating around the clock
  • Bachelor’s degree in computer science, software engineering, or a comparable technical field
  • Advanced proficiency in Python or Go
  • Production experience managing EKS or GKE cluster lifecycles, ingress and egress, networking, RBAC, and GitOps tools such as Argo CD
  • Strong Terraform expertise in complex, multi-account AWS environments, including IAM, VPCs, Transit Gateway, ALB/NLB, and Route 53
  • Experience improving cloud cost efficiency through right-sizing, cost allocation, tagging, workload optimization, and FinOps reporting
  • Experience designing and validating multi-region disaster recovery architectures and automated failover
  • Experience defining SLIs and SLOs, managing PagerDuty rotations, and improving production observability platforms
  • Experience operating enterprise service meshes such as Istio or Linkerd, along with production ingress and proxy systems such as HAProxy or NGINX
  • Ability to lead technical discussions, produce architecture documents and RFCs, and mentor engineering colleagues
  • Strong analytical, communication, and collaboration skills
  • Preferred familiarity with chaos engineering, secrets-management architectures, DevSecOps, identity services, IAM, enterprise directories, or security-focused SaaS
  • Located in India and authorized to work there
  • Fluent written and spoken English

Benefits

  • Work remotely from within India
  • Collaborate with colleagues across more than 15 countries
  • Access professional development and knowledge-sharing opportunities
  • Join a supportive and collaborative workplace
  • Equal opportunity employment

Related Jobs

Knowtion Health

Remote Talent Acquisition Manager

Knowtion Health

Oversee recruiting systems, requisitions, analytics, and contingent workforce operations at Knowtion Health. Help support scalable hiring processes for a growing healthcare company.

Open