RELX
RELX
RELX develops information-based analytics and decision tools for professional and business customers worldwide. Its products combine specialized content, data, and technology to support better decisions and stronger outcomes across Risk, Scientific, Technical & Medical, Legal, and Exhibitions. RELX also operates across consulting, healthcare, and insurance-related markets, helping organizations improve productivity, manage risk, advance research, support legal work, and conduct more effective business transactions.

Site Reliability Engineering Lead at RELX (Remote, United States)

Lead SRE teams and reliability programs for LexisNexis Risk Solutions’ cloud-based risk platforms. Drive Kubernetes, Terraform, Azure, automation, incident response, and operational excellence.

Description

  • Lead, mentor, and develop SREs through regular 1:1s, performance reviews, and career planning
  • Manage hiring, onboarding, team capacity, and resourcing decisions
  • Establish team objectives, prioritize the backlog, and lead planning activities
  • Promote blameless incident learning and collaboration across Development, Security, and Product
  • Direct reliability improvements across infrastructure and services
  • Coordinate incident response and ongoing service enhancements
  • Advance platform-wide automation and operational excellence
  • Help build scalable, secure, and resilient cloud-native environments
  • Manage a small to medium-sized team, including performance, compensation, and recruitment responsibilities
  • Lead post-incident reviews and ensure root-cause analyses are completed promptly
  • Evaluate and resolve issues whose effects extend beyond the immediate team

Requirements

  • Expertise in Kubernetes cluster architecture, upgrades, autoscaling, security hardening, and large-scale troubleshooting
  • Advanced Terraform experience covering modular infrastructure as code, state management, multi-environment provisioning, and policy as code
  • In-depth Azure knowledge spanning compute, networking, Azure AD identity, storage, and cost optimization
  • Experience designing and scaling CI/CD pipelines with GitHub Actions, release strategies, and automated rollback
  • Experience with Prometheus, Grafana, OpenTelemetry, and SLO, SLA, and error-budget practices
  • Strong automation capabilities aimed at reducing toil through self-healing systems and automated infrastructure
  • Advanced Python, Bash, and/or PowerShell skills for tooling and automation
  • Deep understanding of TCP/IP, DNS, load balancing, VPNs, and cloud-native networking
  • Background in SRE, DevOps, or infrastructure engineering, including engineering team leadership
  • Demonstrated success leading incident response and improving service reliability

Benefits

  • Annual incentive bonus
  • Benefits tailored to the employee’s country
  • Support for accommodations or adjustments during the hiring process

Related Jobs