RELX
RELX
RELX develops information-based analytics and decision tools for professional and business customers worldwide. Its products combine specialized content, data, and technology to support better decisions and stronger outcomes across Risk, Scientific, Technical & Medical, Legal, and Exhibitions. RELX also operates across consulting, healthcare, and insurance-related markets, helping organizations improve productivity, manage risk, advance research, support legal work, and conduct more effective business transactions.

Senior Site Reliability Engineer AWS Terraform Remote United States

Improve AWS platform reliability, automation, and observability for Elsevier’s scientific and medical information services. Support secure AI-enabled capabilities and engineering teams in a senior SRE role.

Description

  • Build monitoring queries and define service-level baselines
  • Assist senior engineers with incident response
  • Contribute to post-incident reviews and root-cause investigations
  • Take part in disaster recovery exercises
  • Develop automation and run code in production environments
  • Maintain and expand SRE knowledge documentation
  • Help deploy, monitor, and maintain services that integrate AI tools
  • Assist architects and senior engineers with infrastructure topology diagrams and deployment workflows
  • Evaluate service availability, reliability, and recoverability in non-production environments
  • Lead complex reliability programs and automate work to reduce operational toil
  • Create solutions that increase availability, simplify operations, and strengthen recovery capabilities
  • Work with engineering teams and host-function stakeholders
  • Support handover and capability development so teams can own and operate solutions after squad transition

Requirements

  • Advanced Terraform expertise, including modules, providers, state management, lifecycle controls, drift detection, safe refactoring, remote state, locking, and cross-stack dependencies
  • Production experience managing multi-account, multi-region AWS environments using ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch
  • Experience developing and troubleshooting reusable GitHub Actions CI/CD workflows, OIDC authentication, approval gates, runners, Terraform deployments, application releases, and migration pipelines
  • Knowledge of Docker, ECR, ECS task definitions and services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks
  • Proficiency in AWS networking and security, including VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security practices
  • Incident response and observability skills using logs, metrics, alarms, deployment history, root-cause analysis, rollback decisions, and operational runbooks
  • Strong Linux and Git foundations, with Bash or Python scripting for AWS CLI automation, CI/CD, and operational tooling
  • Production experience integrating and operating AI services and APIs, including monitoring, reliability, and security for AI-powered features
  • Ability to support multiple engineering teams, troubleshoot infrastructure and application layers, document solutions, and enable secure self-service

Benefits

  • Annual performance incentive bonus
  • Benefits tailored to the employee’s country
  • Support for accommodations or adjustments during the hiring process

Related Jobs

Kreato Global | BPO and Language Solutions

Remote English-Spanish OPI/VRI Interpreter

Kreato Global | BPO and Language Solutions
201 – 500 Employees
HealthcareHospitalityLogistics

Interpret remotely between English- and Spanish-speaking people in medical, financial, social service, and customer care settings. Provide language support for Kreato Global across Latin America.

Open