EX Squared
EX Squared
EX Squared is a digital product and experience consultancy serving large B2B enterprises. The company helps organizations turn AI and digital strategies into production-ready products through consulting, UX/UI design, engineering, IT infrastructure, and multi-channel experience development. Its expertise includes generative AI, machine learning, computer vision, and natural language processing, alongside work with enterprise platforms such as Optimizely and Sitecore. Founded in 2000, EX Squared has delivered technology initiatives since 2001 and supports teams working on enterprise-scale digital products and advanced AI solutions.

Senior Site Reliability Engineer Remote in Costa Rica

EX Squared is seeking a Site Reliability Engineer in Costa Rica to automate, monitor, and strengthen critical services. The role focuses on incident management, cloud reliability, and CI/CD deployments.

Description

  • Design and implement automation that reduces repetitive manual work in production environments.
  • Build and improve observability practices using metrics, logs, and traces.
  • Define and maintain SLIs and SLOs in partnership with development teams and business stakeholders.
  • Respond to incidents and participate in on-call rotations, diagnosing issues and restoring services.
  • Lead and document post-incident reviews and root cause analyses, identifying preventive improvements.
  • Promote the “you build it, you run it” operating model.
  • Design and maintain change-control processes and deployments integrated with CI/CD pipelines.
  • Assess service capacity, performance, and load behavior, recommending architectural or configuration improvements.
  • Evaluate and strengthen platform and service resilience, availability, and failure recovery.
  • Document Site Reliability Engineering standards, procedures, runbooks, and best practices.
  • Collaborate closely with development and platform teams.

Requirements

  • Professional experience in Site Reliability Engineering, DevOps, cloud, infrastructure, systems engineering, or operations.
  • Hands-on experience supporting critical services and production environments.
  • Strong Linux knowledge and experience troubleshooting infrastructure and applications.
  • Experience working with cloud computing platforms.
  • Experience with monitoring and observability through metrics, logs, traces, and alerting.
  • Background in production incident response, troubleshooting, and root cause analysis.
  • Knowledge of SLI, SLO, and SLA concepts and their application to reliability management.
  • Experience with automation and scripting.
  • Knowledge of CI/CD and modern deployment practices.
  • Experience with Infrastructure as Code, preferably Terraform.
  • Understanding of high availability, resilience, disaster recovery, and scalability principles.
  • Ability to collaborate with development, infrastructure, and operations teams.
  • AWS and cloud-native architecture experience is preferred.
  • Terraform and infrastructure automation experience is preferred.
  • Docker and Kubernetes experience is preferred.
  • Familiarity with observability tools such as Prometheus, Grafana, Datadog, Splunk, CloudWatch, or equivalent platforms is preferred.
  • Experience with high-availability and disaster-recovery strategies is preferred.
  • Experience analyzing capacity, performance, and scalability is preferred.
  • Experience participating in on-call rotations and managing critical incidents is preferred.
  • Certifications related to AWS, Cloud Operations, or Site Reliability Engineering are preferred.

Benefits

  • Private health insurance.
  • Solidarity association membership.
  • Vacation and public holidays in accordance with Costa Rican law.
  • Professional learning and development opportunities.
  • Participation in high-impact technology projects.

Related Jobs