FactSet
FactSet
FactSet develops integrated financial data, analytics, and software for investment professionals and financial institutions. Its platform brings together information and tools for investment research, portfolio management, asset management, and risk analysis, helping teams work with complex data and evaluate opportunities across asset classes. FactSet also applies artificial intelligence and data technologies to improve how financial firms access insights, analyze markets, and make investment decisions. The company serves asset managers, consultants, hedge funds, and other institutions worldwide.

Lead Site Reliability Engineer, Kubernetes - Hybrid UK

Lead site reliability engineering for FactSet’s financial data and analytics platform, with a focus on scalable Kubernetes infrastructure. Automate operations, resolve incidents, and strengthen production performance and resilience.

Description

  • Monitor and improve the reliability and availability of production systems
  • Resolve incidents and lead post-mortems to help prevent recurrence
  • Establish and monitor Service Level Objectives and Service Level Indicators
  • Partner with development teams to embed reliability into services
  • Create automation that reduces operational toil and improves efficiency
  • Join the on-call rotation for critical systems
  • Support capacity planning and performance optimization
  • Maintain system documentation, processes, and operational runbooks
  • Develop resilient infrastructure and promote sound engineering practices

Requirements

  • Practical experience deploying, managing, and troubleshooting Kubernetes workloads
  • Strong knowledge of Pods, Deployments, Services, ConfigMaps, and Ingress
  • Experience administering and managing Kubernetes clusters
  • Working familiarity with Helm for packaging and deploying applications
  • Understanding of Kubernetes networking, storage, and security practices
  • Bachelor’s degree in computer science or a related field
  • Able to work in a hybrid arrangement
  • Fluent written and spoken English
  • Strong analytical and problem-solving abilities
  • Clear and effective communication skills
  • Proactive approach to automation and continuous improvement
  • Able to perform effectively under pressure during incident response
  • Commitment to blameless practices and ongoing learning
  • Experience with cloud platforms including AWS, GCP, or Azure
  • Experience with CI/CD tools such as GitHub Actions, ArgoCD, or Harness
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Coralogix, or OpenTelemetry
  • Experience using Infrastructure as Code tools such as Terraform or Pulumi
  • Experience with configuration management tools including Ansible, Puppet, or Chef
  • Experience programming or scripting with Python, Go, or Bash
  • Experience contributing to open-source projects
  • Familiarity with SRE principles from the Google SRE handbook
  • Previous experience in DevOps or Platform Engineering

Benefits

  • Hybrid working arrangement
  • Support for on-call coverage of critical systems
  • Emphasis on continuous learning
  • Blameless engineering culture
  • Opportunity to contribute to open-source projects
  • Collaboration with teams across the globe
  • Equal employment opportunity

Related Jobs

HarperCollins Publishers UK

Editorial Assistant, Fantasy and Science Fiction

HarperCollins Publishers UK
1,001 – 5,000 Employees
eCommerceMediaRetail

Support HarperCollins’ HarperVoyager and HarperMagpie imprints with editorial administration and title production. The role spans metadata, proofreading, author and agent liaison, and publishing workflows.

Open