Betfair Romania Development
Betfair Romania Development
1,001 – 5,000 Employees
ConsultingHealthcareMarketing
Betfair Romania Development is Flutter Entertainment’s technology hub in Cluj-Napoca, bringing together more than 1,900 people who support some of the group’s global betting and gaming brands. Its teams work across software development, online gaming, and customer support for products including Betfair, PokerStars, and Paddy Power, serving customers in markets around the world. The company’s work focuses on building engaging sports and casino betting experiences with safety as a central consideration.

Senior Reliability Engineer, Romania (Remote)

Improve the reliability of FanDuel’s sports technology platform through SLOs, observability, incident learning, and automation. Collaborate with application and platform teams in Cluj-Napoca.

Description

  • Collaborate with application and platform teams to map service architecture, dependencies, critical customer journeys, and reliability risks
  • Establish meaningful SLIs and SLOs that connect service performance with customer and business outcomes
  • Help teams adopt error budgets and use reliability data to guide engineering priorities
  • Evaluate services for production readiness across observability, alerting, SLOs, runbooks, dependencies, capacity, and recovery
  • Apply reliability maturity assessments and scorecards to focus improvement efforts where they have the greatest impact
  • Participate in incident response and diagnose complex production problems using logs, metrics, traces, dependencies, and deployment data
  • Analyse incidents and recurring operational problems to uncover systemic risks and lasting solutions
  • Strengthen runbooks, alerting, escalation procedures, and day-to-day operational practices
  • Create automation and self-service tools that reduce engineering toil
  • Develop tooling for reliability processes, operational readiness, investigation, and remediation
  • Work with Observability Engineering to establish the telemetry required for reliability measurement and troubleshooting
  • Turn performance testing, capacity analysis, Gamedays, chaos experiments, and failover exercises into concrete reliability improvements
  • Help prepare for peak events by assessing service health, SLOs, dependencies, capacity exposure, and reliability findings
  • Increase the resilience of critical customer journeys through monitoring, dependency analysis, failure handling, and recovery practices
  • Contribute standards, patterns, documentation, and reusable golden paths for Reliability Engineering
  • Embed reliability practices in CI/CD and developer workflows through SLO-as-Code, telemetry validation, production-readiness checks, and automated controls
  • Apply AI-assisted investigation and automation to speed troubleshooting and reduce manual work
  • Share expertise and help engineers build stronger SRE and production-engineering practices

Requirements

  • Demonstrated hands-on experience in Site Reliability Engineering, Reliability Engineering, Platform Engineering, DevOps, or Production Engineering
  • Solid command of SRE fundamentals, including SLIs, SLOs, error budgets, incident management, operational readiness, toil reduction, and automation
  • Practical experience defining or operating with SLIs and SLOs for production services
  • Strong troubleshooting capability across applications, infrastructure, networks, databases, and service dependencies
  • Working knowledge of distributed systems and reliability patterns including retries, timeouts, circuit breakers, graceful degradation, redundancy, backpressure, and failure isolation
  • Experience with Datadog or a comparable observability platform covering logs, metrics, traces, APM, dashboards, monitors, and synthetic monitoring
  • Experience contributing to incident response, post-incident reviews, and effective follow-up work
  • Hands-on experience with Kubernetes and cloud infrastructure, preferably AWS
  • Experience using infrastructure-as-code tools such as Terraform in modern CI/CD environments
  • Strong automation and software engineering ability, with proficiency in a modern language such as Go, Java, Python, or JavaScript
  • Track record of replacing repetitive operational work with automation or self-service
  • Working knowledge of performance engineering, capacity management, resilience testing, failover, or chaos engineering
  • Ability to interpret application architecture and identify reliability risks across service and infrastructure dependencies
  • Strong analytical thinking and problem-solving ability
  • Clear communication skills, including the ability to explain reliability concepts and recommendations to engineers and engineering leaders
  • Ability to work across teams, manage competing priorities, and deliver measurable outcomes
  • An approach centred on automation, continuous improvement, knowledge sharing, and systemic problem-solving

Benefits

  • Choice of hybrid or remote working
  • €1,000 annual self-development allowance
  • Company share scheme
  • 25 days of annual leave
  • Up to 20 days per year working abroad
  • Five personal days each year
  • Flexible benefits covering travel, sports, and hobbies
  • Extended health, dental, and travel insurance
  • Tailored wellbeing programmes
  • Career growth sessions
  • Access to thousands of Udemy online courses
  • A range of engaging office events

Related Jobs

ReSus Consult GmbH

Sales Director, HVAC and Plumbing (SHK)

ReSus Consult GmbH
DEGermany
€120,000 – €180,000 / year
HybridFull-timeLeadGerman RequiredSales

Lead a regional portfolio of five to twelve SHK trade businesses, with responsibility for budgets and operational development. Build regional collaboration through digitalization, shared capacity, larger projects and best-practice exchange.

Open
smartkündigen OHG

Senior Sales Manager (German-speaking), Remote

smartkündigen OHG
11 – 50 Employees
B2CProductivitySaaS

Advise customers, grow existing accounts, and close sales for smartkündigen’s digital contract cancellation service. Work fully remotely without cold calling.

Open
smartkündigen OHG

Senior Sales Manager, Remote

smartkündigen OHG
11 – 50 Employees
B2CProductivitySaaS

Advise customers, grow existing accounts, and close sales for smartkündigen’s digital contract cancellation service. Work fully remotely worldwide, handling inbound inquiries without cold calling.

Open