NatWest Group
NatWest Group
NatWest Group is a UK-based banking and financial services organisation serving personal, commercial and institutional customers. Its work spans banking, finance and fintech, with services that include commercial banking, wealth management and risk management. The group also highlights its relationships with customers and communities, alongside opportunities for people building careers across its banking and technology-focused teams.

Site Reliability Engineer (AVP) - NatWest Group, Gurugram

Lead reliability engineering for AWS-based financial services applications in Gurugram. Automate operations, improve observability, manage incidents, and support dependable service delivery.

Description

  • Ensure applications remain stable, resilient, and reliable to reduce disruption across customer and colleague journeys.
  • Find opportunities to automate repetitive manual work.
  • Implement observability solutions and build an application-wide understanding of customer and colleague journeys.
  • Partner with feature teams to achieve agreed service-level objectives.
  • Set error budgets that balance operational risk with reliability.
  • Improve the structure and effectiveness of release processes.
  • Scale systems sustainably through automation and reliability-led improvements.
  • Coach colleagues and the broader team, taking the lead when needed.
  • Offer ideas and innovations that support immediate and long-term objectives.
  • Assess, balance, and manage potential risks.
  • Maintain the daily health of production and non-production environments and handle incidents.
  • Apply technical expertise to define appropriate risk tolerance for products and services.
  • Keep teams, customers, and stakeholders informed of incident progress.

Requirements

  • Strong understanding of reliability systems thinking.
  • Professional experience in software engineering.
  • Experience applying data-driven and scientific methods to investigate facts and identify findings.
  • Knowledge of financial services.
  • Ability to assess broader business impact, risk, and opportunity, and connect related outputs and processes.
  • Experience with AWS services including EMR, Airflow, S3, EKS, EC2, EMR Serverless, Lambda, and CloudWatch.
  • Hands-on programming experience with Spark, Python, and shell scripting.
  • Experience using DevOps tools such as GitLab, GitLab CI, and Artifactory.
  • Practical experience with Docker, Prometheus, and Grafana.
  • Ability to troubleshoot complex data problems using Splunk, Spark, and CloudWatch audit logs.
  • Experience using tools and technologies throughout the software development lifecycle.
  • Strong communication skills and a proactive approach to working with diverse stakeholders.

Benefits

  • Access to opportunities for personal and career growth.
  • A supportive working environment.

Related Jobs