Broadridge
Broadridge
Broadridge develops technology and operational solutions for the financial services industry. Its platform and services support asset managers, institutional investors, and other financial organizations with capital markets operations, investor and customer communications, data management, analytics, portfolio administration, regulatory processes, and operational risk. The company also serves adjacent sectors including consumer finance, insurance, and telecommunications, helping organizations coordinate complex information flows, improve oversight, and run essential business processes more efficiently.

Global Site Reliability Engineering Manager - Broadridge (New York Hybrid)

Lead Broadridge’s global SRE function for a financial-services platform, with responsibility for Java, AWS, Kafka, PostgreSQL, automation, and production resilience. Build engineering capability, improve release operations, and strengthen reliability across critical distributed systems.

Description

  • Build and manage a high-performing SRE organization through hiring, coaching, performance management, career development, and succession planning.
  • Set the SRE strategy and delivery roadmap, converting business and platform priorities into measurable gains in reliability, release quality, automation, and scalability.
  • Direct architecture reviews, code reviews, complex troubleshooting, automation initiatives, and tooling decisions.
  • Create reusable AWS infrastructure-as-code patterns, environment configurations, automated provisioning workflows, and recovery capabilities.
  • Improve Kafka reliability and event-processing performance across partitioning, consumer groups, schema evolution, delivery semantics, replay, and failure recovery.
  • Increase PostgreSQL performance and resilience through schema design, indexing, query optimization, transaction management, connection pooling, migrations, and recovery testing.
  • Lead release engineering and deployment readiness by advancing CI/CD, automated testing, security validation, artifact traceability, production checks, and rollback or roll-forward procedures.
  • Establish service-level indicators, service-level objectives, and error-budget practices, using metrics, logs, and distributed traces to improve service health.
  • Lower operational toil with reusable tooling, self-service workflows, and controlled remediation mechanisms.
  • Promote responsible AI use in code and test development, infrastructure reviews, knowledge retrieval, and incident investigations.
  • Coordinate major incident response, communicate business impact and recovery status, facilitate blameless reviews, and ensure corrective actions are completed.
  • Assess capacity, failover, backup restoration, and disaster-recovery capabilities against agreed objectives.
  • Advise senior stakeholders on technical risk, investment trade-offs, production readiness, secure delivery, and dependable service operations.

Requirements

  • At least 10 years of experience in software engineering or closely related production engineering.
  • Deep experience building and running Java services in production environments.
  • Experience managing engineering teams, including hiring, performance management, talent development, resource planning, and accountability for complex technical delivery.
  • Track record of leading senior engineers and developing technical leads.
  • Advanced Java and Spring Boot expertise spanning API design, concurrency, performance tuning, automated testing, and secure coding.
  • Hands-on AWS experience with IAM, networking, observability, infrastructure as code, and ECS, EKS, or Lambda.
  • Production Kafka experience covering event-driven architecture, consumer groups, partitioning, schema evolution, delivery semantics, and failure recovery.
  • Strong PostgreSQL capability in schema design, indexing, query optimization, transactions, migrations, and operational troubleshooting.
  • Ability to develop maintainable automation and operational tools with version control, peer review, automated testing, and reusable design practices.
  • Experience with CI/CD, release engineering, automated testing, monitoring, incident response, and reliable distributed-system design.
  • Knowledge of service-level objectives, capacity planning, production readiness, recovery strategies, and delivery and reliability metrics.
  • Ability to lead technical design reviews, mentor engineers, and explain architectural, operational, and delivery trade-offs.
  • Preferred experience with Kubernetes, Terraform, OpenTelemetry, Kafka Connect, and PostgreSQL replication or high availability.
  • Preferred background working with financial-services platforms or other regulated, business-critical distributed systems.

Benefits

  • Eligibility for bonus compensation.
  • Comprehensive benefits package.
  • Equal employment opportunity.
  • Reasonable accommodations are available throughout the application and hiring process.

Related Jobs

SPERTON - Where Great People Meet

Sales Executive, Elevators and Car Parking Systems

SPERTON - Where Great People Meet
51 – 200 Employees

Drive elevator and car parking system sales across Mumbai’s Western Region. Build client and dealer relationships, and manage deals from initial enquiry through project execution.

Open