Broadridge
Broadridge
Broadridge izstrādā tehnoloģiskus un operacionālus risinājumus finanšu pakalpojumu nozarei. Tās platforma un pakalpojumi atbalsta aktīvu pārvaldītājus, institucionālos investorus un citas finanšu organizācijas kapitāla tirgu operācijās, investoru un klientu saziņā, datu pārvaldībā, analītikā, portfeļu administrēšanā, regulatīvajos procesos un operacionālā riska pārvaldībā. Uzņēmums apkalpo arī saistītās nozares, tostarp patērētāju finansēšanu, apdrošināšanu un telekomunikācijas, palīdzot organizācijām koordinēt sarežģītas informācijas plūsmas, uzlabot pārraudzību un efektīvāk īstenot būtiskus uzņēmējdarbības procesus.

Global Site Reliability Engineering Manager - Broadridge (New York Hybrid)

Lead Broadridge’s global SRE function for a financial-services platform, with responsibility for Java, AWS, Kafka, PostgreSQL, automation, and production resilience. Build engineering capability, improve release operations, and strengthen reliability across critical distributed systems.

Apraksts

  • Build and manage a high-performing SRE organization through hiring, coaching, performance management, career development, and succession planning.
  • Set the SRE strategy and delivery roadmap, converting business and platform priorities into measurable gains in reliability, release quality, automation, and scalability.
  • Direct architecture reviews, code reviews, complex troubleshooting, automation initiatives, and tooling decisions.
  • Create reusable AWS infrastructure-as-code patterns, environment configurations, automated provisioning workflows, and recovery capabilities.
  • Improve Kafka reliability and event-processing performance across partitioning, consumer groups, schema evolution, delivery semantics, replay, and failure recovery.
  • Increase PostgreSQL performance and resilience through schema design, indexing, query optimization, transaction management, connection pooling, migrations, and recovery testing.
  • Lead release engineering and deployment readiness by advancing CI/CD, automated testing, security validation, artifact traceability, production checks, and rollback or roll-forward procedures.
  • Establish service-level indicators, service-level objectives, and error-budget practices, using metrics, logs, and distributed traces to improve service health.
  • Lower operational toil with reusable tooling, self-service workflows, and controlled remediation mechanisms.
  • Promote responsible AI use in code and test development, infrastructure reviews, knowledge retrieval, and incident investigations.
  • Coordinate major incident response, communicate business impact and recovery status, facilitate blameless reviews, and ensure corrective actions are completed.
  • Assess capacity, failover, backup restoration, and disaster-recovery capabilities against agreed objectives.
  • Advise senior stakeholders on technical risk, investment trade-offs, production readiness, secure delivery, and dependable service operations.

Prasības

  • At least 10 years of experience in software engineering or closely related production engineering.
  • Deep experience building and running Java services in production environments.
  • Experience managing engineering teams, including hiring, performance management, talent development, resource planning, and accountability for complex technical delivery.
  • Track record of leading senior engineers and developing technical leads.
  • Advanced Java and Spring Boot expertise spanning API design, concurrency, performance tuning, automated testing, and secure coding.
  • Hands-on AWS experience with IAM, networking, observability, infrastructure as code, and ECS, EKS, or Lambda.
  • Production Kafka experience covering event-driven architecture, consumer groups, partitioning, schema evolution, delivery semantics, and failure recovery.
  • Strong PostgreSQL capability in schema design, indexing, query optimization, transactions, migrations, and operational troubleshooting.
  • Ability to develop maintainable automation and operational tools with version control, peer review, automated testing, and reusable design practices.
  • Experience with CI/CD, release engineering, automated testing, monitoring, incident response, and reliable distributed-system design.
  • Knowledge of service-level objectives, capacity planning, production readiness, recovery strategies, and delivery and reliability metrics.
  • Ability to lead technical design reviews, mentor engineers, and explain architectural, operational, and delivery trade-offs.
  • Preferred experience with Kubernetes, Terraform, OpenTelemetry, Kafka Connect, and PostgreSQL replication or high availability.
  • Preferred background working with financial-services platforms or other regulated, business-critical distributed systems.

Priekšrocības

  • Eligibility for bonus compensation.
  • Comprehensive benefits package.
  • Equal employment opportunity.
  • Reasonable accommodations are available throughout the application and hiring process.

Saistītās vakances

Knowtion Health

Remote Talent Acquisition Manager

Knowtion Health

Oversee recruiting systems, requisitions, analytics, and contingent workforce operations at Knowtion Health. Help support scalable hiring processes for a growing healthcare company.

Atvērt
Napco National

Purchasing Coordinator, Saudi Arabia (Remote)

Napco National

Coordinate purchasing operations for a manufacturing business in Saudi Arabia, from purchase orders and supplier deliveries to customs paperwork and material transfers. Support shipment clearance, invoice processing, supplier claims, and product certificate renewals.

Atvērt
Terumo Medical Corporation

Region Manager, Terumo Interventional Systems Sales

Terumo Medical Corporation

Lead medical device sales across a North Central New Jersey region, managing field teams and hospital relationships. Drive regional revenue, sales performance, and compliant promotion of Terumo Interventional Systems products.

Atvērt