Reap
Reap
201 – 500 Angajați
B2BCriptoFintech
Reap este o companie fintech care dezvoltă infrastructură financiară transfrontalieră pentru companii care operează în mai multe monede și piețe. Fondată în 2018, compania combină plățile transfrontaliere bazate pe stablecoin cu conturi pentru companii, carduri corporative Visa, gestionarea cheltuielilor în mai multe monede și API-uri pentru finanțare integrată. Platforma sa acceptă plăți globale în monede fiduciare, emiterea de carduri, automatizarea plăților și fluxuri de lucru pentru trezorerie, destinate companiilor care caută modalități mai flexibile de a-și gestiona finanțele internaționale. Reap activează în domeniile fintech, cripto și pe piețele B2B și are o echipă de 201–500 de angajați.

Senior Site Reliability Engineer - Reap Mexico Hybrid

Lead reliability engineering for Reap’s global stablecoin-powered payments platform, with a focus on observability, self-service infrastructure, and operational resilience. Work across AWS, Kubernetes, Terraform, and PCI DSS-regulated systems.

Descriere

  • Establish reliability targets with product and engineering teams through SLIs, SLOs, and error budgets.
  • Join the on-call rotation and strengthen incident response and blameless postmortem practices.
  • Modernize legacy infrastructure with Infrastructure as Code while reducing manual provisioning and configuration drift.
  • Unify Terraform into a governed codebase with consistent modules and automated drift detection.
  • Automate account provisioning and environment configuration across regions and services.
  • Create golden paths and self-service tools for provisioning, deployment, and observability.
  • Design production-like ephemeral environments with automated cleanup.
  • Develop observability across logs, metrics, distributed traces, and actionable alerts.
  • Lead cloud operations for PCI DSS-compliant financial systems, covering availability, failover, capacity, disaster recovery, and incident response.
  • Integrate secrets management, least-privilege IAM, network segmentation, and compliance safeguards into infrastructure.
  • Develop secure, auditable infrastructure for AI assistants and agents.
  • Collaborate with product and engineering teams while managing the platform as a product.

Cerințe

  • Experience setting SLOs, managing error budgets, or developing incident and postmortem practices.
  • Strong Linux and computer systems knowledge, including operating-system internals, networking, and administration.
  • Advanced Terraform expertise spanning state management, module architecture, and team-wide standards.
  • Experience leading a large-scale Infrastructure as Code migration or greenfield implementation.
  • Strong AWS background across multi-account, multi-region environments, including Control Tower, IAM, networking, RDS, and cost management.
  • Experience running ECS and Fargate, plus production Kubernetes environments at scale.
  • Understanding of Kubernetes cluster lifecycle management, upgrades, autoscaling, resource limits, and multi-tenant workload isolation.
  • Production experience with Lambda, SQS, and EventBridge, including retries, poison messages, ordering, and idempotency.
  • Experience with Argo CD or an equivalent platform and GitHub Actions CI/CD pipelines.
  • Hands-on observability experience using logging, metrics, and tracing tools such as New Relic, CloudWatch, Datadog, or Prometheus.
  • Ability to use Python, Go, or Bash for automation and tooling.
  • Experience owning services across design, delivery, observability, incident response, capacity planning, upgrades, and security patching.
  • Experience working in regulated environments, with familiarity with PCI DSS and financial compliance.
  • Ability to manage projects independently and build alignment without formal authority.
  • Clear plain-language communication skills and the ability to break complex problems into work for people or AI agents.
  • Substantial use of AI tools for coding, review, automation, incident analysis, and documentation.
  • Comfort responding to production incidents with significant potential impact.
  • Significant SRE, DevOps, or infrastructure engineering experience is a bonus.
  • Experience in fintech, payments, card issuing, or another regulated sector is a bonus.
  • Hands-on PCI DSS experience is a bonus.
  • Kubernetes or EKS experience in a PCI DSS environment is a bonus.
  • Experience designing and operating ephemeral or on-demand environment platforms is a bonus.
  • Experience building an internal developer platform, account factory, or landing zone from scratch is a bonus.
  • Experience with Ansible or similar configuration-management tools is a bonus.
  • AWS certifications are a bonus.
  • Curiosity about stablecoins and the Web2/Web3 intersection is a bonus.

Locuri de muncă similare

Knowtion Health

Remote Talent Acquisition Manager

Knowtion Health

Oversee recruiting systems, requisitions, analytics, and contingent workforce operations at Knowtion Health. Help support scalable hiring processes for a growing healthcare company.

Deschide
Napco National

Purchasing Coordinator, Saudi Arabia (Remote)

Napco National

Coordinate purchasing operations for a manufacturing business in Saudi Arabia, from purchase orders and supplier deliveries to customs paperwork and material transfers. Support shipment clearance, invoice processing, supplier claims, and product certificate renewals.

Deschide
Terumo Medical Corporation

Region Manager, Terumo Interventional Systems Sales

Terumo Medical Corporation

Lead medical device sales across a North Central New Jersey region, managing field teams and hospital relationships. Drive regional revenue, sales performance, and compliant promotion of Terumo Interventional Systems products.

Deschide