Stratus
Stratus
Stratus is a B2B technology and consulting firm focused on modernizing insurance operations for property and casualty insurers and Managing General Agents. Its teams help organizations replace or optimize legacy core systems, implement Guidewire and Insurity platforms, plan and execute Azure cloud migrations, and build stronger data strategies and governance. Stratus also supports AI readiness, application maintenance, quality assurance, platform selection, and embedded IT delivery through on-demand talent pods. The company’s work combines insurance domain expertise with practical technology transformation to help clients improve analytics, streamline operations, and bring new products to market more efficiently.

Senior Site Reliability Engineer - Remote in the United States

Senior SRE responsible for operationalizing reliability at Stratus, whose software digitizes MEP contractor workflows. The role covers SLOs, observability, incident response, and performance engineering across Azure production systems.

Description

  • Own service level indicators, objectives, and error budgets for customer-facing services
  • Develop measurement pipelines that report availability and latency against defined targets
  • Manage production observability through instrumentation standards, dashboards, and actionable alerts
  • Operate observability tooling across Azure AKS, Prometheus, Loki, Tempo, Grafana, Istio, and Flux
  • Establish incident response operations covering on-call rotations, paging, severity, ownership, escalation, and blameless post-mortems
  • Set rollback and recovery standards and validate them through recurring exercises
  • Lead performance and capacity engineering through query and index tuning, connection-pool sizing, capacity modeling, and graceful-degradation design
  • Partner with the team on k6 load, stress, spike, and soak testing
  • Define SLO-based endpoint latency thresholds and use test results as delivery gates
  • Lead production readiness reviews for instrumentation, alerting, failure modes, resource limits, and rollback plans
  • Follow post-mortem remediation through to verified production changes
  • Support business continuity and disaster recovery planning, including backup and restore validation, failover design, and recovery objectives
  • Coach engineering teams in instrumenting and operating their services
  • Create tooling, automation, instrumentation libraries, and production code fixes
  • Collaborate across engineering pods, platform and security teams, and customer-facing groups
  • Report to the Senior Director of Platform Engineering

Requirements

  • At least six years of professional engineering experience
  • Three or more years in a dedicated SRE or production engineering role within a B2B SaaS company
  • Proven ownership of production SLOs and error budgets
  • Advanced hands-on observability experience with Prometheus, Loki, Tempo, Grafana, or comparable tools
  • Practical experience serving as incident commander for significant incidents
  • Experience building on-call and escalation processes from the ground up
  • Strong database performance expertise, including query profiling, index design, connection pooling, and saturation diagnosis under load
  • MongoDB experience preferred; comparable depth with document or relational databases is acceptable
  • Production Kubernetes experience, preferably with AKS
  • Ability to troubleshoot pod scheduling, resource limits, networking, and service-mesh behavior; Istio experience preferred
  • Proficiency in at least one of Go, Python, C#, or TypeScript
  • Comfort working in a C#/.NET codebase
  • Experience with load and performance tools such as k6, JMeter, or Gatling
  • Production experience with Azure or AWS and an understanding of managed-service failure modes
  • Fluency with AI-assisted engineering tools and experience designing AI-enabled workflows
  • Excellent written communication for post-mortems, runbooks, and reliability reports
  • Sound judgment and composure under pressure
  • Experience establishing or maturing an SRE practice preferred
  • Experience with Sentry or a comparable application error-monitoring platform preferred
  • Experience operating event-driven and real-time systems preferred
  • Experience running MongoDB Atlas at production scale preferred
  • Experience with durable workflow orchestration platforms such as Temporal preferred
  • Background in multi-region or multi-zone architecture and disaster recovery design preferred
  • Experience with incident.io, PagerDuty, or comparable incident management platforms preferred
  • Familiarity with DORA metrics and reliability work within SOC 2 or NIST 800-171 environments preferred
  • Experience improving reliability for legacy monoliths preferred
  • Interest in MEP, BIM, AEC, or construction technology preferred
  • Previous experience at a Series B or growth-stage company preferred

Benefits

  • Comprehensive, competitive health benefits
  • 401(k) contributions with employer matching
  • Twenty days of annual paid time off
  • Primarily remote work arrangement
  • Occasional annual team onsite events

Related Jobs

Kreato Global | BPO and Language Solutions

Remote English-Spanish OPI/VRI Interpreter

Kreato Global | BPO and Language Solutions
201 – 500 Employees
HealthcareHospitalityLogistics

Interpret remotely between English- and Spanish-speaking people in medical, financial, social service, and customer care settings. Provide language support for Kreato Global across Latin America.

Open