Rentsync
Rentsync
51 – 200 Employees
ConsultingLogisticsMarketing
Rentsync is a North American proptech SaaS company building software for multifamily property marketing, leasing, and resident management. Its platform brings together marketing automation, AI-powered tools, listing syndication, customizable property websites, digital applications and payments, tenant portals, and analytics. Rentsync also provides agency-led digital marketing services, APIs, and integrations that help property management and multifamily marketing teams manage the tenant journey from attracting prospects through coordinating residents across the lease lifecycle.

Site Reliability Engineer (AWS/Kubernetes) — Canada Remote

Own production reliability for Rentsync’s rental-property software across AWS and Kubernetes. Lead incident response, observability, automation, performance testing, and infrastructure hardening.

Description

  • Respond to production alerts and incidents across services, from initial triage through resolution
  • Diagnose and resolve AWS and Kubernetes issues, including failed pods, resource exhaustion, faulty deployments, networking and DNS, databases, and caches
  • Roll back, scale, reconfigure, or patch infrastructure to restore service promptly
  • Escalate to development teams when code changes are required, with a well-defined diagnosis
  • Manage PagerDuty and business-hours incident response while improving MTTD and MTTR
  • Facilitate blameless post-mortems and coordinate resulting technical improvements
  • Automate runbooks and recurring operational tasks, including AI-assisted triage, investigation, and remediation
  • Develop and maintain Kubernetes and service observability with Prometheus/Mimir, Loki, Tempo, Grafana, and OpenTelemetry
  • Monitor production releases for regressions in latency, errors, and resource consumption
  • Build and maintain synthetic monitoring, smoke tests, health checks, and load or performance tests
  • Work with engineering teams to address performance and reliability issues and establish SLOs, SLIs, and error budgets
  • Strengthen the platform through Terraform, Kubernetes resource tuning, autoscaling, CI/CD controls, secrets management, and IAM improvements
  • Keep service documentation and architecture decisions current

Requirements

  • At least 3 years of experience in cloud engineering, DevOps, or SRE roles supporting production web applications
  • Hands-on experience responding to production incidents and resolving issues directly
  • Strong production AWS experience with EKS, EC2, RDS, VPC networking, IAM, and CloudWatch
  • Advanced production Kubernetes experience, including workload troubleshooting, debugging, and monitoring
  • Experience diagnosing and resolving performance and reliability issues with engineering teams
  • Experience with observability platforms such as Prometheus, Grafana, Loki, Datadog, or CloudWatch
  • Experience using on-call and alerting platforms such as PagerDuty
  • Experience creating automated production reliability checks, including synthetic, smoke, health, or load tests
  • Willingness to join a future after-hours on-call rotation
  • Experience with Terraform infrastructure as code and CI/CD pipelines
  • Strong Linux, networking, and container fundamentals
  • Scripting and automation skills in Bash, Python, or a similar language
  • Clear, composed communication during incidents and cross-team collaboration
  • Preferred experience includes AI tools for SRE, Azure or GCP, multiple technology stacks, AWS certification, LGTM or OpenTelemetry at scale, k6, Locust or JMeter, MySQL or PostgreSQL operations, Redis or Memcached tuning, Cloudflare, and cloud cost or capacity optimization

Benefits

  • Fully remote position
  • Preference may be given to candidates within reasonable commuting distance of one of the offices
  • Equal opportunity employer
  • Interview accommodations are available
  • A criminal background check may be conducted during the final interview stage

Related Jobs