Backblaze
Backblaze
Backblaze provides cloud storage and data protection services for businesses and individuals. Its B2 Cloud Storage platform delivers S3-compatible object storage for managing and protecting data, while its automatic, unlimited computer backup service supports recovery when systems or files are lost. The company also enables integrations with other applications, making its products suitable for teams building data workflows as well as customers seeking straightforward backup solutions.

Site Reliability Engineer III, Database — United States Remote

Lead reliability, automation, incident response, and operational readiness for Backblaze’s Vitess and Cassandra production databases. Support the resilience of its cloud storage platform through database architecture, recovery, monitoring, and infrastructure operations.

Description

  • Design, deploy, and own highly available database architectures for Vitess, distributed MySQL, and Cassandra
  • Create and maintain procedures, runbooks, and escalation guidance for Level 1 and Level 2 SRE Database Engineers
  • Improve database performance through query tuning, indexing, schema design, and capacity planning
  • Own database backup, recovery, replication, and disaster recovery strategies
  • Execute and validate disaster recovery tests and database recovery procedures
  • Lead database security, access control, patching, hardening, and compliance practices
  • Collaborate with DBA and Data Infrastructure teams on resharding, capacity, replication, and sharded MySQL architecture
  • Help maintain the availability and durability of critical production services
  • Track service health with SLIs, SLOs, error budgets, monitoring, logging, and alerting systems
  • Join on-call rotations and contribute to incident response, root cause analysis, and post-incident reviews
  • Act as an escalation resource for complex production database incidents
  • Automate database administration and operational processes
  • Contribute to monitoring, logging, and alerting systems using Prometheus, Grafana, Catchpoint, and ELK
  • Connect operational runbooks and incident-response workflows with FireHydrant
  • Use CI/CD, configuration management, and infrastructure-as-code tools including Terraform, Ansible, and Jenkins
  • Write operational scripts in Bash, Python, Go, or comparable languages
  • Run and troubleshoot containerized production systems with Kubernetes and Docker
  • Lead Production Readiness Reviews and help prepare new database-backed services for operation
  • Create training plans, onboarding resources, and technical documentation for Level 1 and Level 2 SRE Database Engineers
  • Work with Engineering, Product, Operations, and DBA/Data Infrastructure teams on reliability programs
  • Support capacity planning, disaster recovery exercises, database migrations, and infrastructure initiatives
  • Coordinate with vendors and service providers to resolve issues and monitor SLA performance
  • Investigate and resolve production database, infrastructure, and service incidents
  • Troubleshoot and escalate database, Linux, networking, application, and infrastructure problems
  • Address recurring issues with durable corrective actions that strengthen reliability

Requirements

  • Bring 6–8 years of experience in site reliability, systems engineering, infrastructure operations, database engineering, or a related discipline, including substantial production database support
  • Demonstrate deep practical experience with MySQL and distributed or sharded database systems
  • Production experience with Vitess is strongly preferred
  • Have experience administering and supporting NoSQL databases such as Cassandra
  • Know how to design highly available database architectures, replication topologies, backup strategies, and disaster recovery processes
  • Use strong SQL skills for query analysis, indexing, schema design, and troubleshooting
  • Bring solid Linux administration and troubleshooting capabilities
  • Have experience with security-focused operations, including patching, hardening, access control, and vulnerability remediation
  • Understand reliability practices such as monitoring, alerting, incident response, root cause analysis, SLIs, SLOs, and error budgets
  • Have worked with containers and orchestration platforms including Kubernetes and Docker
  • Be comfortable operating Kubernetes and Vitess environments with tools such as kubectl, mysqlsh, and Vitess keyspaces
  • Have experience with infrastructure and configuration management tools including Terraform, Ansible, Jenkins, and HashiCorp products such as Vault and Nomad
  • Be proficient in at least one scripting language, such as Python, Bash, or Go
  • Have established operational procedures, runbooks, documentation, and escalation processes
  • Have mentored, trained, or onboarded engineers in complex technical environments
  • Experience in SaaS, cloud services, service-provider, or large-scale distributed-systems environments is preferred
  • Experience with AWS, GCP, Azure, or comparable cloud platforms is preferred
  • Familiarity with ITIL/OSS practices and SLA/SLO management is preferred
  • Hold a bachelor’s degree in Computer Science, Engineering, or a related field, or offer equivalent professional experience

Benefits

  • Family healthcare coverage, including dental and vision
  • Competitive compensation and a 401(k) plan
  • RSU grants for full-time employees
  • Employee stock purchase program
  • Flexible vacation policy
  • Maternity and paternity leave
  • MacBook Pro for work plus a workstation personalization stipend
  • Childcare bonus for human children
  • Fertility treatment and support
  • Learning and development program
  • Commuter benefits
  • A culture that promotes healthy work-life balance

Related Jobs