Backblaze
Backblaze
Backblaze व्यवसायों और व्यक्तिगत उपयोगकर्ताओं के लिए क्लाउड स्टोरेज और डेटा सुरक्षा सेवाएँ प्रदान करता है। इसका B2 Cloud Storage प्लेटफ़ॉर्म डेटा के प्रबंधन और सुरक्षा के लिए S3-संगत ऑब्जेक्ट स्टोरेज उपलब्ध कराता है, जबकि इसकी स्वचालित और असीमित कंप्यूटर बैकअप सेवा सिस्टम या फ़ाइलें खो जाने पर उन्हें पुनर्प्राप्त करने में सहायता करती है। कंपनी अन्य एप्लिकेशनों के साथ एकीकरण की सुविधा भी देती है, जिससे इसके उत्पाद डेटा वर्कफ़्लो बनाने वाली टीमों और सरल बैकअप समाधान चाहने वाले ग्राहकों, दोनों के लिए उपयुक्त हैं।

Site Reliability Engineer III, Database — United States Remote

Lead reliability, automation, incident response, and operational readiness for Backblaze’s Vitess and Cassandra production databases. Support the resilience of its cloud storage platform through database architecture, recovery, monitoring, and infrastructure operations.

विवरण

  • Design, deploy, and own highly available database architectures for Vitess, distributed MySQL, and Cassandra
  • Create and maintain procedures, runbooks, and escalation guidance for Level 1 and Level 2 SRE Database Engineers
  • Improve database performance through query tuning, indexing, schema design, and capacity planning
  • Own database backup, recovery, replication, and disaster recovery strategies
  • Execute and validate disaster recovery tests and database recovery procedures
  • Lead database security, access control, patching, hardening, and compliance practices
  • Collaborate with DBA and Data Infrastructure teams on resharding, capacity, replication, and sharded MySQL architecture
  • Help maintain the availability and durability of critical production services
  • Track service health with SLIs, SLOs, error budgets, monitoring, logging, and alerting systems
  • Join on-call rotations and contribute to incident response, root cause analysis, and post-incident reviews
  • Act as an escalation resource for complex production database incidents
  • Automate database administration and operational processes
  • Contribute to monitoring, logging, and alerting systems using Prometheus, Grafana, Catchpoint, and ELK
  • Connect operational runbooks and incident-response workflows with FireHydrant
  • Use CI/CD, configuration management, and infrastructure-as-code tools including Terraform, Ansible, and Jenkins
  • Write operational scripts in Bash, Python, Go, or comparable languages
  • Run and troubleshoot containerized production systems with Kubernetes and Docker
  • Lead Production Readiness Reviews and help prepare new database-backed services for operation
  • Create training plans, onboarding resources, and technical documentation for Level 1 and Level 2 SRE Database Engineers
  • Work with Engineering, Product, Operations, and DBA/Data Infrastructure teams on reliability programs
  • Support capacity planning, disaster recovery exercises, database migrations, and infrastructure initiatives
  • Coordinate with vendors and service providers to resolve issues and monitor SLA performance
  • Investigate and resolve production database, infrastructure, and service incidents
  • Troubleshoot and escalate database, Linux, networking, application, and infrastructure problems
  • Address recurring issues with durable corrective actions that strengthen reliability

आवश्यकताएँ

  • Bring 6–8 years of experience in site reliability, systems engineering, infrastructure operations, database engineering, or a related discipline, including substantial production database support
  • Demonstrate deep practical experience with MySQL and distributed or sharded database systems
  • Production experience with Vitess is strongly preferred
  • Have experience administering and supporting NoSQL databases such as Cassandra
  • Know how to design highly available database architectures, replication topologies, backup strategies, and disaster recovery processes
  • Use strong SQL skills for query analysis, indexing, schema design, and troubleshooting
  • Bring solid Linux administration and troubleshooting capabilities
  • Have experience with security-focused operations, including patching, hardening, access control, and vulnerability remediation
  • Understand reliability practices such as monitoring, alerting, incident response, root cause analysis, SLIs, SLOs, and error budgets
  • Have worked with containers and orchestration platforms including Kubernetes and Docker
  • Be comfortable operating Kubernetes and Vitess environments with tools such as kubectl, mysqlsh, and Vitess keyspaces
  • Have experience with infrastructure and configuration management tools including Terraform, Ansible, Jenkins, and HashiCorp products such as Vault and Nomad
  • Be proficient in at least one scripting language, such as Python, Bash, or Go
  • Have established operational procedures, runbooks, documentation, and escalation processes
  • Have mentored, trained, or onboarded engineers in complex technical environments
  • Experience in SaaS, cloud services, service-provider, or large-scale distributed-systems environments is preferred
  • Experience with AWS, GCP, Azure, or comparable cloud platforms is preferred
  • Familiarity with ITIL/OSS practices and SLA/SLO management is preferred
  • Hold a bachelor’s degree in Computer Science, Engineering, or a related field, or offer equivalent professional experience

लाभ

  • Family healthcare coverage, including dental and vision
  • Competitive compensation and a 401(k) plan
  • RSU grants for full-time employees
  • Employee stock purchase program
  • Flexible vacation policy
  • Maternity and paternity leave
  • MacBook Pro for work plus a workstation personalization stipend
  • Childcare bonus for human children
  • Fertility treatment and support
  • Learning and development program
  • Commuter benefits
  • A culture that promotes healthy work-life balance

संबंधित नौकरियाँ

4M Analytics

Senior Computer Vision Algorithm Engineer

4M Analytics

Develop computer vision and machine learning systems for 4M Analytics’ subsurface utility mapping platform. Own algorithms from research through production in partnership with data engineering and product teams.

खोलें
Joom

Enterprise Customer Success Manager, Brazil (Remote)

Joom

Lead renewals, retention, and account growth for JoomPulse enterprise customers. Turn Mercado Livre and Shopee analytics into practical recommendations for marketplace sellers.

खोलें
Cyera

Regional Sales Director, Identity Security

Cyera

Lead Cyera’s Pacific Northwest sales team, grow the territory, and exceed quotas for its AI and data security platform.

खोलें
Terumo Medical Corporation

Region Manager, Terumo Interventional Systems Sales

Terumo Medical Corporation

Lead medical device sales across a North Central New Jersey region, managing field teams and hospital relationships. Drive regional revenue, sales performance, and compliant promotion of Terumo Interventional Systems products.

खोलें
Dream

Senior Program Manager, Sovereign AI (Hybrid, Israel)

Dream

Lead Dream’s Sovereign AI roadmap for government and critical infrastructure customers, coordinating research, engineering, product, launches, and AI-enabled operations.

खोलें
Claroty
SKF Group

Global Product Engineer, Bearings — SKF, Bengaluru

SKF Group

Design bearing products, models, drawings, and variants at SKF, a global provider of rotating equipment solutions. Support modular design, testing, and product lifecycle management.

खोलें
Webbing

Mobile Fraud Analyst – Bucharest Hybrid

Webbing

Investigate and prevent telecom fraud across Webbing’s global MVNO connectivity and IoT services. Analyze traffic, roaming, SIM, voice, SMS, and data activity for anomalies.

खोलें
The Clorox Company

Finance Manager, Manufacturing Finance

The Clorox Company

Manage manufacturing finance activities for Clorox, including cost consolidation, forecasting, and monthly accounting close. Partner with plant sites, Procurement, and Supply Planning in Mexico.

खोलें