Seismic
Seismic
Seismic este o platformă de marketing digital care ajută companiile să creeze conexiuni mai puternice cu publicul lor prin coduri QR, URL-uri scurtate, pagini de destinație optimizate pentru dispozitive mobile și instrumente de gestionare a linkurilor. Funcțiile sale de analiză oferă echipelor vizibilitate asupra interacțiunii, iar instrumentele de conținut personalizabile susțin campanii în industrii precum comerțul cu amănuntul, serviciile medicale și serviciile financiare. Seismic reunește comunicarea cu publicul, experiențele de brand și monitorizarea performanței într-o singură platformă.

Senior Site Reliability Engineer in Hyderabad (On-site)

Senior SRE role focused on improving reliability across AWS, Azure, IBM Cloud, and OCI for Seismic’s AI-powered revenue execution platform. The position covers automation, observability, incident response, and proactive risk reduction.

Descriere

  • Create and maintain automation and operational tooling that minimizes repetitive work
  • Advance reliability practices alongside Product and Engineering leadership
  • Help improve the full incident-management lifecycle, from detection and escalation through mitigation, communication, review, and corrective actions
  • Join the Global SRE team’s 12-hour follow-the-sun on-call rotation
  • Keep incident-management practices customer-focused, evidence-based, blameless, and consistent across teams
  • Analyze alert, incident, support, and SLO trends to replace reactive response with proactive risk reduction
  • Connect incident learnings with engineering standards, service maturity, product priorities, and vendor actions
  • Collaborate with application engineering teams to enhance developer experience and reduce operational toil
  • Work with service owners to establish production-readiness standards before releases
  • Advise on capacity planning, resilience testing, game days, disaster recovery, and modernization of fragile or legacy systems
  • Map critical customer workflows, define service health expectations, identify dependencies, and align reliability work with business priorities
  • Contribute to cross-team reliability initiatives and influence decisions without direct authority
  • Develop strategic vendor partnerships across observability, incident response, and cloud infrastructure
  • Apply AI-assisted and agentic workflows to alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge
  • Maintain qualified human oversight of actions that may affect production
  • Strengthen reliability-workflow context using service metadata, observability data, incident records, runbooks, architecture documentation, and corrective-action quality

Cerințe

  • Professional experience in a production-facing SRE role supporting a complex SaaS environment
  • Sound technical judgment across distributed systems, multi-cloud platforms, Kubernetes, networking, infrastructure technologies, GitOps, and CI/CD
  • Experience creating and maturing SRE practices involving SLOs, error budgets, observability, capacity planning, incident response, and toil reduction
  • Ability to use observability data to resolve severe incidents and conduct root-cause analysis
  • Experience leading critical incidents and communicating with engineers, executives, customer-facing teams, and external vendors
  • Bachelor’s or master’s degree in computer science or a related discipline
  • At least 6 years of software engineering experience
  • At least 4 years in DevOps roles focused on developing CI/CD pipelines
  • Advanced knowledge of Kubernetes, Docker, and orchestration platforms
  • Strong hands-on experience with AWS, Azure, or GCP
  • Proficiency with infrastructure-as-code tools such as Terraform, Chef, or Ansible
  • Practical experience with New Relic, Prometheus, Grafana, or comparable observability platforms
  • Proficiency in Python, Go, Bash, or similar programming and scripting languages
  • Knowledge of event-driven autoscaling and advanced Kubernetes configuration
  • Familiarity with Buildkite, Spinnaker, GitHub Actions, or comparable CI/CD platforms
  • Experience with microservices, containerization, and DevOps operating practices
  • Strong understanding of distributed systems, scalability, high availability, and performance optimization
  • Excellent problem-solving ability in a fast-paced work environment

Beneficii

  • An inclusive workplace culture that supports employee growth and belonging
  • Participation in a 12-hour follow-the-sun on-call rotation

Locuri de muncă similare

Kreato Global | BPO and Language Solutions

Remote English-Spanish OPI/VRI Interpreter

Kreato Global | BPO and Language Solutions

Interpret remotely between English- and Spanish-speaking people in medical, financial, social service, and customer care settings. Provide language support for Kreato Global across Latin America.

Deschide