Cisco
Cisco
Cisco tarmoqlar, xavfsizlik va raqamli infratuzilmaga ixtisoslashgan korporativ texnologiyalar kompaniyasidir. Uning mahsulotlar portfeliga marshrutizatorlar, kommutatorlar, optik transiverlar, dasturlashtiriladigan kremniy yechimlari, chekka hisoblash platformalari, Webex hamkorlik vositalari, kuzatuvchanlik mahsulotlari, sun’iy intellekt asosidagi dasturiy ta’minot va bulut orqali boshqariladigan xizmatlar kiradi. Cisco texnologiyalar, professional xizmatlar, treninglar va doimiy yordam ko‘rsatish orqali korxonalar, xizmat ko‘rsatuvchi provayderlar hamda davlat tashkilotlariga yirik tarmoqlar va ma’lumotlar markazlarini loyihalash, boshqarish va himoyalashda ko‘maklashadi.

Senior Site Reliability Engineer, Observability at Cisco (Hybrid)

Lead observability engineering for Cisco’s cloud infrastructure, covering Splunk, Elasticsearch, distributed tracing, telemetry, and production reliability. Build and operate systems that improve service visibility, incident response, and platform resilience.

Tavsif

  • Design, deploy, operate, and improve enterprise observability platforms
  • Administer Splunk Enterprise and Splunk Cloud infrastructure, including indexers, Search Head Clusters, heavy forwarders, deployment servers, and integrations
  • Run large-scale Elasticsearch clusters supporting log analytics, search, and operational troubleshooting
  • Design and support distributed tracing platforms built with Grafana Tempo and OpenTelemetry
  • Develop and maintain complete telemetry pipelines for logs, metrics, and traces
  • Set instrumentation standards, data-quality practices, retention policies, and observability patterns
  • Scale and optimize Prometheus, Grafana, Kafka, Tempo, OpenTelemetry, and related monitoring platforms
  • Create dashboards, alerts, analytics, and trace visualizations with Splunk SPL, Grafana, Kibana, and Tempo
  • Define monitoring and alerting standards that improve service visibility and reduce incident detection and resolution times
  • Automate infrastructure provisioning, configuration, upgrades, and operational workflows with Terraform and configuration-management tools
  • Contribute to capacity planning, performance tuning, upgrades, patching, disaster recovery, and production-readiness reviews
  • Diagnose complex distributed-systems problems and lead or support incident response and root-cause analysis
  • Work with platform, application, security, database, and network engineering teams to improve reliability and operational consistency
  • Help maintain shared SRE standards, documentation, runbooks, and engineering practices
  • Take part in the shared SRE pager and production on-call rotation
  • Support customer-facing production services and Kubernetes platform workloads such as Nextunnel
  • Improve service visibility, accelerate incident detection and resolution, strengthen platform reliability, and advance operational excellence across Cisco’s cloud environment

Talablar

  • At least 7 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or a related discipline
  • Production experience administering Splunk Enterprise or Splunk Cloud
  • Strong command of Splunk SPL, including production dashboards, alerts, and analytics
  • Experience designing, deploying, and operating Elasticsearch or ELK platforms
  • Hands-on experience with Prometheus, Grafana, Grafana Tempo, OpenTelemetry, distributed tracing, and Kafka
  • Proven ability to implement and operate observability strategies spanning logs, metrics, and traces
  • Experience using Terraform and Infrastructure as Code practices
  • Strong knowledge of Linux, networking, distributed systems, high availability, and production operations
  • Programming or scripting experience with Python, Go, Ruby, Bash, or a comparable language
  • Experience resolving production incidents within an on-call or operational support model
  • Strong communication, collaboration, and problem-solving abilities
  • Must be a U.S. Person, defined as a U.S. citizen or U.S. national
  • Work must be performed within the United States
  • Splunk certification preferred
  • Experience with Kubernetes and containerized production workloads preferred
  • Experience operating AWS, Azure, or GCP cloud infrastructure preferred
  • Experience with Ansible, Consul, CI/CD pipelines, or service mesh technologies preferred
  • Experience with SLOs, SLIs, error budgets, capacity planning, and incident management preferred
  • Experience supporting FedRAMP, government, or other regulated environments preferred
  • Experience building secure, highly available, and compliant observability platforms preferred
  • Experience leading cross-functional technical initiatives from proposal through implementation preferred

Imtiyozlar

  • Medical, dental, and vision insurance
  • 401(k) plan with a Cisco matching contribution
  • Paid parental leave
  • Short- and long-term disability coverage
  • Basic life insurance
  • Cisco restricted stock unit grants may be available
  • 10 paid holidays per full calendar year
  • One floating holiday for non-exempt employees
  • One paid day off for the employee’s birthday
  • Paid year-end holiday shutdown
  • Four paid days off for personal wellness
  • 16 paid vacation days per full calendar year for non-exempt employees
  • Flexible vacation time off with no defined limit for eligible exempt employees
  • 80 hours of sick time provided on the hire date and each January 1 thereafter
  • Up to 80 hours of unused sick time may carry forward annually
  • Additional paid time away for critical or emergency family matters
  • Optional 10 paid volunteer days per full calendar year
  • Annual bonuses for non-sales roles, subject to Cisco policies

O‘xshash ish o‘rinlari