Cisco
Cisco
Cisco एक एंटरप्राइज़ प्रौद्योगिकी कंपनी है, जिसका मुख्य ध्यान नेटवर्किंग, सुरक्षा और डिजिटल अवसंरचना पर है। इसके उत्पादों और सेवाओं में राउटर, स्विच, ऑप्टिकल ट्रांसीवर, प्रोग्रामेबल सिलिकॉन, एज कंप्यूटिंग प्लेटफ़ॉर्म, Webex सहयोग उपकरण, ऑब्ज़र्वेबिलिटी उत्पाद, कृत्रिम बुद्धिमत्ता-सक्षम सॉफ़्टवेयर और क्लाउड-प्रबंधित सेवाएँ शामिल हैं। Cisco उद्यमों, सेवा प्रदाताओं और सरकारों को सेवा देता है तथा प्रौद्योगिकी, पेशेवर सेवाओं, प्रशिक्षण और निरंतर सहायता के माध्यम से बड़े पैमाने के नेटवर्क और डेटा केंद्रों के डिज़ाइन, संचालन और सुरक्षा में सहयोग करता है।

Senior Site Reliability Engineer, Observability at Cisco (Hybrid)

Lead observability engineering for Cisco’s cloud infrastructure, covering Splunk, Elasticsearch, distributed tracing, telemetry, and production reliability. Build and operate systems that improve service visibility, incident response, and platform resilience.

विवरण

  • Design, deploy, operate, and improve enterprise observability platforms
  • Administer Splunk Enterprise and Splunk Cloud infrastructure, including indexers, Search Head Clusters, heavy forwarders, deployment servers, and integrations
  • Run large-scale Elasticsearch clusters supporting log analytics, search, and operational troubleshooting
  • Design and support distributed tracing platforms built with Grafana Tempo and OpenTelemetry
  • Develop and maintain complete telemetry pipelines for logs, metrics, and traces
  • Set instrumentation standards, data-quality practices, retention policies, and observability patterns
  • Scale and optimize Prometheus, Grafana, Kafka, Tempo, OpenTelemetry, and related monitoring platforms
  • Create dashboards, alerts, analytics, and trace visualizations with Splunk SPL, Grafana, Kibana, and Tempo
  • Define monitoring and alerting standards that improve service visibility and reduce incident detection and resolution times
  • Automate infrastructure provisioning, configuration, upgrades, and operational workflows with Terraform and configuration-management tools
  • Contribute to capacity planning, performance tuning, upgrades, patching, disaster recovery, and production-readiness reviews
  • Diagnose complex distributed-systems problems and lead or support incident response and root-cause analysis
  • Work with platform, application, security, database, and network engineering teams to improve reliability and operational consistency
  • Help maintain shared SRE standards, documentation, runbooks, and engineering practices
  • Take part in the shared SRE pager and production on-call rotation
  • Support customer-facing production services and Kubernetes platform workloads such as Nextunnel
  • Improve service visibility, accelerate incident detection and resolution, strengthen platform reliability, and advance operational excellence across Cisco’s cloud environment

आवश्यकताएँ

  • At least 7 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or a related discipline
  • Production experience administering Splunk Enterprise or Splunk Cloud
  • Strong command of Splunk SPL, including production dashboards, alerts, and analytics
  • Experience designing, deploying, and operating Elasticsearch or ELK platforms
  • Hands-on experience with Prometheus, Grafana, Grafana Tempo, OpenTelemetry, distributed tracing, and Kafka
  • Proven ability to implement and operate observability strategies spanning logs, metrics, and traces
  • Experience using Terraform and Infrastructure as Code practices
  • Strong knowledge of Linux, networking, distributed systems, high availability, and production operations
  • Programming or scripting experience with Python, Go, Ruby, Bash, or a comparable language
  • Experience resolving production incidents within an on-call or operational support model
  • Strong communication, collaboration, and problem-solving abilities
  • Must be a U.S. Person, defined as a U.S. citizen or U.S. national
  • Work must be performed within the United States
  • Splunk certification preferred
  • Experience with Kubernetes and containerized production workloads preferred
  • Experience operating AWS, Azure, or GCP cloud infrastructure preferred
  • Experience with Ansible, Consul, CI/CD pipelines, or service mesh technologies preferred
  • Experience with SLOs, SLIs, error budgets, capacity planning, and incident management preferred
  • Experience supporting FedRAMP, government, or other regulated environments preferred
  • Experience building secure, highly available, and compliant observability platforms preferred
  • Experience leading cross-functional technical initiatives from proposal through implementation preferred

लाभ

  • Medical, dental, and vision insurance
  • 401(k) plan with a Cisco matching contribution
  • Paid parental leave
  • Short- and long-term disability coverage
  • Basic life insurance
  • Cisco restricted stock unit grants may be available
  • 10 paid holidays per full calendar year
  • One floating holiday for non-exempt employees
  • One paid day off for the employee’s birthday
  • Paid year-end holiday shutdown
  • Four paid days off for personal wellness
  • 16 paid vacation days per full calendar year for non-exempt employees
  • Flexible vacation time off with no defined limit for eligible exempt employees
  • 80 hours of sick time provided on the hire date and each January 1 thereafter
  • Up to 80 hours of unused sick time may carry forward annually
  • Additional paid time away for critical or emergency family matters
  • Optional 10 paid volunteer days per full calendar year
  • Annual bonuses for non-sales roles, subject to Cisco policies

संबंधित नौकरियाँ

Maxibook OÜ

Scientific Language and AI Adaptation Contributor

Maxibook OÜ

Help Maxibook, an AI education startup, adapt English scientific content for Korean, Japanese, or Arabic audiences. Assess AI-generated material for accurate concepts, natural language, and academic credibility.

खोलें
ZACO Robot

B2B Sales Setter for Commercial Cleaning Robotics

ZACO Robot

Qualify leads and book sales calls for RV-Tech’s professional cleaning robots, selling remotely to business decision-makers. The role focuses on automation and robotics in a growing commercial market.

खोलें
Administration Intelligence AG

Product Manager, Procurement and Tendering Software

Administration Intelligence AG
51 – 200 कर्मचारी
उद्यमपरामर्शसरकार

Shape the roadmap for AI-enabled procurement and tendering software. Consolidate products and develop digital tools for public-sector buyers.

खोलें
Prospeum
MMG Management Consulting

Strategic Sourcing Consultant

MMG Management Consulting

Drive supplier selection, negotiations, and sourcing initiatives for clients at a boutique management consultancy. Work with stakeholders to advance sourcing strategies, processes, and tools.

खोलें
SUXXEED Sales for your Success GmbH

Working Student AI Engineer

SUXXEED Sales for your Success GmbH
201 – 500 कर्मचारी
B2Bपरामर्श

Join SUXXEED’s AI team to build practical AI applications and automation for a B2B sales specialist. Work with colleagues across AI, IT, and business teams in a hybrid role.

खोलें
Peraton
TD Trusted Decisions GmbH

Senior CPM Consultant

TD Trusted Decisions GmbH
11 – 50 कर्मचारी
B2Bपरामर्श

Join a management consultancy delivering business intelligence solutions. Advise clients, implement systems, and provide training in integrated financial planning, consolidation, and reporting.

खोलें
SUXXEED Sales for your Success GmbH

Working Student, Business Administration (AI)

SUXXEED Sales for your Success GmbH
201 – 500 कर्मचारी
B2Bपरामर्श

Support SUXXEED’s B2B sales teams with AI business cases, KPI reporting, and process improvements. Work with AI engineers and business teams in Cologne as part of the AI Center of Excellence.

खोलें
GLORY

Remote Creative Marketing Specialist

GLORY

Create visual assets and campaigns for GLORY, a California-rooted activewear brand. Support digital content, product launches, brand consistency, and creative workflows.

खोलें