Cisco
Cisco
Cisco is an enterprise technology company focused on networking, security, and digital infrastructure. Its portfolio includes routers, switches, optical transceivers, programmable silicon, edge computing platforms, Webex collaboration tools, observability products, AI-enabled software, and cloud-managed services. Cisco serves enterprises, service providers, and governments, supporting the design, operation, and security of large-scale networks and data centers through technology, professional services, training, and ongoing support.

Senior Site Reliability Engineering Technical Leader at Cisco — Remote, United States

Lead the migration of specialized Kubernetes workloads from AWS to Cisco’s internal platforms. Improve the reliability, networking, scalability, and operations of globally distributed services.

Description

  • Assess migration feasibility, develop technical plans, and execute phased transfers of eligible workloads from AWS-hosted Kubernetes clusters to Cisco’s internal Kubernetes platform.
  • Set technical direction and define delivery plans for each stage of the work.
  • Operate and improve specialized Kubernetes workloads with complex load-balancing, networking, and traffic-management needs.
  • Support migrated applications through production readiness, incident response, performance tuning, and ongoing operational improvements.
  • Diagnose and resolve complex problems spanning applications, Kubernetes, Linux, networking, containers, and infrastructure.
  • Strengthen the scalability, reliability, security, performance, and operability of Kubernetes-based services.
  • Collaborate with Node Connectivity, firmware, cloud infrastructure, security, SRE, and product teams.
  • Set support priorities and coordinate changes across interconnected systems.
  • Advance wider Kubernetes SRE and platform-reliability programs.
  • Build and maintain production-grade Go software for distributed, concurrent, and networked systems.
  • Improve operational readiness through SLIs and SLOs, comprehensive testing, on-call support, root-cause analysis, and durable corrective actions.

Requirements

  • At least 10 years of professional experience in software, site reliability, or infrastructure engineering, including technical leadership for substantial production systems.
  • Experience designing, deploying, and operating large-scale distributed services on Kubernetes.
  • At least 5 years of programming experience with Go or a comparable systems programming language.
  • Experience supporting production services through incident response, performance analysis, Kubernetes reliability practices, observability, automation, and operational-readiness work.
  • Working knowledge of Linux and networking, including IPv4/IPv6, TCP, routing, DNS, and TLS.
  • Strong judgment in architecture, incident response, prioritization, and technical tradeoffs.
  • Ability to build alignment across teams without depending on formal authority.
  • Preferred: Experience operating Kubernetes in AWS, on-premises, or hybrid environments with AWS/EKS, Docker, Kustomize, GitLab CI/CD, or comparable deployment tools.
  • Preferred: Experience with Kubernetes networking, ingress, load balancing, service discovery, and traffic management for high-throughput or geographically distributed services.
  • Preferred: Experience with gRPC, Protocol Buffers, mutual TLS, PKI, VPNs, tunneling, or network security.
  • Preferred: Experience with OpenTelemetry, Prometheus, Datadog, or comparable observability tools, as well as distributed routing, packet processing, performance optimization, or failure testing.

Benefits

  • Medical, dental, and vision insurance.
  • 401(k) plan with Cisco matching contributions.
  • Paid parental leave.
  • Short- and long-term disability coverage.
  • Basic life insurance.
  • Cisco restricted stock unit grants may be available based on eligibility and continued employment.
  • 10 paid holidays per full calendar year.
  • One floating holiday for non-exempt employees.
  • Paid birthday day off.
  • Paid year-end holiday shutdown.
  • Four paid personal wellness days.
  • 16 paid vacation days per full calendar year for non-exempt employees.
  • Flexible vacation time with no defined limit for eligible exempt employees.
  • 80 hours of sick time on the hire date and each January 1 thereafter.
  • Carry forward of up to 80 unused sick hours.
  • Additional paid leave for critical or emergency family matters.
  • Optional 10 paid volunteer days per full calendar year.
  • Annual bonuses for non-sales roles, subject to Cisco policies.
  • Incentive compensation for sales-plan employees, subject to the applicable Cisco plan.

Related Jobs

CONVACT - Agentur für Suchmaschinenoptimierung

Remote Sales Representative – SEO and Google Ads

CONVACT - Agentur für Suchmaschinenoptimierung
1 – 10 Employees
B2BMarketing

Follow up with leads and cold-call prospective clients for Convact’s SEO and Google Ads agency. Book appointments with online shops and e-commerce businesses across the DACH region.

Open