F5
F5
F5 is a cybersecurity and enterprise SaaS company focused on application services and security. Its technology helps organizations manage digital applications across cloud environments, with capabilities designed to support application performance, availability, control, and protection.

Site Reliability Engineer II at F5 (Hybrid, San Jose)

Operate F5’s FedRAMP-compliant AWS, Kubernetes, and bare-metal infrastructure. Support continuous monitoring, incident response, security operations, and GitOps-based delivery.

Description

  • Monitor bare-metal RHEL servers, EKS clusters, and AWS environments continuously.
  • Triage critical Prometheus and Alertmanager alerts, respond to incidents, and investigate root causes.
  • Analyze and troubleshoot logs with Elasticsearch and Kibana.
  • Lead incident-resolution bridges in Slack and coordinate with on-call, security, and infrastructure teams.
  • Record incident timelines, root-cause analyses, post-incident reports, and post-mortems.
  • Manage configuration, maintenance, lifecycle administration, and hardware operations for bare-metal RHEL servers.
  • Diagnose network interface bonding, link aggregation, VLAN tagging, and connectivity problems.
  • Coordinate vendor hardware dispatches and replacement of physical components.
  • Carry out FedRAMP security operations, including vulnerability patching, system hardening, FIPS configuration, RBAC, SSH key management, and security-boundary enforcement.
  • Operate and maintain Amazon EKS clusters in AWS Commercial and GovCloud environments.
  • Troubleshoot ALB, NLB, ingress controllers, SSL/TLS termination, and traffic routing.
  • Administer Linux systems, including kernel tuning, storage expansion, log rotation, and general troubleshooting.
  • Implement Terraform infrastructure changes through change-control procedures.
  • Monitor AWS VPCs, Security Groups, EC2, IAM, S3, and KMS.
  • Run, monitor, and troubleshoot GitLab CI/CD pipelines, ArgoCD synchronization and rollouts, and Argo Workflows.
  • Coordinate releases and configuration rollouts with Helm, Kustomize, and GitOps workflows.
  • Verify build compliance, container security scans, and image signatures before deployment.
  • Keep operational runbooks, standard operating procedures, and triage workflows up to date.
  • Reduce operational toil by developing Bash and Python scripts for monitoring, hardware checks, and maintenance.

Requirements

  • U.S. citizenship is required because of FedRAMP and AWS GovCloud security requirements.
  • Three to five years of hands-on experience in SRE, DevOps, systems administration, or hybrid infrastructure support.
  • Availability and ability to work a rotating 24/7 schedule, including nights, weekends, and holidays.
  • Strong hands-on administration experience with bare-metal RHEL servers, including hardware diagnostics, out-of-band management through IPMI, iDRAC, or iLO, RAID, LVM, and physical NIC bonding.
  • Direct experience querying and analyzing Kibana and Elasticsearch log streams using KQL or Lucene, index management, and log-pattern matching.
  • Practical experience troubleshooting Layer 4 and Layer 7 secure load balancing, mTLS, TLS termination, health checks, and secure traffic routing.
  • Solid operational experience with AWS core services, including EC2, VPC, IAM, S3, and KMS, plus familiarity with AWS GovCloud operating models.
  • Hands-on experience operating Amazon EKS and Kubernetes, including kubectl, Helm, pod lifecycle management, and node-pool maintenance.
  • Experience with Prometheus, PromQL, Alertmanager, and Grafana.
  • Hands-on experience deploying and troubleshooting with GitLab CI/CD, ArgoCD, and Argo Workflows.
  • Experience using Slack for operational messaging, ChatOps commands, and incident war rooms.
  • Working knowledge of FedRAMP and NIST SP 800-53 controls, vulnerability patching cycles, and DISA STIG hardening.
  • Python or Bash automation scripting experience is preferred.
  • Basic working knowledge of Terraform is preferred.
  • RHCSA or RHCE certification is preferred.
  • AWS Certified SysOps Administrator or CKA certification is preferred.

Benefits

  • Incentive compensation
  • Bonus
  • Restricted stock units
  • Benefits package
  • Reasonable accommodations for candidates

Related Jobs

CONVACT - Agentur für Suchmaschinenoptimierung

Remote Sales Representative – SEO and Google Ads

CONVACT - Agentur für Suchmaschinenoptimierung
1 – 10 Employees
B2BMarketing

Follow up with leads and cold-call prospective clients for Convact’s SEO and Google Ads agency. Book appointments with online shops and e-commerce businesses across the DACH region.

Open