Cloudflare
Cloudflare
Cloudflare provides a connectivity cloud that helps organizations secure and connect employees, devices, applications, networks, and data across on-premise environments, public clouds, SaaS platforms, and the internet. Its platform brings together application security, network protection, performance optimization, and visibility tools, helping teams reduce technical complexity while building and delivering software more efficiently. Cloudflare’s work spans cloud and edge delivery, secure networking, and resilient digital infrastructure, making it relevant to professionals working across cybersecurity, networking, cloud platforms, and application performance.

Senior Forward Deployed DevOps Engineer - Cloudflare

Lead the deployment of Cloudflare compute capacity for strategic customers, from infrastructure planning and hardware validation through production handover. Automate fleet operations, establish reliability standards, and coordinate customer and infrastructure teams.

Description

  • Bring new compute capacity online for strategic customers through planning, provisioning, validation, and production handover
  • Define acceptance criteria and validation workflows for burn-in, benchmarking, and health checks across compute, storage, and networking
  • Partner with on-site data center teams to perform checks and resolve operational issues
  • Enforce build, configuration, security, and observability standards before capacity launches
  • Create monitoring, alerting, and runbooks for production systems
  • Coordinate Cloudflare infrastructure, hardware engineering, data center operations, supply chain, network, and customer engineering teams
  • Act as the primary technical contact for customer infrastructure while shaping account strategy and technical direction
  • Take part in standups, capacity reviews, change management, and incident response
  • Lead blameless root-cause analysis and implement post-incident improvements
  • Build production-grade automation and tooling for provisioning, validation, and fleet management
  • Forecast customer capacity requirements and identify constraints proactively
  • Escalate operational issues and platform opportunities to Infrastructure, Hardware, and Product teams
  • Develop trusted technical relationships with Staff+ engineers, Directors, and VPs across both organizations
  • Work regularly on-site at customer offices

Requirements

  • Demonstrated infrastructure, SRE, or production engineering experience running and scaling large production systems
  • Hands-on experience bringing server capacity online, including provisioning, imaging, configuration management, and fleet-scale lifecycle operations
  • Working knowledge of server hardware, firmware, BMC/IPMI/Redfish, and diagnostic processes
  • Ability to establish validation standards, troubleshoot issues remotely, and collaborate with on-site technicians
  • Strong knowledge of Linux and core networking, including TCP/IP, BGP, DNS, load balancing, and data center topologies
  • Proficiency in Go, Python, or Rust
  • Experience with infrastructure-as-code and configuration management tools such as Terraform, Ansible, or Salt
  • Background supporting mission-critical services through on-call rotations, incident response, SLO management, and reliability engineering
  • Familiarity with monitoring, metrics, and logging tools such as Prometheus, Grafana, or ClickHouse
  • Exposure to Kubernetes or comparable workload-scheduling platforms
  • Interest or experience using modern AI tools for automation, debugging, and documentation
  • Ability to coordinate complex delivery across multiple teams, regions, and time zones
  • Ability to conduct strategic technical discussions with executive stakeholders
  • Experience at a hyperscaler, cloud provider, or large-scale data center is a plus
  • Familiarity with Cloudflare’s platform and network architecture is a plus
  • Experience in customer-embedded or partner-facing technical infrastructure roles is a plus
  • Technical writing, internal mentorship, open-source work, or community presentations are additional advantages
  • Applicants may require authorization to receive software or technology controlled under U.S. export laws without sponsorship for an export license

Benefits

  • Participation in an equity plan
  • Medical and prescription drug insurance
  • Dental insurance
  • Vision insurance
  • Flexible spending accounts
  • Commuter spending accounts
  • Fertility and family-forming benefits
  • On-demand mental health support and an Employee Assistance Program
  • Global travel medical insurance
  • Short- and long-term disability insurance
  • Life and accident insurance
  • 401(k) retirement savings plan
  • Employee stock participation plan
  • Flexible paid time off for vacation and sick leave
  • Leave programs covering parental, pregnancy health, medical, and bereavement leave

Related Jobs

Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Employees
Artificial IntelligenceB2B

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Open
General Motors

Summer AI Research Intern, Embodied AI

General Motors

Research machine learning, robotics, and multimodal models for autonomous driving at General Motors. Run experiments and collaborate with research and engineering teams in a hybrid internship based in Sunnyvale, California.

Open
Mobileye

Senior Deep Learning Researcher, Autonomous Vehicles

Mobileye

Research and develop deep learning systems for Mobileye’s autonomous vehicles, including large-scale neural networks built for its EyeQ chip. Improve model performance and bring research into production.

Open