Platform.sh
Platform.sh
Platform.sh provides a collaborative cloud application platform for full-stack web development. Its SaaS offering brings application code, frontend and backend services, APIs, databases, security, and deployment workflows together, helping development teams build, preview, launch, scale, and iterate on web applications without managing the underlying infrastructure. The platform also supports eCommerce projects and API-driven applications, with built-in observability and on-demand preview environments designed to make collaboration and delivery more efficient.

Senior Site Reliability Engineer Australia Remote

Senior Site Reliability Engineer responsible for advancing Upsun’s cloud application platform through stronger reliability, observability, and automation. Build resilient multi-cloud infrastructure as part of a globally distributed remote software team.

Description

  • Advance Upsun’s cloud application platform from traditional operations toward a proactive, automation-led SRE model
  • Lead key engineering initiatives that improve reliability, scalability, and operational efficiency across multi-cloud environments
  • Collaborate with engineering, product, and platform teams to integrate reliability and performance into the full software delivery lifecycle
  • Identify architectural bottlenecks and advance infrastructure-as-code practices
  • Define observability standards that support sustained stability and uptime
  • Design monitoring, alerting, and logging solutions using Prometheus, Grafana, and ELK Stack
  • Create actionable SLIs and SLOs connected to core business metrics
  • Build resilient, automated infrastructure and workflows with Terraform and Ansible across AWS, GCP, and Azure
  • Improve CI/CD pipeline architecture to support rapid, secure, zero-downtime releases
  • Coordinate high-priority incident response and facilitate blameless post-mortems
  • Introduce preventative controls that strengthen system resilience
  • Work with product and software engineering teams to bring SRE practices into product planning
  • Investigate performance constraints and assess technologies including eBPF and container orchestration
  • Work a four-week rotation combining engineering delivery with operational troubleshooting and innovation
  • Join the on-call rotation approximately one week every 4–5 weeks from 02:00–10:00 UTC, including a weekend shift

Requirements

  • At least five years of experience in Site Reliability Engineering, cloud operations, or DevOps
  • Demonstrated ownership of reliability for production platforms operating at scale
  • Advanced Go or Python skills for developing automation tools, custom controllers, or SRE platform components
  • Extensive practical knowledge of Linux internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting
  • Deep experience with AWS, GCP, Azure, or OpenStack
  • Experience developing custom tools around cloud provider SDKs
  • Hands-on experience with declarative infrastructure tools such as Terraform
  • Ability to anticipate operational risks, weigh architectural trade-offs, and lead infrastructure initiatives with limited direction
  • Excellent cross-functional communication and a history of building alignment within an inclusive engineering culture
  • Legal authorization to work in Western Australia; visa sponsorship is not available
  • Successful completion of a background check
  • Bonus: experience building orchestration, edge, storage, and operational tools
  • Bonus: experience with Docker and production Kubernetes cluster management or containerized deployment architectures
  • Bonus: familiarity with PaaS architectures or developer-facing cloud platforms

Benefits

  • Flexible paid time off
  • Company stock options
  • Budget for professional development
  • Budget for office equipment
  • Wellness budget
  • Annual team gatherings
  • Internet cost reimbursement
  • Inclusive parental leave
  • Program supporting remote-work travel
  • Flexible, open, and inclusive work environment
  • Accommodations available throughout the hiring process

Related Jobs