Qutwo
Qutwo
Qutwo is an AI lab and software company working at the intersection of artificial intelligence, quantum computing, and enterprise optimization. Its Qutwo OS platform helps teams develop quantum and hybrid AI algorithms, connect high-performance computing, GPUs, and different quantum processing unit modalities, and operate production workloads with data pipelines, monitoring, and resilience. Qutwo also supports customers through research partnerships, forward-deployed scientists, and integration services designed to turn advanced AI and quantum-computing experiments into practical enterprise applications. The international team includes AI and quantum scientists and engineers, with hiring across AI, MLOps and DevOps, quantum algorithms, and platform engineering.

Senior Site Reliability Engineer in Helsinki (Hybrid)

Lead reliable, secure multi-cloud Kubernetes infrastructure for Qutwo’s quantum-AI platform. Operate distributed machine-learning workloads while advancing observability, GitOps, and CI/CD.

Description

  • Design and run Kubernetes infrastructure across multiple cloud platforms
  • Develop metrics, logging, tracing, and alerting capabilities
  • Establish practical service-level objectives and incident-response processes
  • Automate reproducible, reviewable environments with infrastructure as code and GitOps
  • Manage scheduling, autoscaling, and cost efficiency for distributed training and inference, including Ray clusters on Kubernetes
  • Protect platforms through secrets management, network policies, workload identity, and supply-chain security
  • Create CI/CD and deployment platforms for engineers and coding agents
  • Maintain documentation for infrastructure and architectural decisions

Requirements

  • Seven or more years of experience building and operating production systems, including substantial work in SRE, platform, or infrastructure engineering
  • Advanced practical knowledge of Kubernetes, including cluster operations, networking, storage, and troubleshooting
  • Experience operating platforms across multiple cloud providers and evaluating their trade-offs
  • Strong infrastructure-as-code and automation capabilities with tools such as Terraform, Helm, and GitOps
  • Proficiency in Python or Go for developing engineering tools
  • Sound understanding of distributed systems, failure scenarios, and reliability trade-offs
  • Security-focused approach to trust boundaries, identity, secrets, and software supply chains
  • Ability to use AI-assisted development tools and critically review coding-agent output, including infrastructure changes
  • Fluent English communication skills
  • Demonstrated effectiveness in agile, cross-functional teams

Benefits

  • Hybrid working arrangement in Helsinki
  • Chance to establish reliability practices and infrastructure foundations
  • Exposure to distributed machine-learning workloads, quantum-AI systems, and multi-cloud environments
  • Collaborative, agile, cross-functional working culture

Related Jobs

Walmart

Principal Software Engineer, Cloud and Distributed Systems

Walmart

Lead Walmart’s cloud-native, AI-enabled mGPS/Compass platform, shaping distributed services and indoor routing. Improve in-store operations through resilient architecture and technical leadership.

Open