GrooveTech
GrooveTech
51 – 200 Employees
B2BEnterpriseRecruitment
GrooveTech is a Brazilian IT services firm that helps organizations extend and strengthen their technology capabilities through managed staff augmentation, dedicated delivery squads, software quality assurance, 24/7 NOC monitoring, and strategic IT consulting. Its services include digital due diligence for mergers and acquisitions, technology leadership and delivery roles, technical and business partnerships, and tailored engagement models. GrooveTech combines daily reporting with test automation and on-premises or cloud monitoring to support reliable delivery, reduce turnover, and improve project performance.

Mid-Level Cloud Engineer / Site Reliability Engineer – São Paulo Hybrid

Support GrooveTech’s critical platform on GCP and Kubernetes, with a focus on reliability, observability, automation, incident response, and cloud cost optimization.

Description

  • Manage and improve GrooveTech’s GCP environment across GKE, Compute Engine, Cloud SQL, Storage, VPC, Load Balancers, Pub/Sub, and BigQuery
  • Use SLI, SLO, and Error Budget practices to support high availability and configure autoscaling strategies
  • Strengthen observability with New Relic, OpenTelemetry, Prometheus, Grafana, and Cloud Monitoring
  • Develop dashboards and alerts that provide clear, actionable operational insight
  • Provision and manage infrastructure as code with Terraform
  • Support deployment workflows built with GitHub Actions and ArgoCD
  • Automate recurring operational tasks using Go, Python, and/or Bash
  • Investigate production issues across Linux systems, networks, and Kubernetes
  • Facilitate root-cause analysis and blameless postmortems
  • Enforce governance controls such as IAM/RBAC and Secret Manager policies
  • Deliver initiatives that reduce and optimize cloud costs
  • Help improve the reliability, resilience, and scalability of a critical platform within a strategic project

Requirements

  • At least three years of experience managing production cloud infrastructure
  • Hands-on production experience with Google Cloud Platform and Kubernetes, including GKE
  • Advanced Terraform skills, including module development, maintenance, and state management
  • Experience automating CI/CD with GitHub Actions and implementing GitOps workflows with ArgoCD
  • Strong scripting ability in Go, Python, and/or Bash
  • Working knowledge of TCP/IP, DNS, HTTP/HTTPS, TLS, Linux, and load balancers for network troubleshooting
  • Experience with FinOps, disaster recovery, or service mesh technologies such as Istio or Linkerd is a technical advantage
  • Experience creating developer portals or self-service pipelines in a platform engineering context is a technical advantage
  • GCP certifications such as Professional Cloud Architect or Professional Cloud DevOps Engineer, or Kubernetes certifications such as CKA or CKAD, are a technical advantage
  • Demonstrated resilience and adaptability
  • Clear communication and effective collaboration skills
  • Systems thinking with a results-oriented approach
  • Strong ownership and initiative

Benefits

  • Independent Contractor (PJ) engagement
  • Access to the Wellhub (Gympass) package
  • Corporate life insurance
  • Remote work arrangement with occasional in-person attendance in São Paulo

Related Jobs

SupportNinja

Technical Support Representative II – IoT and Networking | Italian, French, English | Romania Remote

SupportNinja
1,001 – 5,000 Employees
B2BSaaS

Provide remote technical support for IoT devices, networking, and cloud connectivity to SupportNinja’s global customers. Assist field engineers and resolve tickets within SLA commitments.

Open