Verity Group
Verity Group
Verity Group is a digital transformation and innovation consulting firm helping companies turn technology into practical business outcomes. Its work spans artificial intelligence migration, hyperautomation, cybersecurity, and digital engineering, combining strategic guidance with hands-on delivery from initial planning through implementation. With a focus on depth and tailored solutions, Verity Group works with ambitious organizations to improve efficiency and advance their technology capabilities.

SRE Engineer - Remote

Improve reliability, observability, and resilience at Verity Group, a digital transformation and engineering consultancy. Automate cloud, Kubernetes, and incident-management operations across remote environments.

Description

  • Define and monitor SLIs, SLOs, SLAs, MTTR, and MTTD.
  • Build observability, monitoring, alerting, and APM capabilities.
  • Track latency, traffic, errors, saturation, availability, and system performance.
  • Prevent, detect, and resolve operational incidents.
  • Perform root cause analyses and establish measures to prevent recurring issues.
  • Assess operational risks, bottlenecks, and single points of failure.
  • Help design resilient, scalable, and highly available solutions.
  • Automate routine operations and minimize manual work.
  • Manage and enhance Kubernetes and Docker environments.
  • Contribute to capacity planning, business continuity, and disaster recovery strategies.
  • Take part in deployments and application stabilization efforts.
  • Partner with teams to embed reliability from the solution-design phase.
  • Develop and maintain dashboards, alerts, procedures, and operational documentation.
  • Foster reliability, observability, and continuous improvement across the organization.

Requirements

  • Professional experience as an SRE, Site Reliability Engineer, or in a comparable role.
  • Hands-on experience with cloud platforms including GCP, AWS, and/or Azure.
  • Working knowledge of Kubernetes and Docker.
  • Experience implementing observability, monitoring, alerting, and APM solutions.
  • Understanding of SRE metrics and practices, including SLI, SLO, SLA, MTTR, and MTTD.
  • Experience managing, investigating, and resolving incidents.
  • Knowledge of application and infrastructure troubleshooting techniques.
  • Experience administering Linux environments.
  • Knowledge of networking, security, performance, and high-availability principles.
  • Experience with automation and Infrastructure as Code.
  • Experience working with CI/CD pipelines.
  • Strong communication skills and the ability to collaborate across functions.
  • Analytical, proactive, collaborative, and prevention-focused approach.
  • Experience with GKE, EKS, or AKS.
  • Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or comparable tools.
  • Experience with the ELK Stack, Elasticsearch, and Kibana.
  • Knowledge of Terraform and Ansible.
  • Experience supporting critical environments and distributed systems.
  • Experience in financial institutions or regulated environments.
  • Experience with capacity management and cloud cost optimization.
  • Knowledge of disaster recovery and business continuity practices.
  • Experience defining and managing error budgets.
  • Cloud, Kubernetes, or SRE certifications.

Benefits

  • Meal allowance.
  • Food allowance.
  • Home office allowance.
  • Health insurance.
  • Dental insurance.
  • Life insurance.
  • Birthday day off.
  • Total Pass / Wellhub access.
  • Boon Saúde app access.
  • Discount partnerships.
  • Discounts at partner establishments and educational institutions.
  • Welcome kit.
  • Onboarding program.
  • Verity Learning.
  • Verity Break.
  • #VerityComVocê.
  • Access to professional development courses.
  • Great Place to Work-certified workplace.

Related Jobs