Entarian
Entarian
1,001 – 5,000 Employees
AerospaceDefenseGovernment
Entarian is an engineering and federal-technology company serving aerospace, defense, and civilian government missions. Formerly known as ERT, it combines science, engineering, and mission-support expertise across satellite operations and ground systems, space and Earth science data, digital engineering, data analytics, AI/ML, network integration, command and control, and cyber defense. Its work with U.S. federal agencies and defense partners includes satellite calibration and validation, search-and-rescue operations, weather forecasting support, and IT modernization. Entarian rebranded following its acquisition of Sev1Tech and is backed by Macquarie Capital.

Senior Site Reliability Engineer (Windows) - Remote in Virginia

Senior Site Reliability Engineer responsible for dependable Windows production infrastructure supporting Entarian’s mission-critical engineering solutions. The role focuses on automation, service monitoring, incident response, and deployment reliability.

Description

  • Design and maintain scalable infrastructure through Infrastructure as Code practices
  • Create and support PowerShell, Python, Ruby, and other automation scripts for provisioning, configuration management, and operations
  • Deploy and administer Terraform, Puppet, and/or Chef across hybrid environments
  • Track system health, performance, and availability using established monitoring practices
  • Define and uphold SLAs, SLOs, and error budgets for production services
  • Join the on-call rotation, restore services quickly, and perform root cause analysis
  • Partner with development teams to strengthen deployment pipelines and release workflows
  • Maintain operational procedures, runbooks, and architecture decision records
  • Lead post-mortem reviews and apply corrective measures to prevent repeat incidents
  • Safeguard the reliability, availability, and performance of Windows production environments
  • Connect development and operations teams to provide highly available services and operational excellence

Requirements

  • At least 5 years of experience in systems administration, DevOps, or site reliability engineering
  • Advanced experience with Windows Server 2016 or later, including Active Directory, IIS, and Microsoft SQL Server
  • Strong cloud experience, preferably with AWS
  • Advanced scripting ability, including module development and REST API integration
  • Practical Terraform experience for infrastructure provisioning
  • Practical Puppet or Chef experience for configuration management
  • Experience with observability platforms such as Prometheus, Grafana, Datadog, or New Relic
  • Strong networking knowledge covering DNS, TCP/IP, load balancing, and VPNs
  • Demonstrated ability to diagnose complex problems across multiple technology layers
  • Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent professional experience
  • Certifications such as AWS Solutions Architect, Microsoft certifications, or HashiCorp Certified: Terraform Associate
  • Experience working with Docker and Kubernetes
  • Familiarity with GitLab Pipelines, Jenkins, or GitHub Actions
  • Knowledge of security best practices and compliance frameworks
  • Experience with the ELK Stack or Splunk
  • Ability to obtain the required security clearance

Benefits

  • Medical, dental, and vision coverage
  • Life, AD&D, and disability insurance
  • Paid time off
  • Eleven company holidays
  • 401(k) plan with company matching
  • Additional employee benefits and wellness resources

Related Jobs

Knowtion Health

Remote Talent Acquisition Manager

Knowtion Health

Oversee recruiting systems, requisitions, analytics, and contingent workforce operations at Knowtion Health. Help support scalable hiring processes for a growing healthcare company.

Open