MCCi
MCCi
51 – 200 Employees
ConsultingGovernmentLogistics
MCCi is a GovTech and consulting company that helps public sector organizations modernize how they manage information and deliver services. Its work includes Laserfiche enterprise content management implementations, records request and permitting platforms, government websites, intelligent document processing, robotic process automation, electronic signatures, system integrations, managed cloud services, and professional services. MCCi serves local and state governments, law enforcement organizations, and education agencies, with a focus on reducing manual work and improving operational efficiency.

Cloud Operations Engineer at MCCi (Remote, United States)

MCCi is hiring a Cloud Operations Engineer to strengthen reliability, observability, and automation for its JustFOIA public-records SaaS platform. This remote role operates Azure infrastructure and coordinates incident response for government-facing services.

Description

  • Provide site reliability engineering for JustFOIA, MCCi’s SaaS platform for government public-records requests
  • Develop and manage monitoring, alerting, dashboards, logging, and operational reporting across application, database, and infrastructure layers
  • Examine logs, metrics, traces, and service-health data to detect developing and recurring problems
  • Advance the platform’s reliability, availability, performance, scalability, and maintainability
  • Join the on-call rotation and direct incident triage, diagnosis, escalation, and service recovery
  • Troubleshoot production faults spanning applications, IIS/.NET, SQL Server, networking, and cloud services
  • Facilitate blameless incident reviews, document root causes, and implement follow-up remediation
  • Manage and optimize Azure and Azure Government environments
  • Create automation that limits repetitive tasks, human error, and configuration drift
  • Maintain and expand Infrastructure-as-Code, configuration management, deployment, and operations tooling
  • Assist with CI/CD workflows and application release activities
  • Assess performance constraints, capacity exposure, and scalability challenges
  • Establish and monitor service-level indicators, objectives, error budgets, or comparable reliability metrics
  • Assist with backup verification, disaster recovery, failover exercises, capacity planning, and operational readiness assessments
  • Help address vulnerabilities, strengthen operational security controls, and support compliance work
  • Write and update runbooks, diagnostic instructions, recovery processes, and operational records
  • Work with development teams to strengthen application reliability, observability, and supportability
  • Contribute to release automation, infrastructure enhancements, and the continued development of DevOps practices
  • May provide Cloud Operations support for additional MCCi SaaS products as responsibilities expand

Requirements

  • At least six years of total information technology experience
  • At least three years of site reliability engineering experience with production SaaS workloads in a public cloud, including accountability for availability and performance
  • Direct production SRE experience supporting a SaaS application
  • Practical experience establishing and using SLIs, SLOs, and error budgets
  • Experience implementing observability through metrics, centralized logs, distributed tracing, dashboards, and application performance monitoring
  • Ability to design useful alerts and reduce unnecessary alert volume
  • Background leading incident response, root cause investigations, and blameless post-incident reviews
  • Record of delivering reliability improvements through completion
  • Ability to develop automation with PowerShell, Python, Bash, or a similar language
  • Experience with Infrastructure as Code and configuration management practices
  • Experience running workloads on Microsoft Azure or another leading public cloud
  • Experience supporting customer-facing, business-critical software
  • Strong diagnostic, communication, teamwork, and technical writing abilities
  • Preferred: production support for Windows Server and IIS/.NET applications
  • Preferred: SQL Server monitoring and performance diagnosis
  • Preferred: CI/CD workflow and deployment automation experience
  • Preferred: performance testing and capacity planning experience
  • Preferred: network traffic handling, load balancing, web application firewalls, and content delivery experience
  • Preferred: backup, replication, failover, disaster recovery, and recovery testing experience
  • Preferred: container and orchestration technology experience
  • Familiarity with Datadog, Azure Application Insights, Elastic, Terraform, Ansible, PowerShell, Azure DevOps, or Cloudflare is advantageous

Benefits

  • Remote work supported by virtual inclusion for distributed teammates
  • Accessible managers and leadership
  • A trust-oriented culture with limited bureaucracy
  • Casual, comfortable dress expectations
  • An inclusive, diverse, and respectful workplace
  • Opportunities for collaboration, relationship building, and recognition

Related Jobs

Knowtion Health

Remote Talent Acquisition Manager

Knowtion Health

Oversee recruiting systems, requisitions, analytics, and contingent workforce operations at Knowtion Health. Help support scalable hiring processes for a growing healthcare company.

Open