Mozn
Mozn
201 – 500 Employees
ConsultingHealthcareLogistics
Mozn is an AI company focused on Arabic-native generative AI and enterprise intelligence for organizations across the MENA region. Its products include OSOS, an Arabic-first generative AI platform, and FOCAL, a platform for financial crime and fraud detection. Mozn also develops tailored solutions in language intelligence, risk intelligence, operational AI, data management, geospatial intelligence, and AI centers, serving sectors such as healthcare, finance, and government. Its work emphasizes culturally relevant AI, data security, and regulatory compliance.

Distributed Systems Engineer III, Remote in Egypt

Build and operate scalable cloud, Kubernetes, Kafka, and data platforms for Mozn’s Enterprise AI products. Support resilient, data-intensive, and AI workloads in production environments.

Description

  • Build and continuously improve production cloud-native platforms for distributed workloads
  • Operate Kubernetes platforms, including upgrades, node pools, workload lifecycles, troubleshooting, and routine operations
  • Deploy and manage services with ArgoCD, Helm, GitOps, Terraform, and automation
  • Design systems that prioritize scalability, availability, resilience, performance, and operational simplicity
  • Resolve complex issues spanning Kubernetes, cloud infrastructure, networking, storage, applications, and distributed services
  • Run and troubleshoot Apache Kafka in high-throughput production environments
  • Manage Kafka topics, partitions, replication, consumer groups, retention, throughput, latency, and recovery from failures
  • Connect Kafka with databases and applications through Kafka Connect, Debezium, or comparable CDC and event-streaming tools
  • Operate MySQL and/or PostgreSQL, covering replication, high availability, backups, recovery, performance, and migrations
  • Support data-intensive and analytical platforms such as StarRocks, ClickHouse, Apache Doris, or comparable technologies
  • Design and operate platforms that remain resilient through node, service, zone, and infrastructure failures
  • Implement and test backup, recovery, disaster-recovery, failover, and business-continuity capabilities
  • Apply multi-zone, multi-region, and active-active architecture patterns when appropriate
  • Develop multi-tenant platforms with suitable isolation, scaling, resource controls, and reliability
  • Take part in disaster-recovery exercises, failure simulations, migrations, and broader resilience programs
  • Automate infrastructure and platform lifecycle tasks with Terraform, Python, Bash, Go, or similar technologies
  • Develop dependable deployment and GitOps workflows that reduce manual operational work
  • Implement monitoring, logging, alerting, and observability for distributed workloads
  • Join production incident response, root-cause analysis, and long-term reliability initiatives
  • Provide infrastructure for AI, machine-learning, and data-intensive workloads
  • Advance cloud and Kubernetes platforms for AI workloads, data pipelines, model serving, and platform services
  • Address compute, GPU, networking, storage, data movement, observability, and workload-isolation needs across AI platforms
  • Create reusable platform capabilities with engineering teams to operate AI and data workloads reliably at scale
  • Keep current with infrastructure patterns spanning AI platforms, distributed data systems, and cloud-native technologies

Requirements

  • Bring 4–7 years of experience in platform engineering, infrastructure engineering, distributed systems, SRE, backend engineering, data infrastructure, or a related discipline
  • Demonstrate strong, mandatory production experience with Apache Kafka
  • Demonstrate hands-on, mandatory production experience with Kubernetes
  • Have strong, mandatory experience with either MySQL or PostgreSQL
  • Understand distributed-systems principles including replication, partitioning, consistency, availability, fault tolerance, scalability, and failure recovery
  • Have experience with ArgoCD, GitOps, and infrastructure as code such as Terraform
  • Have operated workloads on a public cloud platform such as GCP, OCI, AWS, or Azure
  • Bring strong production troubleshooting and incident-resolution abilities
  • Use Python, Bash, Go, Java, or comparable languages for automation and scripting
  • Understand high-availability, disaster-recovery, and multi-tenant architecture fundamentals
  • Have experience with observability and operations tools such as Prometheus, Grafana, OpenSearch/ELK, LGTM, or equivalents
  • Understand infrastructure and networking fundamentals in cloud-native environments
  • Preferred: experience with Kafka Connect, Debezium, Kafka Streams, or CDC platforms
  • Preferred: experience with distributed analytical databases such as StarRocks, ClickHouse, Apache Doris, or equivalents
  • Preferred: experience with Flink, Spark, or other distributed data-processing technologies
  • Preferred: experience delivering large-scale data, database, application, or infrastructure migrations
  • Preferred: experience with active-active, multi-zone, or multi-region systems
  • Preferred: experience operating stateful workloads on Kubernetes
  • Preferred: experience supporting AI/ML infrastructure or GPU-based workloads
  • Preferred: experience with cloud networking, service mesh, ingress, load balancing, or storage platforms
  • Preferred: contributions to Kubernetes, Kafka, distributed-systems, or other open-source infrastructure projects

Benefits

  • Competitive pay
  • Premium health insurance
  • A collaborative and enabling culture
  • Meaningful responsibility supported by trust
  • Autonomy in decision-making
  • A lively and engaging work environment
  • The chance to work with leading AI professionals
  • An inclusive workplace that values diversity

Related Jobs

MDY Contact Center

Remote Customer Service Representative – Claro Peru | MDY Contact Center

MDY Contact Center
5,001 – 10,000 Employees
LogisticsMarketingTravel

MDY Contact Center is hiring a full-time remote customer service representative to support Claro Peru customers. The role involves handling inbound calls and resolving inquiries without sales responsibilities.

Open
General Dynamics Information Technology

VMware Engineer with TS/SCI and Polygraph Clearance

General Dynamics Information Technology

Designs and operates secure VMware virtualization environments for General Dynamics Information Technology’s U.S. government customers. This full-time onsite role supports vSphere infrastructure, systems engineering, automation, and operations in Maryland or Virginia.

Open