CodiLime
CodiLime
201 – 500 Employees
B2BHardwareTelecommunications
CodiLime is a software and network engineering services company founded in 2011, working with networking hardware vendors, software providers, telecommunications firms, technology startups, and established industry companies. Its teams help clients validate ideas through proofs of concept, build network-focused products, and support systems in production. The company’s expertise includes network automation, low-level systems programming, observability, DevOps, and cybersecurity, with project experience spanning the US, Japan, Israel, and Europe. CodiLime’s hiring context is centered on engineering work for B2B technology and telecommunications projects.

Senior RDMA Performance Engineer (Remote, Egypt)

CodiLime is hiring a senior engineer to validate NVIDIA-based AI data-center networks. The role focuses on RDMA fabrics, congestion control, benchmarking, and network automation.

Description

  • Configure, troubleshoot, and maintain NVIDIA/Mellanox Spectrum switches and ConnectX network interface cards (NICs)
  • Implement and tune congestion-control mechanisms for AI workloads, including RoCEv2, PFC, and ECN
  • Optimize network environments for high throughput, low latency, and reduced Job Completion Time (JCT)
  • Benchmark and validate AI-ready data-center fabrics using RDMA and NVIDIA Collective Communication Library (NCCL)
  • Apply AI agents and specialized tools to automate network operations, testing, and telemetry analysis
  • Create and maintain technical documentation for network designs, configurations, and benchmarking outcomes
  • Collaborate in an Agile startup-style project team supporting a US-based client

Requirements

  • At least 7 years of professional network engineering experience
  • Preferably 3 or more years specializing in data-center network design and architecture
  • Hands-on experience configuring and troubleshooting BGP, EVPN/VXLAN, LACP, and ECMP
  • Deep knowledge of AI-focused data-center architectures, including Rail-Optimized Design (ROD) and Rail-Unified Design (RUD), with practical deployment experience
  • Demonstrated ability to configure, manage, and troubleshoot NVIDIA/Mellanox Spectrum switches and ConnectX network interface cards (NICs)
  • Strong understanding of RoCEv2, Priority-based Flow Control (PFC), and Explicit Congestion Notification (ECN)
  • Ability to tune network environments through switch-buffer optimization and Quality of Service (QoS) policies
  • Experience benchmarking and validating AI-ready data-center fabrics with RDMA and NVIDIA Collective Communication Library (NCCL)
  • Ability to use AI agents and tools for task automation, testing, and analysis
  • Strong technical writing skills and professional English communication
  • Professional certifications such as CCNP/CCIE Enterprise or Data Center, JNCIP, or equivalent are desirable
  • Hands-on experience with pyATS, NAPALM, Batfish, Nornir, or NetBox is desirable
  • Familiarity with GitLab CI, GitHub Actions, Jenkins, and infrastructure-as-code concepts is desirable

Benefits

  • Choose flexible working arrangements, including fully remote, office-based, or hybrid work
  • Access internal training sessions and a professional development budget
  • Receive structured, hands-on onboarding
  • Join a collaborative team of professionals who care about their work
  • Have the opportunity to move between projects

Related Jobs