CodiLime
CodiLime
CodiLime — 2011-yilda tashkil etilgan dasturiy ta’minot va tarmoq muhandisligi xizmatlari kompaniyasi. Kompaniya tarmoq uskunalari ishlab chiqaruvchilari, dasturiy ta’minot provayderlari, telekommunikatsiya firmalari, texnologik startaplar va sohada faoliyat yuritadigan yirik kompaniyalar bilan ishlaydi. Uning jamoalari mijozlarga g‘oyalarni konsepsiyaning amaliy isboti orqali tekshirish, tarmoqqa yo‘naltirilgan mahsulotlar yaratish va ishlab turgan tizimlarni qo‘llab-quvvatlashda yordam beradi. Kompaniyaning ekspertiza yo‘nalishlariga tarmoqni avtomatlashtirish, quyi darajadagi tizimli dasturlash, kuzatuvchanlik, DevOps va kiberxavfsizlik kiradi. Loyihalar bo‘yicha tajriba AQSh, Yaponiya, Isroil va Yevropani qamrab oladi. CodiLime kompaniyasining ishga yollash yo‘nalishi B2B texnologiyalari va telekommunikatsiya loyihalaridagi muhandislik ishlariga qaratilgan.

Senior RDMA Performance Engineer (Remote, Egypt)

CodiLime is hiring a senior engineer to validate NVIDIA-based AI data-center networks. The role focuses on RDMA fabrics, congestion control, benchmarking, and network automation.

Tavsif

  • Configure, troubleshoot, and maintain NVIDIA/Mellanox Spectrum switches and ConnectX network interface cards (NICs)
  • Implement and tune congestion-control mechanisms for AI workloads, including RoCEv2, PFC, and ECN
  • Optimize network environments for high throughput, low latency, and reduced Job Completion Time (JCT)
  • Benchmark and validate AI-ready data-center fabrics using RDMA and NVIDIA Collective Communication Library (NCCL)
  • Apply AI agents and specialized tools to automate network operations, testing, and telemetry analysis
  • Create and maintain technical documentation for network designs, configurations, and benchmarking outcomes
  • Collaborate in an Agile startup-style project team supporting a US-based client

Talablar

  • At least 7 years of professional network engineering experience
  • Preferably 3 or more years specializing in data-center network design and architecture
  • Hands-on experience configuring and troubleshooting BGP, EVPN/VXLAN, LACP, and ECMP
  • Deep knowledge of AI-focused data-center architectures, including Rail-Optimized Design (ROD) and Rail-Unified Design (RUD), with practical deployment experience
  • Demonstrated ability to configure, manage, and troubleshoot NVIDIA/Mellanox Spectrum switches and ConnectX network interface cards (NICs)
  • Strong understanding of RoCEv2, Priority-based Flow Control (PFC), and Explicit Congestion Notification (ECN)
  • Ability to tune network environments through switch-buffer optimization and Quality of Service (QoS) policies
  • Experience benchmarking and validating AI-ready data-center fabrics with RDMA and NVIDIA Collective Communication Library (NCCL)
  • Ability to use AI agents and tools for task automation, testing, and analysis
  • Strong technical writing skills and professional English communication
  • Professional certifications such as CCNP/CCIE Enterprise or Data Center, JNCIP, or equivalent are desirable
  • Hands-on experience with pyATS, NAPALM, Batfish, Nornir, or NetBox is desirable
  • Familiarity with GitLab CI, GitHub Actions, Jenkins, and infrastructure-as-code concepts is desirable

Imtiyozlar

  • Choose flexible working arrangements, including fully remote, office-based, or hybrid work
  • Access internal training sessions and a professional development budget
  • Receive structured, hands-on onboarding
  • Join a collaborative team of professionals who care about their work
  • Have the opportunity to move between projects

O‘xshash ish o‘rinlari

Terumo Medical Corporation

Territory Manager, Interventional Systems — Philadelphia

Terumo Medical Corporation

Lead sales of Terumo Interventional Systems devices to hospitals and outpatient facilities in the Philadelphia area. Build accounts, support procedures, educate clinicians, and meet territory sales goals.

Ochish
Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Xodimlar
B2BSun’iy intellekt

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Ochish