CodiLime
CodiLime
201 – 500 Employees
B2BHardwareTelecommunications
CodiLime is a software and network engineering services company founded in 2011, working with networking hardware vendors, software providers, telecommunications firms, technology startups, and established industry companies. Its teams help clients validate ideas through proofs of concept, build network-focused products, and support systems in production. The company’s expertise includes network automation, low-level systems programming, observability, DevOps, and cybersecurity, with project experience spanning the US, Japan, Israel, and Europe. CodiLime’s hiring context is centered on engineering work for B2B technology and telecommunications projects.

Senior RDMA Performance Engineer — Poland Remote

Senior RDMA Performance Engineer role at CodiLime, focused on optimizing NVIDIA AI data-center fabrics for US-based clients. The position covers RDMA benchmarking, congestion-control tuning, and network performance engineering.

Description

  • Configure, troubleshoot, and maintain NVIDIA/Mellanox Spectrum switches and ConnectX Network Interface Cards (NICs)
  • Implement and tune congestion-control mechanisms for AI workloads, including RoCEv2, Priority-based Flow Control (PFC), and Explicit Congestion Notification (ECN)
  • Optimize network environments for high throughput, low latency, and reduced Job Completion Time (JCT)
  • Run performance benchmarks and validation tests on AI-ready data-center fabrics using RDMA and NVIDIA Collective Communication Library (NCCL)
  • Use AI agents and specialized tools to automate network tasks, perform testing, and analyze network telemetry
  • Create and maintain technical documentation for network designs, configurations, and benchmarking results
  • Collaborate within an Agile project team in a startup environment serving a US-based client

Requirements

  • At least 7 years of professional network-engineering experience
  • Preferably 3 or more years focused on data-center network design and architecture
  • Hands-on experience configuring and troubleshooting BGP, EVPN/VXLAN, LACP, and ECMP
  • Deep knowledge of AI-dedicated data-center designs, including Rail-Optimized Design (ROD) and Rail-Unified Design (RUD) topologies, with practical configuration and deployment experience
  • Experience configuring, managing, and troubleshooting NVIDIA/Mellanox Spectrum switches and ConnectX Network Interface Cards (NICs)
  • Deep knowledge of RoCEv2, Priority-based Flow Control (PFC), and Explicit Congestion Notification (ECN)
  • Ability to tune and optimize network environments, including switch buffers and Quality of Service (QoS) policies
  • Experience benchmarking and validating AI-ready data-center fabrics with RDMA and NVIDIA Collective Communication Library (NCCL)
  • Ability to use AI agents and tools for task automation, testing, and analysis
  • Strong technical-writing skills and professional English communication
  • Professional-level certification such as CCNP/CCIE Enterprise or Data Center, JNCIP, or equivalent preferred
  • Hands-on experience with network automation and validation frameworks such as pyATS, NAPALM, Batfish, Nornir, or NetBox preferred
  • Exposure to CI/CD pipelines, including GitLab CI, GitHub Actions, or Jenkins, and infrastructure-as-code concepts preferred

Benefits

  • Flexible working arrangements, including fully remote, office-based, or hybrid work
  • Professional development through internal training sessions and a training budget
  • Structured, hands-on onboarding to support a smooth start
  • A collaborative environment with professionals who are passionate about their work
  • The opportunity to change projects

Related Jobs

Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Employees
Artificial IntelligenceB2B

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Open