CodiLime
CodiLime
201 – 500 Employees
B2BHardwareTelecommunications
CodiLime is a software and network engineering services company founded in 2011, working with networking hardware vendors, software providers, telecommunications firms, technology startups, and established industry companies. Its teams help clients validate ideas through proofs of concept, build network-focused products, and support systems in production. The company’s expertise includes network automation, low-level systems programming, observability, DevOps, and cybersecurity, with project experience spanning the US, Japan, Israel, and Europe. CodiLime’s hiring context is centered on engineering work for B2B technology and telecommunications projects.

Senior RDMA Performance Engineer — Remote Poland

CodiLime is hiring a senior engineer to validate SONiC and NVIDIA Spectrum AI data-center networks in Poland. The role focuses on RDMA performance, congestion control, benchmarking, and network automation.

Description

  • Configure, troubleshoot, and maintain NVIDIA/Mellanox Spectrum switches and ConnectX network interface cards.
  • Implement and tune RoCEv2, PFC, and ECN congestion controls for AI workloads.
  • Optimize network environments for throughput, latency, and Job Completion Time.
  • Benchmark and validate AI-ready data-center fabrics using RDMA and NCCL.
  • Automate network operations, testing, and telemetry analysis with AI agents and specialized tools.
  • Create and maintain documentation for network designs, configurations, and benchmark results.
  • Collaborate in an Agile startup team validating SONiC on NVIDIA Spectrum AI networking platforms.

Requirements

  • At least seven years of professional network engineering experience.
  • Preferably three or more years designing and architecting data-center networks.
  • Hands-on expertise with BGP, EVPN/VXLAN, LACP, and ECMP configuration and troubleshooting.
  • Strong practical knowledge of AI-focused data-center architectures, including Rail-Optimized Design and Rail-Unified Design topologies.
  • Demonstrated experience managing and troubleshooting NVIDIA/Mellanox Spectrum switches and ConnectX network interface cards.
  • Advanced knowledge of RoCEv2, Priority-based Flow Control, and Explicit Congestion Notification.
  • Experience tuning switch buffers and deploying Quality of Service policies.
  • Experience benchmarking and validating AI-ready data-center fabrics with RDMA and NVIDIA Collective Communication Library.
  • Ability to use AI agents and tools to automate engineering tasks, testing, and analysis.
  • Strong technical writing and documentation abilities.
  • Professional English communication skills.
  • Professional certifications such as CCNP, CCIE Enterprise or Data Center, or JNCIP are appreciated.
  • Hands-on experience with pyATS, NAPALM, Batfish, Nornir, or NetBox is appreciated.
  • Familiarity with GitLab CI, GitHub Actions, Jenkins, and infrastructure-as-code concepts is appreciated.

Benefits

  • Flexible working arrangements, with fully remote, office-based, and hybrid options.
  • Professional development through internal training and a dedicated training budget.
  • Structured onboarding with practical, hands-on support.
  • A collaborative environment of professionals who care about their work.
  • The opportunity to move between projects.

Related Jobs

Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Employees
Artificial IntelligenceB2B

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Open