CodiLime
CodiLime
201 – 500 Employees
B2BHardwareTelecommunications
CodiLime is a software and network engineering services company founded in 2011, working with networking hardware vendors, software providers, telecommunications firms, technology startups, and established industry companies. Its teams help clients validate ideas through proofs of concept, build network-focused products, and support systems in production. The company’s expertise includes network automation, low-level systems programming, observability, DevOps, and cybersecurity, with project experience spanning the US, Japan, Israel, and Europe. CodiLime’s hiring context is centered on engineering work for B2B technology and telecommunications projects.

Senior Platform Validation Engineer, RDMA — Remote Poland

CodiLime is hiring a senior Platform Validation Engineer to test and optimize NVIDIA AI data-center networks using RDMA technologies. The role focuses on congestion control, performance benchmarking, network automation, and telemetry analysis.

Description

  • Configure, troubleshoot, and maintain NVIDIA/Mellanox Spectrum switches and ConnectX network interface cards.
  • Implement and tune RoCEv2, Priority-based Flow Control, and Explicit Congestion Notification for AI workloads.
  • Tune network environments to maximize throughput, reduce latency, and minimize Job Completion Time.
  • Benchmark and validate AI-ready data-center fabrics using RDMA and the NVIDIA Collective Communication Library.
  • Apply AI agents and specialized tools to automate network operations, testing, and telemetry analysis.
  • Create and maintain documentation for network architectures, configurations, and benchmarking outcomes.
  • Collaborate within an Agile project team in a startup environment supporting a US-based client.

Requirements

  • At least seven years of professional network engineering experience.
  • Preferably three or more years focused on data-center network design and architecture.
  • Hands-on expertise configuring and troubleshooting BGP, EVPN/VXLAN, LACP, and ECMP.
  • Strong knowledge of AI-focused data-center architectures, including Rail-Optimized Design and Rail-Unified Design topologies.
  • Experience configuring, administering, and troubleshooting NVIDIA/Mellanox Spectrum switches and ConnectX network interface cards.
  • Advanced understanding of RoCEv2, Priority-based Flow Control, and Explicit Congestion Notification.
  • Experience tuning network environments, including switch buffers and Quality of Service policies.
  • Experience benchmarking and validating AI-ready data-center fabrics with RDMA and the NVIDIA Collective Communication Library.
  • Ability to use AI agents and tools for task automation, testing, and analysis.
  • Strong technical writing skills and professional English communication.
  • Professional certification such as CCNP, CCIE Enterprise or Data Center, JNCIP, or an equivalent credential is desirable.
  • Hands-on experience with network automation and validation frameworks such as pyATS, NAPALM, Batfish, Nornir, or NetBox is desirable.
  • Exposure to CI/CD pipelines, including GitLab CI, GitHub Actions, or Jenkins, and infrastructure-as-code concepts is desirable.

Benefits

  • Flexible working arrangements with fully remote, office-based, and hybrid options.
  • Professional development through internal training and a dedicated training budget.
  • Structured, hands-on onboarding.
  • A collaborative atmosphere with colleagues passionate about their work.
  • Opportunities to change projects over time.

Related Jobs

Aleph Alpha

Senior AI Researcher, Foundation Model Pre-Training — Aleph Alpha, Heidelberg

Aleph Alpha
51 – 200 Employees
Artificial IntelligenceB2B

Lead architecture and large-scale pre-training work for Aleph Alpha’s European foundation models, shaping training methods across thousands of GPUs. Build and refine PyTorch training recipes with the Heidelberg-based hybrid team.

Open