Designworks Talent LLC
Designworks Talent LLC
Designworks Talent LLC is a recruitment and talent advisory firm serving high-growth startups and global enterprises. Founded in 2009, the company combines experienced recruiters with AI-enabled tools to support faster, more targeted hiring. Its work spans full-service searches, project-based recruiting, and flexible on-demand models, alongside scalable hiring strategy, recruiting operations, enterprise talent acquisition leadership, data-informed sourcing, and candidate experience. The team’s approach is designed to give organizations adaptable support as their hiring needs evolve.

Senior AI Training Infrastructure Engineer

Lead distributed GPU training infrastructure for large-scale AI models on a next-generation cloud platform. Improve reliability, efficiency, fault tolerance, and production readiness across training systems.

Description

  • Build and expand distributed training infrastructure for large AI models across extensive GPU clusters
  • Create and refine systems that improve training reliability, efficiency, and resource utilization
  • Develop fault-tolerant solutions for checkpointing, recovery, and large-scale training operations
  • Integrate AI models into production training pipelines alongside platform, orchestration, and performance engineering teams
  • Investigate and resolve problems affecting training throughput, stability, reliability, and cost efficiency
  • Develop tools and automation that enhance the experience of AI researchers and engineers
  • Define best practices for training infrastructure, operational workflows, and platform reliability
  • Help shape the AI infrastructure platform as an early member of the engineering team
  • Partner with infrastructure, orchestration, performance, and machine learning teams

Requirements

  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure
  • Experience supporting large AI models, foundation models, post-training workflows, or comparable machine learning systems
  • Strong understanding of the reliability, scalability, and efficiency challenges of multi-node GPU training
  • Experience connecting training systems to production machine learning pipelines
  • Strong programming ability and experience working with complex distributed systems
  • Ability to take independent ownership of technically demanding projects
  • Comfort working with substantial ownership and limited process overhead
  • Preferred: Experience with PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies
  • Preferred: Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows
  • Preferred: Experience operating AI training infrastructure at scale for a hyperscaler, AI research organization, cloud provider, or GPU cloud environment
  • Preferred: Experience optimizing GPU utilization, training performance, or distributed-system reliability
  • Preferred: Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms
  • U.S. work authorization is required
  • Visa sponsorship is not currently available

Benefits

  • Some roles are eligible for merit-based increases
  • Some roles include annual bonus eligibility
  • Some roles offer long-term incentives
  • Medical insurance is available to U.S.-based employees
  • Dental insurance is available to U.S.-based employees
  • Vision insurance is available to U.S.-based employees
  • 401(k) plan
  • Company 401(k) matching
  • Paid holidays each calendar year

Related Jobs