Designworks Talent LLC
Designworks Talent LLC
Designworks Talent LLC tez o‘sayotgan startaplar va global korxonalarga xizmat ko‘rsatuvchi rekruting hamda iste’dodlar bo‘yicha maslahat kompaniyasidir. 2009-yilda tashkil etilgan kompaniya tezroq va aniqroq ishga yollashni qo‘llab-quvvatlash uchun tajribali rekruterlarni sun’iy intellekt asosidagi vositalar bilan birlashtiradi. Uning faoliyati to‘liq xizmat ko‘rsatish asosidagi qidiruvlar, loyiha asosidagi rekruting va talabga ko‘ra moslashuvchan modellarni, shuningdek, kengaytiriladigan ishga yollash strategiyasi, rekruting operatsiyalari, korporativ iste’dodlarni jalb qilish bo‘yicha rahbarlik, ma’lumotlarga asoslangan nomzodlarni izlash va nomzod tajribasini qamrab oladi. Jamoaning yondashuvi tashkilotlarning ishga yollash ehtiyojlari rivojlanib borgani sari moslashuvchan yordam ko‘rsatish uchun ishlab chiqilgan.

Senior AI Training Infrastructure Engineer

Lead distributed GPU training infrastructure for large-scale AI models on a next-generation cloud platform. Improve reliability, efficiency, fault tolerance, and production readiness across training systems.

Tavsif

  • Build and expand distributed training infrastructure for large AI models across extensive GPU clusters
  • Create and refine systems that improve training reliability, efficiency, and resource utilization
  • Develop fault-tolerant solutions for checkpointing, recovery, and large-scale training operations
  • Integrate AI models into production training pipelines alongside platform, orchestration, and performance engineering teams
  • Investigate and resolve problems affecting training throughput, stability, reliability, and cost efficiency
  • Develop tools and automation that enhance the experience of AI researchers and engineers
  • Define best practices for training infrastructure, operational workflows, and platform reliability
  • Help shape the AI infrastructure platform as an early member of the engineering team
  • Partner with infrastructure, orchestration, performance, and machine learning teams

Talablar

  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure
  • Experience supporting large AI models, foundation models, post-training workflows, or comparable machine learning systems
  • Strong understanding of the reliability, scalability, and efficiency challenges of multi-node GPU training
  • Experience connecting training systems to production machine learning pipelines
  • Strong programming ability and experience working with complex distributed systems
  • Ability to take independent ownership of technically demanding projects
  • Comfort working with substantial ownership and limited process overhead
  • Preferred: Experience with PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies
  • Preferred: Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows
  • Preferred: Experience operating AI training infrastructure at scale for a hyperscaler, AI research organization, cloud provider, or GPU cloud environment
  • Preferred: Experience optimizing GPU utilization, training performance, or distributed-system reliability
  • Preferred: Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms
  • U.S. work authorization is required
  • Visa sponsorship is not currently available

Imtiyozlar

  • Some roles are eligible for merit-based increases
  • Some roles include annual bonus eligibility
  • Some roles offer long-term incentives
  • Medical insurance is available to U.S.-based employees
  • Dental insurance is available to U.S.-based employees
  • Vision insurance is available to U.S.-based employees
  • 401(k) plan
  • Company 401(k) matching
  • Paid holidays each calendar year

O‘xshash ish o‘rinlari

General Dynamics Information Technology

VMware Engineer with TS/SCI and Polygraph Clearance

General Dynamics Information Technology

Designs and operates secure VMware virtualization environments for General Dynamics Information Technology’s U.S. government customers. This full-time onsite role supports vSphere infrastructure, systems engineering, automation, and operations in Maryland or Virginia.

Ochish
Hewlett Packard Enterprise

Aruba Courseware Developer – Mexico Remote

Hewlett Packard Enterprise

Develop customer-focused courseware and technical content for HPE Aruba Networking cloud, automation, and management products. Turn complex networking technologies into practical documentation and learning materials.

Ochish
Ellison Institute of Technology Oxford

Data Engineer, Autonomous Systems — Oxford Hybrid

Ellison Institute of Technology Oxford

Build edge-to-cloud data infrastructure for autonomous scientific labs at the Ellison Institute of Technology Oxford. Support robotics telemetry, knowledge graphs, and machine-learning-ready datasets.

Ochish
Ellison Institute of Technology Oxford

Forward-Deployed Data Engineer – Oxford Hybrid

Ellison Institute of Technology Oxford

Build reproducible biological data pipelines as a Forward-Deployed Data Engineer at the Ellison Institute of Technology Oxford. Support AI, robotics, and life-science research by turning scientific data into practical resources.

Ochish
Akamai Technologies

Senior Software Engineer, Security Services (Israel Remote)

Akamai Technologies

Build browser security, data loss prevention, and AI capabilities at Akamai. Lead cloud-native microservices that protect millions of enterprise workspaces.

Ochish
Family Care Network, Inc.

Bilingual ECM Case Manager – San Luis Obispo Hybrid

Family Care Network, Inc.

Coordinate medical, behavioral health, and social services as an ECM Case Manager. Support children, adults, and families across California’s Central Coast through FCNI’s trauma-informed nonprofit programs.

Ochish