Designworks Talent LLC
Designworks Talent LLC
Designworks Talent LLC is a recruitment and talent advisory firm serving high-growth startups and global enterprises. Founded in 2009, the company combines experienced recruiters with AI-enabled tools to support faster, more targeted hiring. Its work spans full-service searches, project-based recruiting, and flexible on-demand models, alongside scalable hiring strategy, recruiting operations, enterprise talent acquisition leadership, data-informed sourcing, and candidate experience. The team’s approach is designed to give organizations adaptable support as their hiring needs evolve.

Staff Inference Engineer - Bellevue Hybrid

Lead the development of production model-serving systems for an AI infrastructure company. Improve GPU utilization, latency, throughput, scalability, and reliability across large-scale AI workloads.

Description

  • Build and operate production model-serving and inference systems for high-throughput, low-latency AI workloads
  • Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across diverse models and workloads
  • Design systems that use GPUs efficiently while delivering consistent performance and reliability
  • Advance inference-platform scalability and operational maturity as customer demand increases
  • Work with AI training, GPU performance, orchestration, and infrastructure teams to move models from development into production
  • Establish monitoring, alerting, and operational practices that support dependable inference services
  • Diagnose and resolve performance, reliability, and capacity issues across inference workloads
  • Shape architecture decisions and engineering standards as the platform develops

Requirements

  • Experience building and operating production machine learning inference or model-serving systems at scale
  • Strong understanding of latency, throughput, memory utilization, and cost-efficiency trade-offs involved in serving large AI models
  • Experience designing reliable distributed systems or production infrastructure
  • Understanding of GPU-backed AI workloads and the challenges of scaling inference systems
  • Strong engineering fundamentals and the ability to independently own complex technical problems
  • Comfort working in a fast-moving environment where systems and processes are built from the ground up
  • Experience with modern inference-serving frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or comparable technologies
  • Experience optimizing LLM inference workloads or large-scale AI serving platforms
  • Background operating API-based AI products or high-volume production services
  • Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms
  • Familiarity with quantization, batching, caching, or performance tuning
  • Experience at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization
  • U.S. work authorization is required
  • Visa sponsorship is not currently available
  • Willingness to work in a hybrid arrangement in downtown Bellevue, Washington, with at least three days per week in the office once the permanent office is established

Benefits

  • Competitive base pay for the Bellevue market
  • Merit-based increases
  • Annual bonus
  • Long-term incentives
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • 401(k) plan
  • Company 401(k) match
  • Paid holidays
  • Startup ownership and meaningful technical impact
  • Collaboration with an experienced AI infrastructure team

Related Jobs

RockstarDevelopers GmbH

Senior Full Stack Developer, Java and Angular

RockstarDevelopers GmbH
11 – 50 Employees
ConsultingLogisticsMarketing

Build Java and Angular solutions for enterprise and regulated-sector clients at RockstarDevelopers GmbH. Work remotely from Germany with modern AI coding tools.

Open