Baseten
Baseten
Baseten provides model inference infrastructure for companies deploying machine learning and AI applications. Its platform is built for fast, scalable serving, combining high-throughput inference, rapid deployment, autoscaling, secure enterprise model serving, and support for open-source model packaging. Baseten helps engineering and machine learning teams manage the operational demands of model infrastructure so they can focus on developing domain-specific models. The company operates across artificial intelligence, SaaS, and enterprise technology, with a team of 11–50 employees.

Engineering Manager, GPU Inference Performance at Baseten

Lead GPU inference performance engineering at Baseten, an AI infrastructure company. Guide runtime optimization, model efficiency, and the growth of a specialized technical team.

Description

  • Manage, mentor, and develop inference performance engineers through regular one-to-ones, feedback, career planning, and performance reviews
  • Recruit GPU and inference engineers while fostering a collaborative team culture
  • Own the runtime performance roadmap and its execution
  • Review technical designs, direct profiling and optimization work, and help engineers assess time and memory use
  • Productionize quantization, speculative decoding, KV-cache reuse, chunked prefill, and custom scheduling
  • Translate performance gains into measurable improvements in tokens per GPU-hour, utilization, latency, and cost
  • Bring new model architectures onto emerging hardware and tune their performance
  • Collaborate with Infrastructure, Inference Platform, Kernels, Model APIs, and customer-facing teams to prioritize work, coordinate launches, and deliver improvements
  • Establish engineering standards for quality, benchmarking, operational excellence, and incident response
  • Advance agentic inference optimization, agentic kernels, speculative-decoding model training, and the Baseten Inference Stack

Requirements

  • Bachelor’s, master’s, or doctoral degree in computer science, engineering, mathematics, or a related discipline
  • Management experience covering engineering hiring, mentoring, feedback, and performance reviews
  • Experience leading or closely supporting GPU optimization teams in training, inference, or recommendation systems
  • Deep technical knowledge of GPU workloads, architecture, and performance tradeoffs
  • Familiarity with machine learning libraries including PyTorch, TensorRT, or TensorRT-LLM
  • Demonstrated ability to set roadmaps and deliver complex technical projects with a team
  • Strong written and verbal communication skills for aligning stakeholders across teams
  • Familiarity with inference engines such as vLLM, SGLang, or TensorRT-LLM
  • Experience applying LLM optimization methods such as quantization, speculative decoding, and continuous batching
  • Experience developing or working with GPU kernels using CUDA, Triton, or CUTLASS
  • Experience scaling an engineering team during rapid startup growth
  • Prior hands-on experience as a performance or systems engineer before moving into management

Benefits

  • Competitive compensation package with meaningful equity
  • For U.S. employees, full medical, dental, and vision coverage for employees and dependents
  • Flexible paid time off, including a company-wide winter break
  • Paid parental leave
  • Fertility and family-building support through Carrot
  • For U.S. employees, a company-facilitated 401(k) plan
  • Opportunities to work with a range of machine learning startups and build industry connections

Related Jobs

Pillsbury Winthrop Shaw Pittman LLP

Senior Business Manager, AI Legal Operations (Onsite, Nashville)

Pillsbury Winthrop Shaw Pittman LLP

Oversee Pillsbury’s AI-enabled legal production service, from intake and staffing through delivery, reporting, and billing coordination. Improve workflows and coordinate teams supporting assignments across U.S. offices.

Open
Avalon Healthcare Solutions

Senior AI Engineer, Florida Remote | Avalon Healthcare Solutions

Avalon Healthcare Solutions

Build and lead production AI systems—including agentic workflows, retrieval-augmented generation, and automation—for Avalon Healthcare Solutions’ diagnostic intelligence platform. Help improve healthcare operations and clinical outcomes through reliable, governed AI solutions.

Open
Avalon Healthcare Solutions

Senior AI Engineer, Generative AI and Agentic Systems

Avalon Healthcare Solutions

Build and lead production generative AI and agentic systems for Avalon Healthcare Solutions’ diagnostic intelligence platform. Guide deployment, automation, evaluation, and technical direction.

Open