Baseten
Baseten
Baseten provides model inference infrastructure for companies deploying machine learning and AI applications. Its platform is built for fast, scalable serving, combining high-throughput inference, rapid deployment, autoscaling, secure enterprise model serving, and support for open-source model packaging. Baseten helps engineering and machine learning teams manage the operational demands of model infrastructure so they can focus on developing domain-specific models. The company operates across artificial intelligence, SaaS, and enterprise technology, with a team of 11–50 employees.

Inference Performance Software Engineer at Baseten (Hybrid)

Baseten is hiring an inference performance engineer to optimize LLM runtimes, GPU kernels, and model-serving systems. The role focuses on improving production inference speed, hardware utilization, and cost efficiency.

Description

  • Productionize advanced inference methods such as quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling or routing.
  • Analyze the full inference pipeline, optimizing kernel overhead, memory placement, request scheduling, prefill/decode separation, and cache-aware routing.
  • Increase tokens per GPU-hour and utilization while balancing latency, throughput, and cost.
  • Deploy and optimize emerging model architectures on new hardware platforms.
  • Create benchmarks spanning model types, batch sizes, sequence lengths, and hardware setups.
  • Submit improvements to open-source inference projects including vLLM, SGLang, and TensorRT-LLM.
  • Work with model, infrastructure, and customer-facing teams to deliver measurable performance gains.

Requirements

  • Degree at the bachelor's, master's, or doctoral level in computer science, engineering, mathematics, or a related discipline.
  • Professional experience with at least one general-purpose programming language, including Python or C++.
  • Working knowledge of LLM optimization methods such as quantization, speculative decoding, and continuous batching.
  • Strong command of machine-learning libraries, particularly PyTorch, TensorRT, or TensorRT-LLM.
  • Demonstrated interest in and practical experience with large language models.
  • Thorough understanding of GPU architecture.

Benefits

  • Competitive pay package with meaningful equity participation.
  • For U.S. employees, full medical, dental, and vision insurance coverage extends to employees and their dependents.
  • Flexible paid time off, plus a company-wide winter break.
  • Paid parental leave.
  • Fertility and family-building support through a Carrot stipend.
  • For U.S. employees, a company-facilitated 401(k) plan.
  • Opportunities to work with a range of machine-learning startups, supporting substantial learning and professional networking.

Related Jobs