Baseten
Baseten
Baseten mashinaviy o‘rganish va sun’iy intellekt ilovalarini ishga tushiruvchi kompaniyalar uchun model inferensiyasi infratuzilmasini taqdim etadi. Uning platformasi yuqori o‘tkazuvchanlikdagi inferensiya, tezkor joriy etish, avtomatik masshtablash, xavfsiz korporativ model taqdimoti va ochiq kodli modellarni paketlashni qo‘llab-quvvatlashni birlashtirib, tez va masshtablanadigan xizmat ko‘rsatish uchun yaratilgan. Baseten muhandislik hamda mashinaviy o‘rganish jamoalariga model infratuzilmasining operatsion talablarini boshqarishda yordam beradi, shunda ular muayyan sohaga mos modellarni ishlab chiqishga e’tibor qaratishi mumkin. Kompaniya sun’iy intellekt, SaaS va korporativ texnologiyalar sohalarida faoliyat yuritadi, jamoasida 11–50 nafar xodim bor.

Inference Performance Software Engineer at Baseten (Hybrid)

Baseten is hiring an inference performance engineer to optimize LLM runtimes, GPU kernels, and model-serving systems. The role focuses on improving production inference speed, hardware utilization, and cost efficiency.

Tavsif

  • Productionize advanced inference methods such as quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling or routing.
  • Analyze the full inference pipeline, optimizing kernel overhead, memory placement, request scheduling, prefill/decode separation, and cache-aware routing.
  • Increase tokens per GPU-hour and utilization while balancing latency, throughput, and cost.
  • Deploy and optimize emerging model architectures on new hardware platforms.
  • Create benchmarks spanning model types, batch sizes, sequence lengths, and hardware setups.
  • Submit improvements to open-source inference projects including vLLM, SGLang, and TensorRT-LLM.
  • Work with model, infrastructure, and customer-facing teams to deliver measurable performance gains.

Talablar

  • Degree at the bachelor's, master's, or doctoral level in computer science, engineering, mathematics, or a related discipline.
  • Professional experience with at least one general-purpose programming language, including Python or C++.
  • Working knowledge of LLM optimization methods such as quantization, speculative decoding, and continuous batching.
  • Strong command of machine-learning libraries, particularly PyTorch, TensorRT, or TensorRT-LLM.
  • Demonstrated interest in and practical experience with large language models.
  • Thorough understanding of GPU architecture.

Imtiyozlar

  • Competitive pay package with meaningful equity participation.
  • For U.S. employees, full medical, dental, and vision insurance coverage extends to employees and their dependents.
  • Flexible paid time off, plus a company-wide winter break.
  • Paid parental leave.
  • Fertility and family-building support through a Carrot stipend.
  • For U.S. employees, a company-facilitated 401(k) plan.
  • Opportunities to work with a range of machine-learning startups, supporting substantial learning and professional networking.

O‘xshash ish o‘rinlari

Terumo Medical Corporation

Territory Manager, Interventional Systems — Philadelphia

Terumo Medical Corporation

Lead sales of Terumo Interventional Systems devices to hospitals and outpatient facilities in the Philadelphia area. Build accounts, support procedures, educate clinicians, and meet territory sales goals.

Ochish