Baseten
Baseten
Baseten машиналық оқыту және жасанды интеллект қолданбаларын іске қосатын компанияларға модель инференсі инфрақұрылымын ұсынады. Оның платформасы жоғары өткізу қабілеті бар инференсті, жылдам орналастыруды, автоматты масштабтауды, қауіпсіз корпоративтік модельдерді ұсынуды және ашық бастапқы кодты модельдерді қаптамаға жинауға қолдауды біріктіріп, жылдам әрі ауқымды қызмет көрсетуге арналған. Baseten инженерлік және машиналық оқыту командаларына модель инфрақұрылымының операциялық талаптарын басқаруға көмектеседі, осылайша олар нақты салаға арналған модельдерді әзірлеуге назар аудара алады. Компания жасанды интеллект, SaaS және корпоративтік технологиялар салаларында жұмыс істейді, командасында 11–50 қызметкер бар.

Inference Performance Software Engineer at Baseten (Hybrid)

Baseten is hiring an inference performance engineer to optimize LLM runtimes, GPU kernels, and model-serving systems. The role focuses on improving production inference speed, hardware utilization, and cost efficiency.

Сипаттама

  • Productionize advanced inference methods such as quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation, and custom scheduling or routing.
  • Analyze the full inference pipeline, optimizing kernel overhead, memory placement, request scheduling, prefill/decode separation, and cache-aware routing.
  • Increase tokens per GPU-hour and utilization while balancing latency, throughput, and cost.
  • Deploy and optimize emerging model architectures on new hardware platforms.
  • Create benchmarks spanning model types, batch sizes, sequence lengths, and hardware setups.
  • Submit improvements to open-source inference projects including vLLM, SGLang, and TensorRT-LLM.
  • Work with model, infrastructure, and customer-facing teams to deliver measurable performance gains.

Талаптар

  • Degree at the bachelor's, master's, or doctoral level in computer science, engineering, mathematics, or a related discipline.
  • Professional experience with at least one general-purpose programming language, including Python or C++.
  • Working knowledge of LLM optimization methods such as quantization, speculative decoding, and continuous batching.
  • Strong command of machine-learning libraries, particularly PyTorch, TensorRT, or TensorRT-LLM.
  • Demonstrated interest in and practical experience with large language models.
  • Thorough understanding of GPU architecture.

Артықшылықтар

  • Competitive pay package with meaningful equity participation.
  • For U.S. employees, full medical, dental, and vision insurance coverage extends to employees and their dependents.
  • Flexible paid time off, plus a company-wide winter break.
  • Paid parental leave.
  • Fertility and family-building support through a Carrot stipend.
  • For U.S. employees, a company-facilitated 401(k) plan.
  • Opportunities to work with a range of machine-learning startups, supporting substantial learning and professional networking.

Ұқсас жұмыс орындары

Telepatia AI

AI Deployment Strategist, Mexico (Hybrid)

Telepatia AI

Lead Telepatia AI deployments across hospitals, insurers, and clinic networks in Mexico. Guide clinical setup, physician adoption, pilot performance, and deployment teams.

Ашу
Mercor

AI Safety Red Teamer

Mercor
51 – 200 Қызметкерлер

Mercor is seeking an AI Safety Red Teamer to probe advanced AI systems with adversarial prompts. The role involves documenting risks and working with researchers to strengthen model safety and alignment.

Ашу
Maxon

Cloud and Systems Integration Developer

Maxon
201 – 500 Қызметкерлер
МедиаОйындар

Build Azure infrastructure and enterprise integrations supporting Maxon’s creative software operations. Automate deployments, identity workflows, and system connections across the organization.

Ашу
Fanatics

Staff Android Engineer

Fanatics

Lead Android development for Fanatics’ global sports platform, shaping scalable Kotlin and Jetpack Compose architecture. Guide technical standards, incident response, and engineer mentorship.

Ашу
Vega

FP&A Lead at Vega, Hybrid in Israel

Vega

Build Vega’s financial planning, forecasting, and board reporting for its AI-native cybersecurity platform. Work with Sales and Marketing while creating the company’s finance function.

Ашу
AudioCodes

Senior Business Analyst, Oracle Fusion Order Management

AudioCodes

Lead Oracle Fusion order management and global order-to-cash transformation at AudioCodes. Drive automation, system integrations, and AI initiatives across the enterprise voice technology business.

Ашу
iCareManager

Senior AI Engineer, Healthcare Documentation

iCareManager

Lead development of emrgen.ai, iCareManager’s AI clinical documentation platform. Build and evaluate safe healthcare AI for U.S. providers and Medicaid workflows.

Ашу
Terumo Medical Corporation

Territory Manager, Interventional Systems — Philadelphia

Terumo Medical Corporation

Lead sales of Terumo Interventional Systems devices to hospitals and outpatient facilities in the Philadelphia area. Build accounts, support procedures, educate clinicians, and meet territory sales goals.

Ашу
Averna

Senior Account Manager, Test and Quality Solutions

Averna

Grow Averna’s test and quality solutions business across the UK and Ireland. Build customer relationships, manage Salesforce opportunities and negotiate complex technical sales.

Ашу
Manulife

Manulife Operations Supervisor, Hybrid Philippines

Manulife

Lead insurance operations at Manulife, overseeing service levels, workflows, audits, and team performance. This hybrid position supports teams in Quezon City or Lapu-Lapu City.

Ашу