Baseten
Baseten
Baseten provides model inference infrastructure for companies deploying machine learning and AI applications. Its platform is built for fast, scalable serving, combining high-throughput inference, rapid deployment, autoscaling, secure enterprise model serving, and support for open-source model packaging. Baseten helps engineering and machine learning teams manage the operational demands of model infrastructure so they can focus on developing domain-specific models. The company operates across artificial intelligence, SaaS, and enterprise technology, with a team of 11–50 employees.

Software Engineer, LLM Inference Platform - Baseten Hybrid

Build Baseten’s distributed LLM inference platform for AI companies. Work across Kubernetes orchestration, model APIs, routing, autoscaling, and production observability.

Description

  • Build infrastructure and orchestration for large-scale distributed LLM inference, covering routing, autoscaling, scheduling, and runtime management.
  • Develop and operate Model APIs supporting structured outputs, tool and function calling, and multimodal serving.
  • Implement API versioning, request validation, usage metering, quotas, and authentication.
  • Create instrumentation for metrics, traces, and logs, along with repeatable benchmarks for speed, reliability, and quality.
  • Define strong practices for testing, release automation, and operational excellence.
  • Diagnose and strengthen production systems across Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Collaborate with Inference Performance engineers and partner teams to deliver customer-facing optimizations.
  • Lead projects from architecture and deployment through monitoring and iteration informed by customer feedback.
  • Balance system performance, reliability, operational simplicity, and developer experience.

Requirements

  • Bachelor’s, master’s, or doctoral degree in computer science, engineering, or a related field, or equivalent practical experience.
  • At least three years of experience building and operating distributed systems, backend infrastructure, or large-scale APIs where reliability, latency, and scale are critical.
  • Demonstrated ownership of low-latency, dependable backend services involving rate limiting, authentication, quotas, metering, and migrations.
  • Strong infrastructure judgment, including experience with profiling, tracing, capacity planning, and SLO management.
  • Ability to diagnose performance and reliability problems across application, runtime, and infrastructure layers.
  • A strong focus on developer experience.
  • Willingness to learn new languages, frameworks, and systems, with an interest in inference engineering.
  • Excellent written communication and collaboration skills, including clear design documentation and effective cross-functional work.
  • Experience with or contributions to LLM inference engines or frameworks such as vLLM, SGLang, TensorRT-LLM, TGI, or Dynamo is a plus.
  • Deep Kubernetes experience, including operators and custom resources; familiarity with service meshes or API gateways is a plus.
  • Experience with distributed scheduling, autoscaling, or service orchestration is a plus.
  • Experience running GPU workloads in production is a plus.
  • Background in developer-facing infrastructure or APIs, or contributions to open-source infrastructure or machine-learning systems, is a plus.
  • Familiarity with observability tools, CI/CD systems, or release automation is a plus.

Benefits

  • Competitive compensation with meaningful equity.
  • For U.S. employees and dependents, full coverage of medical, dental, and vision insurance.
  • Flexible paid time off, including a company-wide Winter Break.
  • Paid parental leave.
  • Fertility and family-building support through a Carrot stipend.
  • For U.S. employees, a company-facilitated 401(k) plan.
  • Exposure to a range of ML startups, creating substantial opportunities for learning and professional networking.

Related Jobs