Baseten
Baseten
Baseten oferă infrastructură pentru inferența modelelor destinată companiilor care implementează aplicații de învățare automată și inteligență artificială. Platforma sa este concepută pentru servirea rapidă și scalabilă a modelelor, combinând inferența cu debit mare, implementarea rapidă, scalarea automată, servirea securizată a modelelor pentru companii și suportul pentru împachetarea modelelor open-source. Baseten ajută echipele de inginerie și învățare automată să gestioneze cerințele operaționale ale infrastructurii pentru modele, astfel încât să se poată concentra pe dezvoltarea unor modele specifice domeniului. Compania activează în domeniile inteligenței artificiale, software-ului ca serviciu și tehnologiei pentru întreprinderi și are o echipă de 11–50 de angajați.

Software Engineer, LLM Inference Platform - Baseten Hybrid

Build Baseten’s distributed LLM inference platform for AI companies. Work across Kubernetes orchestration, model APIs, routing, autoscaling, and production observability.

Descriere

  • Build infrastructure and orchestration for large-scale distributed LLM inference, covering routing, autoscaling, scheduling, and runtime management.
  • Develop and operate Model APIs supporting structured outputs, tool and function calling, and multimodal serving.
  • Implement API versioning, request validation, usage metering, quotas, and authentication.
  • Create instrumentation for metrics, traces, and logs, along with repeatable benchmarks for speed, reliability, and quality.
  • Define strong practices for testing, release automation, and operational excellence.
  • Diagnose and strengthen production systems across Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Collaborate with Inference Performance engineers and partner teams to deliver customer-facing optimizations.
  • Lead projects from architecture and deployment through monitoring and iteration informed by customer feedback.
  • Balance system performance, reliability, operational simplicity, and developer experience.

Cerințe

  • Bachelor’s, master’s, or doctoral degree in computer science, engineering, or a related field, or equivalent practical experience.
  • At least three years of experience building and operating distributed systems, backend infrastructure, or large-scale APIs where reliability, latency, and scale are critical.
  • Demonstrated ownership of low-latency, dependable backend services involving rate limiting, authentication, quotas, metering, and migrations.
  • Strong infrastructure judgment, including experience with profiling, tracing, capacity planning, and SLO management.
  • Ability to diagnose performance and reliability problems across application, runtime, and infrastructure layers.
  • A strong focus on developer experience.
  • Willingness to learn new languages, frameworks, and systems, with an interest in inference engineering.
  • Excellent written communication and collaboration skills, including clear design documentation and effective cross-functional work.
  • Experience with or contributions to LLM inference engines or frameworks such as vLLM, SGLang, TensorRT-LLM, TGI, or Dynamo is a plus.
  • Deep Kubernetes experience, including operators and custom resources; familiarity with service meshes or API gateways is a plus.
  • Experience with distributed scheduling, autoscaling, or service orchestration is a plus.
  • Experience running GPU workloads in production is a plus.
  • Background in developer-facing infrastructure or APIs, or contributions to open-source infrastructure or machine-learning systems, is a plus.
  • Familiarity with observability tools, CI/CD systems, or release automation is a plus.

Beneficii

  • Competitive compensation with meaningful equity.
  • For U.S. employees and dependents, full coverage of medical, dental, and vision insurance.
  • Flexible paid time off, including a company-wide Winter Break.
  • Paid parental leave.
  • Fertility and family-building support through a Carrot stipend.
  • For U.S. employees, a company-facilitated 401(k) plan.
  • Exposure to a range of ML startups, creating substantial opportunities for learning and professional networking.

Locuri de muncă similare