Baseten
Baseten
Baseten उन कंपनियों के लिए मॉडल इन्फरेंस इंफ्रास्ट्रक्चर उपलब्ध कराता है, जो मशीन लर्निंग और एआई एप्लिकेशन तैनात करती हैं। इसका प्लेटफ़ॉर्म तेज़ और बड़े पैमाने पर मॉडल सर्विंग के लिए बनाया गया है, जिसमें उच्च-थ्रूपुट इन्फरेंस, तेज़ परिनियोजन, ऑटोस्केलिंग, सुरक्षित एंटरप्राइज़ मॉडल सर्विंग और ओपन-सोर्स मॉडल पैकेजिंग का समर्थन शामिल है। Baseten इंजीनियरिंग और मशीन लर्निंग टीमों को मॉडल इंफ्रास्ट्रक्चर की परिचालन संबंधी चुनौतियों को संभालने में मदद करता है, ताकि वे डोमेन-विशिष्ट मॉडल विकसित करने पर ध्यान केंद्रित कर सकें। कंपनी कृत्रिम बुद्धिमत्ता, SaaS और एंटरप्राइज़ तकनीक के क्षेत्रों में काम करती है और इसमें 11–50 कर्मचारी हैं।

Software Engineer, LLM Inference Platform - Baseten Hybrid

Build Baseten’s distributed LLM inference platform for AI companies. Work across Kubernetes orchestration, model APIs, routing, autoscaling, and production observability.

विवरण

  • Build infrastructure and orchestration for large-scale distributed LLM inference, covering routing, autoscaling, scheduling, and runtime management.
  • Develop and operate Model APIs supporting structured outputs, tool and function calling, and multimodal serving.
  • Implement API versioning, request validation, usage metering, quotas, and authentication.
  • Create instrumentation for metrics, traces, and logs, along with repeatable benchmarks for speed, reliability, and quality.
  • Define strong practices for testing, release automation, and operational excellence.
  • Diagnose and strengthen production systems across Kubernetes, distributed runtimes, networking, and GPU workloads.
  • Collaborate with Inference Performance engineers and partner teams to deliver customer-facing optimizations.
  • Lead projects from architecture and deployment through monitoring and iteration informed by customer feedback.
  • Balance system performance, reliability, operational simplicity, and developer experience.

आवश्यकताएँ

  • Bachelor’s, master’s, or doctoral degree in computer science, engineering, or a related field, or equivalent practical experience.
  • At least three years of experience building and operating distributed systems, backend infrastructure, or large-scale APIs where reliability, latency, and scale are critical.
  • Demonstrated ownership of low-latency, dependable backend services involving rate limiting, authentication, quotas, metering, and migrations.
  • Strong infrastructure judgment, including experience with profiling, tracing, capacity planning, and SLO management.
  • Ability to diagnose performance and reliability problems across application, runtime, and infrastructure layers.
  • A strong focus on developer experience.
  • Willingness to learn new languages, frameworks, and systems, with an interest in inference engineering.
  • Excellent written communication and collaboration skills, including clear design documentation and effective cross-functional work.
  • Experience with or contributions to LLM inference engines or frameworks such as vLLM, SGLang, TensorRT-LLM, TGI, or Dynamo is a plus.
  • Deep Kubernetes experience, including operators and custom resources; familiarity with service meshes or API gateways is a plus.
  • Experience with distributed scheduling, autoscaling, or service orchestration is a plus.
  • Experience running GPU workloads in production is a plus.
  • Background in developer-facing infrastructure or APIs, or contributions to open-source infrastructure or machine-learning systems, is a plus.
  • Familiarity with observability tools, CI/CD systems, or release automation is a plus.

लाभ

  • Competitive compensation with meaningful equity.
  • For U.S. employees and dependents, full coverage of medical, dental, and vision insurance.
  • Flexible paid time off, including a company-wide Winter Break.
  • Paid parental leave.
  • Fertility and family-building support through a Carrot stipend.
  • For U.S. employees, a company-facilitated 401(k) plan.
  • Exposure to a range of ML startups, creating substantial opportunities for learning and professional networking.

संबंधित नौकरियाँ

Knowtion Health
Napco National

Purchasing Coordinator, Saudi Arabia (Remote)

Napco National

Coordinate purchasing operations for a manufacturing business in Saudi Arabia, from purchase orders and supplier deliveries to customs paperwork and material transfers. Support shipment clearance, invoice processing, supplier claims, and product certificate renewals.

खोलें
Artemys

IT Talent Acquisition Specialist

Artemys

Recruit IT infrastructure professionals for Artemys, a digital transformation company. Lead sourcing and hiring while contributing to employer branding initiatives.

खोलें
4M Analytics

Senior Computer Vision Algorithm Engineer

4M Analytics

Develop computer vision and machine learning systems for 4M Analytics’ subsurface utility mapping platform. Own algorithms from research through production in partnership with data engineering and product teams.

खोलें
Joom

Enterprise Customer Success Manager, Brazil (Remote)

Joom

Lead renewals, retention, and account growth for JoomPulse enterprise customers. Turn Mercado Livre and Shopee analytics into practical recommendations for marketplace sellers.

खोलें
Cyera

Regional Sales Director, Identity Security

Cyera

Lead Cyera’s Pacific Northwest sales team, grow the territory, and exceed quotas for its AI and data security platform.

खोलें
Terumo Medical Corporation

Region Manager, Terumo Interventional Systems Sales

Terumo Medical Corporation

Lead medical device sales across a North Central New Jersey region, managing field teams and hospital relationships. Drive regional revenue, sales performance, and compliant promotion of Terumo Interventional Systems products.

खोलें
Dream

Senior Program Manager, Sovereign AI (Hybrid, Israel)

Dream

Lead Dream’s Sovereign AI roadmap for government and critical infrastructure customers, coordinating research, engineering, product, launches, and AI-enabled operations.

खोलें
Claroty