adaption
adaption
adaption dezvoltă sisteme și instrumente adaptive de inteligență artificială, concepute pentru a adapta inteligența la utilizatori individuali, limbi, industrii și situații-limită specializate. Activitatea sa acoperă Date Adaptive, Inteligență Adaptivă și Interfețe Adaptive — abordări care ajută modelele să se ajusteze continuu și să ofere experiențe de utilizare mai receptive, fără a se baza la fel de mult pe instrucțiuni manuale sau pe un singur model monolitic. Tehnologia companiei este relevantă pentru organizațiile care explorează aplicații practice ale inteligenței artificiale în domenii precum consultanța, logistica și marketingul.

Inference Engineer — Adaption (Hybrid, Bay Area)

Optimize inference infrastructure for throughput, cost, and latency at Adaption, an AI company developing adaptive intelligence systems. Improve model serving across GPUs, inference engines, and production workloads.

Descriere

  • Lead ownership of inference-stack cost and performance.
  • Increase throughput and reduce cost and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads against real production traffic.
  • Adjust routing between internal infrastructure and external providers according to cost, capacity, and performance.
  • Use serving engines including vLLM, SGLang, and TensorRT-LLM, working beneath the framework when necessary.
  • Develop profiling and measurement systems to locate time, memory, and compute usage.
  • Partner closely with engineers responsible for the serving fleet.

Cerințe

  • Bring at least five years of experience in ML systems, inference infrastructure, or performance engineering, with measurable gains in cost or latency.
  • Demonstrate deep knowledge of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Have production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Use Python confidently and be proficient in C++, Rust, or another systems programming language.
  • Understand GPU performance topics including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
  • Be able to work in person in the Bay Area within a hybrid setup.

Beneficii

  • Combine in-person Bay Area collaboration with a distributed, global-first team and team offsites.
  • Receive an annual Adaption Passport travel stipend for exploring a country you have not previously visited.
  • Get a weekly allowance for takeout meals or grocery delivery.
  • Access comprehensive medical coverage and generous paid time off.

Locuri de muncă similare

Nivoda

Senior Full-Stack Engineer, Growth

Nivoda

Build onboarding, lifecycle, and checkout systems for Nivoda’s global jewellery marketplace. Improve activation, conversion, and retention with event-driven services and experimentation.

Deschide
Capital.com

Technical Product Operations Manager at Capital.com (Hybrid, Poland)

Capital.com

Advance a global trading platform by coordinating engineering programmes and improving how teams work. Lead migrations, data-informed process changes, automation, and practical AI adoption across engineering.

Deschide
Kyndryl

Senior SQL Server Database Administrator

Kyndryl

Administer and improve Kyndryl’s enterprise SQL Server infrastructure in Bangalore or Chennai, supporting performance, security, availability, and disaster recovery. The role covers database operations, lifecycle management, monitoring, and support for application teams.

Deschide