Hewlett Packard Enterprise
Hewlett Packard Enterprise
Hewlett Packard Enterprise гибридті бұлт пен жасанды интеллектке негізделген операцияларды жүргізетін ұйымдарға арналған корпоративтік технологияларды әзірлейді. Оның портфеліне HPE ProLiant серверлері мен HPE Cray жоғары өнімді есептеу жүйелері, HPE Alletra сақтау жүйелері, HPE Aruba желілік шешімдері, жасанды интеллект платформалары мен дайын AI фабрикалары, сондай-ақ қызмет ретінде ұсынылатын GreenLake шешімдері кіреді. HPE сонымен қатар ірі ұйымдарға инфрақұрылымын жаңғыртуға, деректер орталығы орталарын басқаруға және жасанды интеллектіні ауқымды түрде енгізуге көмектесетін Zero Trust және SASE сияқты қауіпсіздік мүмкіндіктерін, сондай-ақ кәсіби, консультациялық және қолдау қызметтерін ұсынады.

Senior Inference Software Engineer at HPE Hybrid

Lead development of HPE AI Essentials’ enterprise LLM inference runtime for customer-owned infrastructure. Improve GPU-efficient serving, distributed execution, and Kubernetes orchestration.

Сипаттама

  • Own core components of the LLM serving deployment, including engine integration, continuous batching, KV-cache management and reuse, and quantized execution.
  • Work with inference engineering teams to reduce time to first token, inter-token latency, GPU throughput, and P95/P99 tail latency.
  • Develop and operate distributed execution features such as disaggregated prefill and decode, tensor and pipeline parallelism, and KV-cache offload across GPU memory, host memory, and RDMA-attached storage.
  • Assess emerging runtimes, quantization methods, speculative decoding, and mixture-of-experts serving, then recommend technologies for adoption.
  • Help build the orchestration layer covering model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling.
  • Resolve customer issues end to end, determine root causes, and strengthen related systems and processes.
  • Review code and designs, mentor colleagues, and model strong engineering practices.

Талаптар

  • At least 8 years of software engineering experience.
  • One to two or more years of direct experience with LLM inference runtimes or production model serving.
  • A degree in computer science or a related discipline.
  • Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including experience modifying engine internals.
  • Strong knowledge of continuous batching, paged attention, KV-cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
  • Working knowledge of tensor and pipeline parallelism, NCCL collective operations, GPU memory hierarchy, and interconnect behavior.
  • Advanced proficiency with Kubernetes platform architecture, including operators, custom resources, controllers, and scheduling.
  • Strong programming skills in Go and Python.
  • Ability to read, debug, and profile C++ and CUDA with tools such as Nsight.
  • Familiarity with debugging and profiling multi-tier workloads, including RAG and agent applications.
  • Excellent analytical, debugging, and problem-solving skills.
  • Preferred: Contributions to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe.
  • Preferred: Experience with disaggregated prefill and decode serving or large-scale KV-cache offload and reuse.
  • Preferred: Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE.
  • Preferred: Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation.
  • Preferred: Experience delivering software for on-premises, air-gapped, or regulated enterprise environments.

Артықшылықтар

  • Comprehensive benefits supporting physical health, financial security, and emotional wellbeing.
  • Programs supporting personal and professional growth.
  • Flexibility for work arrangements and personal needs.
  • An inclusive workplace culture.
  • Reasonable accommodations during the application or interview process for qualified applicants with disabilities.

Ұқсас жұмыс орындары

Yudrio, Inc.

Senior Salesforce Technical Architect

Yudrio, Inc.

Lead secure Service Cloud and Experience Cloud architecture for Yudrio’s federal Salesforce implementation. Guide integrations, platform governance, releases, and Agile delivery.

Ашу
EBARA Precision Machinery Europe (EPME)

Field Service Engineer, Semiconductor Equipment

EBARA Precision Machinery Europe (EPME)
201 – 500 Қызметкерлер
КонсалтингЛогистикаӨндіріс

Install, troubleshoot, and maintain semiconductor equipment for EBARA Precision Machinery Europe, providing on-site customer support across international sites.

Ашу
SPERTON - Where Great People Meet

Smart Home Automation Business Development Executive

SPERTON - Where Great People Meet
51 – 200 Қызметкерлер

Drive sales of Smartwitz Marketing Solutions’ smart home automation products in Navi Mumbai. Manage client relationships, demonstrations, proposals, and opportunities through installation.

Ашу
Personalberatung Pillong

Senior International Sales Manager, PV Mounting Systems

Personalberatung Pillong

Lead international customer acquisition and key accounts for a global industrial group’s photovoltaic mounting systems. Develop solar project markets and partnerships worldwide.

Ашу
Inkster GmbH

Marketing Manager, Hybrid in Hamburg

Inkster GmbH

Lead campaign creative development, social media content, and performance analysis at Inkster, a European provider of temporary fruit-based tattoos. Test and optimize marketing creatives using campaign insights and market trends.

Ашу
Salonkee

Sales Representative – Salon Software, Braunschweig

Salonkee
51 – 200 Қызметкерлер
SaaSМаркетплейсСұлулық

Help Salonkee grow its salon software business in Braunschweig through prospecting, product demonstrations and managing sales from first contact to close.

Ашу
Salonkee

Salonkee Sales Representative (Career Changer) – Hannover

Salonkee
51 – 200 Қызметкерлер
SaaSМаркетплейсСұлулық

Grow Salonkee’s salon software business in Hannover by building relationships with salon owners and guiding them through product demonstrations. Manage the sales process from prospecting to close while planning your week independently.

Ашу
Red Hat

Consulting Cloud Architect, OpenShift

Red Hat

Design and deliver Red Hat OpenShift, Kubernetes, and hybrid cloud solutions for enterprise customers. Lead customer engagements focused on infrastructure, automation, and application delivery.

Ашу