Hewlett Packard Enterprise
Hewlett Packard Enterprise
Hewlett Packard Enterprise develops enterprise technology for organizations running hybrid cloud and AI-driven operations. Its portfolio spans HPE ProLiant servers and HPE Cray high-performance computing, HPE Alletra storage, HPE Aruba networking, AI platforms and turnkey AI factories, and GreenLake as-a-service solutions. HPE also provides security capabilities such as Zero Trust and SASE, alongside professional, advisory, and support services that help large organizations modernize infrastructure, manage data-center environments, and deploy AI at scale.

Senior Inference Software Engineer at HPE Hybrid

Lead development of HPE AI Essentials’ enterprise LLM inference runtime for customer-owned infrastructure. Improve GPU-efficient serving, distributed execution, and Kubernetes orchestration.

Description

  • Own core components of the LLM serving deployment, including engine integration, continuous batching, KV-cache management and reuse, and quantized execution.
  • Work with inference engineering teams to reduce time to first token, inter-token latency, GPU throughput, and P95/P99 tail latency.
  • Develop and operate distributed execution features such as disaggregated prefill and decode, tensor and pipeline parallelism, and KV-cache offload across GPU memory, host memory, and RDMA-attached storage.
  • Assess emerging runtimes, quantization methods, speculative decoding, and mixture-of-experts serving, then recommend technologies for adoption.
  • Help build the orchestration layer covering model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling.
  • Resolve customer issues end to end, determine root causes, and strengthen related systems and processes.
  • Review code and designs, mentor colleagues, and model strong engineering practices.

Requirements

  • At least 8 years of software engineering experience.
  • One to two or more years of direct experience with LLM inference runtimes or production model serving.
  • A degree in computer science or a related discipline.
  • Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including experience modifying engine internals.
  • Strong knowledge of continuous batching, paged attention, KV-cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
  • Working knowledge of tensor and pipeline parallelism, NCCL collective operations, GPU memory hierarchy, and interconnect behavior.
  • Advanced proficiency with Kubernetes platform architecture, including operators, custom resources, controllers, and scheduling.
  • Strong programming skills in Go and Python.
  • Ability to read, debug, and profile C++ and CUDA with tools such as Nsight.
  • Familiarity with debugging and profiling multi-tier workloads, including RAG and agent applications.
  • Excellent analytical, debugging, and problem-solving skills.
  • Preferred: Contributions to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe.
  • Preferred: Experience with disaggregated prefill and decode serving or large-scale KV-cache offload and reuse.
  • Preferred: Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE.
  • Preferred: Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation.
  • Preferred: Experience delivering software for on-premises, air-gapped, or regulated enterprise environments.

Benefits

  • Comprehensive benefits supporting physical health, financial security, and emotional wellbeing.
  • Programs supporting personal and professional growth.
  • Flexibility for work arrangements and personal needs.
  • An inclusive workplace culture.
  • Reasonable accommodations during the application or interview process for qualified applicants with disabilities.

Related Jobs