Hewlett Packard Enterprise
Hewlett Packard Enterprise
Hewlett Packard Enterprise gibrid bulut va sun’iy intellektga asoslangan operatsiyalarni yurituvchi tashkilotlar uchun korporativ texnologiyalarni ishlab chiqadi. Uning portfeli HPE ProLiant serverlari va HPE Cray yuqori unumli hisoblash tizimlari, HPE Alletra ma’lumot saqlash yechimlari, HPE Aruba tarmoq texnologiyalari, sun’iy intellekt platformalari va tayyor AI fabrikalari, shuningdek xizmat sifatida taqdim etiladigan GreenLake yechimlarini qamrab oladi. HPE, shuningdek, yirik tashkilotlarga infratuzilmasini modernizatsiya qilish, ma’lumotlar markazi muhitlarini boshqarish va sun’iy intellektni keng miqyosda joriy etishda yordam beradigan Zero Trust va SASE kabi xavfsizlik imkoniyatlari bilan bir qatorda professional, maslahat va qo‘llab-quvvatlash xizmatlarini ham taqdim etadi.

Senior Inference Software Engineer at HPE Hybrid

Lead development of HPE AI Essentials’ enterprise LLM inference runtime for customer-owned infrastructure. Improve GPU-efficient serving, distributed execution, and Kubernetes orchestration.

Tavsif

  • Own core components of the LLM serving deployment, including engine integration, continuous batching, KV-cache management and reuse, and quantized execution.
  • Work with inference engineering teams to reduce time to first token, inter-token latency, GPU throughput, and P95/P99 tail latency.
  • Develop and operate distributed execution features such as disaggregated prefill and decode, tensor and pipeline parallelism, and KV-cache offload across GPU memory, host memory, and RDMA-attached storage.
  • Assess emerging runtimes, quantization methods, speculative decoding, and mixture-of-experts serving, then recommend technologies for adoption.
  • Help build the orchestration layer covering model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling.
  • Resolve customer issues end to end, determine root causes, and strengthen related systems and processes.
  • Review code and designs, mentor colleagues, and model strong engineering practices.

Talablar

  • At least 8 years of software engineering experience.
  • One to two or more years of direct experience with LLM inference runtimes or production model serving.
  • A degree in computer science or a related discipline.
  • Familiarity with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including experience modifying engine internals.
  • Strong knowledge of continuous batching, paged attention, KV-cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
  • Working knowledge of tensor and pipeline parallelism, NCCL collective operations, GPU memory hierarchy, and interconnect behavior.
  • Advanced proficiency with Kubernetes platform architecture, including operators, custom resources, controllers, and scheduling.
  • Strong programming skills in Go and Python.
  • Ability to read, debug, and profile C++ and CUDA with tools such as Nsight.
  • Familiarity with debugging and profiling multi-tier workloads, including RAG and agent applications.
  • Excellent analytical, debugging, and problem-solving skills.
  • Preferred: Contributions to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe.
  • Preferred: Experience with disaggregated prefill and decode serving or large-scale KV-cache offload and reuse.
  • Preferred: Experience with RDMA, GPUDirect Storage, InfiniBand, or RoCE.
  • Preferred: Experience with MIG, fractional GPU allocation, and multi-tenant GPU isolation.
  • Preferred: Experience delivering software for on-premises, air-gapped, or regulated enterprise environments.

Imtiyozlar

  • Comprehensive benefits supporting physical health, financial security, and emotional wellbeing.
  • Programs supporting personal and professional growth.
  • Flexibility for work arrangements and personal needs.
  • An inclusive workplace culture.
  • Reasonable accommodations during the application or interview process for qualified applicants with disabilities.

O‘xshash ish o‘rinlari

Walmart

Principal Software Engineer, Cloud and Distributed Systems

Walmart

Lead Walmart’s cloud-native, AI-enabled mGPS/Compass platform, shaping distributed services and indoor routing. Improve in-store operations through resilient architecture and technical leadership.

Ochish
Cartonplast Group

Senior ERP Software Development Team Lead – Cartonplast, Hybrid Germany

Cartonplast Group

Lead Cartonplast’s ERP development team and shape commercial, logistics, and production workflows. The role combines team leadership, business coordination, and hands-on technical input in a hybrid setup in Germany.

Ochish
Walmart

Staff Software Engineer, Cloud and AI Platforms — Walmart

Walmart

Lead the development of cloud-native infrastructure, engineering tools, and AI/ML platforms for Walmart’s critical retail systems. Guide architecture and delivery across the software lifecycle while advancing engineering practices and mentoring engineers.

Ochish
Northrop Grumman

Senior Global Trade Technical Advisor, Space Sector

Northrop Grumman

Lead export classification, authorization, and training programs for Northrop Grumman’s Space Sector. Guide trade compliance across aerospace and defense systems.

Ochish
sifamo GmbH

Software Developer, Angular and Spring Boot AI – Hybrid Germany

sifamo GmbH

Build Angular and Spring Boot applications at sifamo, including AI-powered solutions. Work with cloud platforms, Kubernetes and cognitive technologies.

Ochish