Senior Software Engineer, Inference

Hewlett Packard EnterpriseDurham, NC
$137,000 - $315,000Hybrid

About The Position

Hewlett Packard Enterprise (HPE) is seeking a Senior Software Engineer for its Private Cloud AI organization. This role focuses on building and evolving the model runtime within HPE AI Essentials, an inference platform designed for enterprises to operate large language models (LLMs) on their own infrastructure, including air-gapped and sovereign environments. The primary engineering challenge is ensuring sustained execution efficiency, characterized by low tail latency and high GPU utilization on diverse customer-owned hardware. The engineer will design and implement key components of the runtime, such as engine integration, batching, KV cache management, and distributed execution, all supported by a Kubernetes orchestration layer. While the primary work location is in the US, remote options will be considered.

Requirements

  • Familiar with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals.
  • Strong understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding.
  • Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them.
  • Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling.
  • Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight.
  • Familiar with debugging/profiling multi-tier application workloads such as RAG, Agents.
  • Excellent analytical, debugging, and problem-solving abilities.
  • Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving.
  • Degree in Computer Science or related field.

Nice To Haves

  • Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe.
  • Disaggregated prefill/decode serving, or KV cache offload and reuse at scale.
  • RDMA, GPUDirect Storage, InfiniBand, or RoCE.
  • MIG, fractional GPU allocation, and multi-tenant GPU isolation.
  • On-premises, air-gapped, or regulated enterprise software delivery.

Responsibilities

  • Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution.
  • Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency.
  • Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage.
  • Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption.
  • Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling.
  • Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence.
  • Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team.

Benefits

  • Health & Wellbeing comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.
  • Personal & Professional Development programs catered to helping you reach any career goals.
  • Unconditional Inclusion, flexibility to manage work and personal needs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service