Hewlett Packard Enterprise (HPE) is seeking a Senior Software Engineer for its Private Cloud AI organization. This role focuses on building and evolving the model runtime within HPE AI Essentials, an inference platform designed for enterprises to operate large language models (LLMs) on their own infrastructure, including air-gapped and sovereign environments. The primary engineering challenge is ensuring sustained execution efficiency, characterized by low tail latency and high GPU utilization on diverse customer-owned hardware. The engineer will design and implement key components of the runtime, such as engine integration, batching, KV cache management, and distributed execution, all supported by a Kubernetes orchestration layer. While the primary work location is in the US, remote options will be considered.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior