AI Performance Engineering Intern

Hewlett Packard Enterprise•Spring, TX
•$35 - $41•Onsite

About The Position

This role has been designed as ‘Onsite’ with an expectation that you will primarily work from an HPE office. Who We Are: Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you. Open up opportunities with HPE.

Requirements

  • Working toward B.S. in Computer Science, Applied Mathematics, Software Engineering or equivalent
  • Coursework and experience with GPU (NVIDIA/AMD) hardware architectures and parallel programming techniques, including parallel programming methods such as NCCL/RCCL, NVSHMEM, or OpenSHMEM
  • Knowledge and Skills:
  • Background with AI applications and their programming models/languages, including PyTorch/TensorFlow
  • Proficiency in Python
  • Deep understanding of LLM parallelization methods TP/PP/DP, and especially EP (MoE) is a plus
  • Experience with using inference serving engines (e.g. vLLM, Ollama) is a plus
  • Experience with AI clusters using job schedulers (e.g., PBS, Slurm) is a plus
  • Experience on GPU-based AI clusters with Docker or Apptainer Containers is a plus
  • Experience with GPU profiling using NVIDIA Nsight or AMD rocprof is a plus
  • Proficiency with CUDA/ROCm is a plus
  • Proficiency with the Linux operating systems, scripting, and knowledge of performance monitoring tools
  • Experience with AI industry standard benchmarks (MLPerf) is a plus
  • Self-driven, solid work ethic, and motivation
  • Ability to multi-task with strong organizational skills and attention to detail
  • Excellent analytical and problem-solving skills
  • Integrity, customer focus, innovation, teamwork, and accountability in daily work
  • Ability to work in a fast-paced, multitasking environment
  • Excellent written and verbal communication skills. Mastery of the English language is required.

Nice To Haves

  • Deep understanding of LLM parallelization methods TP/PP/DP, and especially EP (MoE) is a plus
  • Experience with using inference serving engines (e.g. vLLM, Ollama) is a plus
  • Experience with AI clusters using job schedulers (e.g., PBS, Slurm) is a plus
  • Experience on GPU-based AI clusters with Docker or Apptainer Containers is a plus
  • Experience with GPU profiling using NVIDIA Nsight or AMD rocprof is a plus
  • Proficiency with CUDA/ROCm is a plus
  • Experience with AI industry standard benchmarks (MLPerf) is a plus

Responsibilities

  • Work with team members to design and create test plans to understand an AI system's performance.
  • Create and execute experiments to understand the performance characteristics of the system under test (SUT).
  • Use profiling and/or monitoring tools to assist with the performance characterization of the SUT.
  • Collaborate with team members regarding status, project progress, and issue resolution.

Benefits

  • Health & Wellbeing
  • Personal & Professional Development
  • Unconditional Inclusion
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service