AI/HPC Systems Engineer

Saige PartnersSan Jose, CA

About The Position

We are seeking an AI/HPC Systems Engineer to build, deploy, and operate the compute infrastructure supporting high-performance computing and AI development workloads. This role will focus on deploying, automating, and maintaining GPU-enabled environments across on-premises and cloud platforms while delivering reliable, scalable, and cost-effective computing resources for engineering and R&D teams. The ideal candidate brings hands-on experience with Linux infrastructure, GPU/HPC environments, cloud platforms, automation, and modern AI/ML infrastructure.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • 3+ years of hands-on experience in IT infrastructure, cloud engineering, platform engineering, HPC, or a related field.
  • Hands-on experience with Linux-based infrastructure and public cloud platforms such as AWS, Azure, or GCP.
  • Experience deploying, configuring, or operating GPU and/or HPC environments.
  • Experience with workload scheduling or orchestration technologies such as Kubernetes, Slurm, or similar platforms.
  • Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization.
  • Strong understanding of compute, storage, networking, virtualization, and container technologies.
  • Strong problem-solving, collaboration, and communication skills.
  • Ability to work effectively across engineering, R&D, and IT teams.

Nice To Haves

  • Experience supporting AI/ML infrastructure or workloads is a plus.

Responsibilities

  • Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments.
  • Support capacity planning, performance tuning, and infrastructure optimization for AI training, inference, and compute-intensive workloads.
  • Monitor system performance, availability, and resource utilization to ensure reliable operations.
  • Deploy and maintain computing environments across on-premises infrastructure and public cloud platforms.
  • Support infrastructure modernization, expansion, and scaling initiatives for HPC and AI workloads.
  • Help evaluate and implement solutions that improve scalability, reliability, and cost efficiency.
  • Implement infrastructure-as-code and automated provisioning solutions.
  • Develop and maintain monitoring, logging, alerting, and observability capabilities.
  • Automate routine infrastructure tasks and identify opportunities to improve resource utilization and operational efficiency.
  • Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by engineering teams.
  • Assist with the infrastructure requirements and operational support of AI/ML workloads.
  • Collaborate with engineering teams to ensure AI platforms are reliable, accessible, and scalable.
  • Troubleshoot and resolve infrastructure, networking, compute, storage, and platform issues.
  • Support day-to-day IT and infrastructure operations across engineering environments.
  • Develop and maintain technical documentation, standards, procedures, and operational runbooks.
  • Collaborate with engineering, IT, and other stakeholders to deliver reliable infrastructure solutions.

Benefits

  • benefit package
  • convenient weekly payment solutions
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service