AI & HPC Infrastructure Engineer

Saxon GlobalSan Jose, CA

About The Position

We are looking for an AI & HPC Infrastructure Engineer to design, build, and operate the compute platforms that power our AI and High-Performance Computing (HPC) workloads. In this role, you will deploy, automate, and manage GPU-enabled infrastructure across on-premises and cloud environments, enabling scalable, reliable, and cost-effective compute resources for engineering and R&D teams. You will work at the intersection of AI infrastructure, cloud platforms, automation, and operations to support next-generation AI and engineering applications.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical discipline.
  • 5+ years of hands-on experience in Infrastructure Engineering, Platform Engineering, Cloud Operations, or HPC environments.
  • Strong experience with Linux-based infrastructure administration.
  • Experience working with public cloud platforms such as AWS, Azure, or GCP.
  • Hands-on experience with GPU/HPC environments and workload orchestration platforms such as Kubernetes or Slurm.
  • Experience with automation, infrastructure provisioning, monitoring, and performance optimization.
  • Solid understanding of compute, storage, networking, virtualization, and container technologies.

Nice To Haves

  • Experience supporting AI/ML workloads and GPU-based infrastructure.
  • Knowledge of Kubernetes, Docker, Infrastructure-as-Code tools, and observability platforms.
  • Familiarity with AI platforms, LLM deployment, and modern engineering productivity tools.

Responsibilities

  • Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments.
  • Support capacity planning, performance tuning, and resource optimization for AI training, inference, and compute-intensive workloads.
  • Monitor infrastructure health and ensure high availability and performance.
  • Deploy and manage compute environments across on-premises and public cloud platforms (AWS, Azure, or GCP).
  • Contribute to infrastructure modernization, scalability, and resiliency initiatives.
  • Support cloud adoption and hybrid computing strategies.
  • Implement Infrastructure-as-Code (IaC) and automation frameworks for provisioning and operations.
  • Develop monitoring, logging, and alerting solutions to improve platform reliability.
  • Drive continuous improvements in operational efficiency and resource utilization.
  • Deploy, integrate, and support AI services, including LLM APIs, coding assistants, and AI/agent platforms.
  • Collaborate with engineering teams to enable AI-driven development workflows.
  • Support AI/ML infrastructure requirements and best practices.
  • Troubleshoot and resolve infrastructure, networking, and platform issues.
  • Create and maintain technical documentation, standards, and operational runbooks.
  • Partner with engineering, IT, and platform teams to deliver secure and scalable solutions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service