HPC Systems Engineer

KLAMilpitas, CA
$136,300 - $231,700Onsite

About The Position

As an AI Infrastructure Engineer, you will research, evaluate, and develop next-generation hardware and software technologies that enable large-scale AI and machine learning workloads across KLA. You will help design and build the infrastructure that powers model training, inference, networking, storage, data protection, and emerging agentic AI systems. This role sits at the intersection of systems engineering, AI platform development, MLOps, and DevOps, ensuring that AI teams have reliable, scalable, high-performance infrastructure to develop, train, deploy, and operate AI solutions. You will also define and develop reference architectures and platform standards that can be embraced across KLA products and engineering organizations, enabling secure, cost-effective, and repeatable AI infrastructure deployments.

Requirements

  • Degree in Computer Science, Computer Engineering, or related field
  • 5–8 years in systems engineering, DevOps, or ML infrastructure
  • Hands‑on experience building AI/GPU cluster
  • Doctorate (Academic) Degree and 0 years related work experience; Master's Level Degree and related work experience of 3 years; Bachelor's Level Degree and related work experience of 5 years
  • Linux administration
  • Kubernetes/Docker
  • Slurm
  • Ray
  • TensorFlow
  • PyTorch
  • GPU/TPU optimization
  • containerization
  • orchestration
  • programming skills
  • High speed networking
  • storage systems
  • virtualization
  • server management
  • Understanding of model training, fine‑tuning, RAG pipelines, and agentic AI systems
  • Bash
  • Python
  • CI/CD pipelines
  • infrastructure as code (IaC)
  • TPM-based encryption
  • Kubernetes security (RBAC, OPA/Gatekeeper)
  • container security

Responsibilities

  • Design & Build AI Infrastructure: Architect, deploy, and operate scalable, secure, and cost-effective AI platforms, distributed training environments, and model serving systems.
  • Research & Innovation: Evaluate emerging technologies in AI infrastructure, compute, storage, networking, orchestration, and security, and develop reference designs for enterprise adoption.
  • Hardware & Software Orchestration: Bring up, integrate, and lead AI compute infrastructure while diagnosing hardware, firmware, operating system, and platform issues.
  • Data & Storage Management: Design and maintain high-performance storage, networking, and data pipelines supporting large-scale AI workloads.
  • Performance & Reliability: Develop observability and monitoring solutions, optimize AI workload performance, and ensure infrastructure meets reliability and availability targets.
  • Security & Compliance: Implement security-by-design principles including encryption, identity and access management, secrets management, and AI data security. Contribute to the architecture and implementation capabilities for protecting AI workloads, intellectual property, and sensitive data at KLA.
  • Collaboration: Partner with AI researchers, data scientists, software engineers, IT, and security teams to align platform capabilities with business and product needs.

Benefits

  • medical
  • dental
  • vision
  • life
  • 401(K) including company matching
  • employee stock purchase program (ESPP)
  • student debt assistance
  • tuition reimbursement program
  • development and career growth opportunities and programs
  • financial planning benefits
  • wellness benefits including an employee assistance program (EAP)
  • paid time off
  • paid company holidays
  • family care and bonding leave

Stand Out From the Crowd

Upload your resume and get instant feedback on how well it matches this job.

Upload and Match Resume

What This Job Offers

Job Type

Full-time

Career Level

Entry Level

Education Level

Ph.D. or professional degree

© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service