Technology Enablement Engineer

KLAAnn Arbor, MI
$129,600 - $190,067Onsite

About The Position

To make electronics, you need chips, wafers, transistors, reticles, and... To make these, you must see, test and manufacture them at scale—faster and better than ever before. That's where KLA comes in. Whether you're early in your career or an experienced professional, you'll solve complex challenges, work alongside brilliant minds and help shape the future of technology. KLA's IT group supports business growth and productivity by connecting people, process and technology around the world. We work to improve the technology that drives our business to thrive and focus on empowering employee use of technology. This integrated approach to customer service, creativity and technological excellence enhances employee productivity, business analytics and process excellence. In this role, you will play a key part in advancing business priorities by delivering high-impact work across your area of expertise.

Requirements

  • Bachelor's Degree and eight (8) years of Software Engineering experience
  • Four (4) years in software, cloud, platform, HPC, or infrastructure engineering, including two (2) years supporting distributed AI/ML workloads.
  • Proven hands-on experience building GPU clusters from the ground up and operating them at production scale.
  • Deep Kubernetes expertise, including operators, CRDs, Helm, networking, storage, scheduling, and cluster lifecycle management.
  • Production experience with Ray and at least two of the following: NVIDIA NeMo RL, vLLM, SGLang, or NVIDIA Dynamo.
  • Strong knowledge of NVIDIA GPUs, CUDA, NCCL, GPU Operator, DCGM, MIG, RDMA, and multi-node collective communication.
  • Experience debugging training and inference performance across GPU, CPU, memory, network, and storage layers.
  • Proficiency in Python and Linux, plus experience with containers, CI/CD, GitOps, infrastructure-as-code, and platform observability.
  • Demonstrated open-source contribution, maintainership, or meaningful participation in an AI infrastructure project.

Nice To Haves

  • Experience with Google TPUs and TPU-oriented frameworks or distributed workloads.
  • Experience with large language model training, fine-tuning, RLHF or agentic reinforcement learning.
  • Knowledge of TensorRT-LLM, Triton Inference Server, DeepSpeed, Megatron-LM, or similar performance-oriented frameworks.
  • Experience operating secure, multi-tenant AI platforms in enterprise or regulated environments.

Responsibilities

  • Design and deploy scalable, multi-node GPU clusters on Kubernetes, including compute, networking, storage, scheduling, security, and observability.
  • Build distributed training and reinforcement learning platforms using Ray and NVIDIA NeMo RL, supporting frameworks such as PyTorch and JAX.
  • Deploy and optimize high-throughput LLM inference using vLLM, SGLang, and NVIDIA Dynamo.
  • Implement GPU scheduling, quotas, isolation, autoscaling, health monitoring, and capacity management for multi-tenant environments.
  • Profile and troubleshoot GPU workloads using NVIDIA Nsight Systems, Nsight Compute, DCGM, CUDA, NCCL, and related diagnostics.
  • Automate cluster provisioning, upgrades, workload deployment, and operational recovery through infrastructure-as-code and GitOps practices.
  • Contribute to or actively participate in relevant open-source AI infrastructure communities and bring upstream best practices into the platform.

Benefits

  • medical
  • dental
  • vision
  • life
  • 401(K) including company matching
  • employee stock purchase program (ESPP)
  • student debt assistance
  • tuition reimbursement program
  • development and career growth opportunities and programs
  • financial planning benefits
  • wellness benefits including an employee assistance program (EAP)
  • paid time off
  • paid company holidays
  • family care and bonding leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service