Sr. Technology Enablement Engineer

KLA•Ann Arbor, MI
•$129,600 - $190,067•Onsite

About The Position

In this role, you will play a key part in advancing business priorities by delivering high-impact work across your area of expertise. We are seeking a highly skilled Sr. Software Engineer with specialized expertise to design, build, and operate production-grade AI infrastructure for large-scale GPU training and inference. The role owns cluster architecture from ground zero, distributed workload orchestration, accelerator management, performance engineering, and open-source integration. TPU experience is optional.

Requirements

  • Bachelor's Degree and eight (8) years of Software Engineering experience
  • Four (4) years in software, cloud, platform, HPC, or systems engineering, including two (2) years supporting distributed AI/ML workloads.
  • Proven hands-on experience building GPU clusters from the ground up and operating them at production scale.
  • Deep Kubernetes expertise, including operators, CRDs, Helm, networking, storage, scheduling, and cluster lifecycle management.
  • Production experience with Ray and at least two of the following: NVIDIA NeMo RL, vLLM, SGLang, or NVIDIA Dynamo.
  • Strong knowledge of NVIDIA GPUs, CUDA, NCCL, GPU Operator, DCGM, MIG, RDMA, and multi-node collective communication.
  • Experience debugging training and inference performance across GPU, CPU, memory, network, and storage layers.
  • Proficiency in Python and Linux, plus experience with containers, CI/CD, GitOps, infrastructure-as-code, and platform observability.
  • Demonstrated open-source contribution, maintainership, or meaningful participation in an AI infrastructure project.

Nice To Haves

  • Experience with Google TPUs and TPU-oriented frameworks or distributed workloads.
  • Experience with large language model training, fine-tuning, RLHF or agentic reinforcement learning.
  • Knowledge of TensorRT-LLM, Triton Inference Server, DeepSpeed, Megatron-LM, or similar performance-oriented frameworks.
  • Experience operating secure, multi-tenant AI platforms in enterprise or regulated environments.

Responsibilities

  • Design and deploy scalable, multi-node GPU clusters on Kubernetes, including compute, networking, storage, scheduling, security, and observability.
  • Build distributed training and reinforcement learning platforms using Ray and NVIDIA NeMo RL, supporting frameworks such as PyTorch and JAX.
  • Deploy and optimize high-throughput LLM inference using vLLM, SGLang, and NVIDIA Dynamo.
  • Implement GPU scheduling, quotas, isolation, autoscaling, health monitoring, and capacity management for multi-tenant environments.
  • Profile and troubleshoot GPU workloads using NVIDIA Nsight Systems, Nsight Compute, DCGM, CUDA, NCCL, and related diagnostics.
  • Automate cluster provisioning, upgrades, workload deployment, and operational recovery through infrastructure-as-code and GitOps practices.
  • Contribute to or actively participate in relevant open-source AI infrastructure communities and bring upstream best practices into the platform.

Benefits

  • medical
  • dental
  • vision
  • life
  • 401(K) including company matching
  • employee stock purchase program (ESPP)
  • student debt assistance
  • tuition reimbursement program
  • development and career growth opportunities and programs
  • financial planning benefits
  • wellness benefits including an employee assistance program (EAP)
  • paid time off
  • paid company holidays
  • family care and bonding leave
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service