Research Engineer - ML Infrastructure

Chai Discovery•San Francisco, CA

About The Position

Research Engineers on ML Infra make our models train and run performantly, reliably, and at scale by owning the distributed systems that sit underneath every model our researchers ship. As a ML Infra Research Engineer, you will: Architect, debug, and optimize the distributed ML training stack across the model, layer, and kernel levels — eliminating runtime and reliability bottlenecks on large GPU clusters. Profile end-to-end training runs to find bottlenecks across compute, communication, and storage, and build tooling to monitor throughput, utilization, and uptime across clusters. Optimize ML workloads through parallelism strategies, quantization, and custom CUDA/Triton kernels. Work closely with Research Scientists to ensure new model architectures and training recipes scale efficiently, from early experiments to frontier-scale runs. Own reliability of the training stack: fault tolerance, checkpointing, and deterministic orchestration for long-running, large-scale jobs.

Requirements

  • 4+ years of industry experience working within AI/ML infrastructure teams.
  • Proficiency in Python and PyTorch or JAX.
  • Strong software systems design skills, with comfort operating across the stack from model code down to kernels.
  • Experience with orchestrating GPU clusters and large-scale model training.
  • Experience with optimizing ML workloads: parallelism, quantization, CUDA/Triton kernels.

Responsibilities

  • Architect, debug, and optimize the distributed ML training stack across the model, layer, and kernel levels — eliminating runtime and reliability bottlenecks on large GPU clusters.
  • Profile end-to-end training runs to find bottlenecks across compute, communication, and storage, and build tooling to monitor throughput, utilization, and uptime across clusters.
  • Optimize ML workloads through parallelism strategies, quantization, and custom CUDA/Triton kernels.
  • Work closely with Research Scientists to ensure new model architectures and training recipes scale efficiently, from early experiments to frontier-scale runs.
  • Own reliability of the training stack: fault tolerance, checkpointing, and deterministic orchestration for long-running, large-scale jobs.

Benefits

  • The opportunity to work at the vanguard of AI research and frontier biology, with world-class people, on a mission that matters.
  • We protect & promote a culture of high velocity and ownership.
  • We compensate our team accordingly.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service