Member of Technical Staff - ML Infra

Causal LabsSan Francisco, CA

About The Position

This role involves designing, deploying, and maintaining large distributed ML training and inference clusters. The position requires developing efficient, scalable end-to-end pipelines for managing petabyte-scale datasets and model training throughout the entire ML lifecycle. Responsibilities also include researching and testing various training approaches, analyzing and debugging low-level GPU operations for performance optimization, and staying current with research to introduce new ideas.

Requirements

  • Relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains
  • Strong grasp of state-of-the-art techniques for optimizing training and inference workloads
  • Demonstrated proficiency with distributed training frameworks (e.g. FSDP, DeepSpeed) to train large foundation models
  • Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
  • Familiarity with containerization and orchestration frameworks (e.g., Kubernetes, Docker)
  • Background working on distributed task management systems and scalable model serving & deployment architectures
  • Understanding of monitoring, logging, observability, and version control best practices for ML systems

Responsibilities

  • Design, deploy, and maintain large distributed ML training and inference clusters
  • Develop efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and model training throughout the entire ML lifecycle
  • Research and test various training approaches including parallelization techniques and numerical precision trade-offs across different model scales
  • Analyze, profile and debug low-level GPU operations to optimize performance
  • Stay up-to-date on research to bring new ideas to work
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service