ML Infrastructure Engineer

Recruiting From ScratchSan Francisco, CA
Onsite

About The Position

As an ML Infrastructure Engineer, you'll own critical training and inference infrastructure powering production machine learning systems. You'll work closely with ML researchers and product engineers to ensure models are trained, deployed, monitored, and served efficiently while continuously improving latency, reliability, scalability, and cost across production AI workloads. This is an exceptional opportunity to join a fast-growing AI-native company where ML Infrastructure Engineers build the foundational systems enabling next-generation AI products at massive production scale.

Requirements

  • 3+ years of ML Infrastructure Engineering experience
  • Experience building production ML model serving and inference infrastructure
  • Experience working at AI-native startups or organizations training production ML models
  • Experience deploying and maintaining large-scale production ML systems
  • Experience supporting ML researchers with production infrastructure
  • Strong startup ownership mentality with demonstrated engineering execution
  • Experience building highly scalable infrastructure supporting AI products
  • Experience collaborating closely with ML, Product, and Infrastructure teams
  • Experience optimizing production inference systems
  • Strong software engineering fundamentals across distributed systems and infrastructure
  • Expert-level Python programming skills
  • Strong Kubernetes, Docker, Helm, and cloud infrastructure experience
  • Deep understanding of GPU optimization, model serving, and inference infrastructure
  • Experience with PyTorch and modern ML deployment workflows
  • Experience building observability, monitoring, and logging systems for ML infrastructure
  • Strong understanding of model deployment, training pipelines, and production inference
  • Experience working with 1–3 node model training and single/double-node serving
  • Familiarity with cloud platforms including AWS, Azure, or GCP
  • Ability to rapidly debug production infrastructure and performance bottlenecks
  • Bachelor's degree in Computer Science, Machine Learning, Engineering, Mathematics, or another technical discipline preferred
  • Strong engineering background with demonstrated infrastructure expertise

Nice To Haves

  • Previous experience working with GPU-intensive workloads strongly preferred
  • Top technical universities, competitive programming backgrounds, or exceptional engineering accomplishments strongly preferred
  • Publications at top ML conferences (NeurIPS, ICML, ICLR, CVPR) are a strong plus

Responsibilities

  • Build and maintain production ML model serving infrastructure powering millions of document processing requests
  • Design and improve training infrastructure supporting models ranging from hundreds of millions to tens of billions of parameters
  • Optimize inference latency, throughput, reliability, and GPU utilization across production systems
  • Develop observability, monitoring, logging, and alerting across the ML infrastructure stack
  • Build internal tooling and data pipelines enabling ML researchers to rapidly deploy models into production
  • Architect infrastructure that intelligently routes inference workloads across multiple cloud providers
  • Optimize infrastructure for model accuracy, latency, reliability, and operational cost
  • Collaborate closely with ML researchers, software engineers, and product teams to accelerate experimentation
  • Improve Kubernetes-based infrastructure supporting large-scale ML workloads
  • Debug complex production infrastructure issues involving GPUs, distributed systems, and model serving
  • Design scalable systems supporting rapid AI model deployment and iteration
  • Champion engineering excellence across ML infrastructure, automation, and production reliability

Benefits

  • Competitive Equity Package
  • Visa Sponsorship Available
  • On-site work model (5 days/week in San Francisco)
  • Opportunity to build core infrastructure powering one of the fastest-growing AI platforms
  • Direct collaboration with ML researchers and engineering leadership
  • Significant ownership across production AI infrastructure
  • High-impact engineering culture with exceptional career growth
  • Well-funded Series B company backed by leading venture firms
  • Opportunity to build foundational infrastructure enabling next-generation AI systems
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service