About The Position

We are sharing a specialised full-time consulting opportunity for US-based MLOps and ML systems engineers with production experience in JAX, PyTorch, distributed training infrastructure, and custom GPU kernel development using Pallas or Triton. This role supports a high-impact generative AI initiative focused on developing and evaluating advanced ML infrastructure tasks for frontier model training. Selected engineers will design technically challenging problems, produce rigorous solutions, assess model-generated outputs, and help establish evaluation standards across training pipelines, distributed systems, framework-level optimisation, and GPU kernel performance.

Requirements

  • At least 2 years of dedicated professional experience in MLOps, ML infrastructure, or ML systems engineering
  • Production experience with JAX, PyTorch, or both at meaningful scale
  • Hands-on experience writing or optimising custom GPU kernels using Pallas or Triton
  • Strong knowledge of model-training pipelines, distributed systems, accelerators, and performance optimisation
  • Experience working within a recognised technology, AI research, or high-performance engineering organisation
  • Demonstrable professional growth and increasing technical responsibility
  • Strong written communication and the ability to explain complex engineering decisions clearly
  • Reliable availability for a full-time, 40-hour weekday schedule

Nice To Haves

  • Experience supporting large language model or generative AI training environments
  • Familiarity with distributed training frameworks, accelerator orchestration, and multi-host systems
  • Knowledge of XLA, CUDA, compiler optimisation, or low-level performance engineering
  • Experience benchmarking GPU workloads and diagnosing training-performance bottlenecks
  • Familiarity with model-evaluation pipelines, technical annotation, or structured training-data development
  • Previous involvement in technical review, engineering mentorship, or rubric development
  • Experience collaborating with research scientists and infrastructure engineering teams

Responsibilities

  • Analyse and improve machine learning training infrastructure, deployment workflows, and model-development systems
  • Guide research and engineering teams on MLOps, distributed training, and ML framework-level challenges
  • Evaluate training-pipeline architecture, scalability, reliability, and performance
  • Identify technical gaps affecting model training, experimentation, and infrastructure efficiency
  • Design challenging tasks covering MLOps, ML systems, training infrastructure, and framework-level engineering
  • Write accurate, technically rigorous, and well-structured solutions
  • Develop realistic scenarios involving distributed systems, accelerator utilisation, and production ML workflows
  • Ensure tasks reflect practical engineering challenges encountered in advanced AI environments
  • Evaluate technical work involving JAX and PyTorch at production scale
  • Review custom GPU kernels written or optimised using Pallas or Triton
  • Assess kernel correctness, memory access patterns, computational efficiency, and hardware utilisation
  • Analyse framework-level implementation choices and identify opportunities for performance improvement
  • Compare alternative technical solutions and determine which approach is more accurate and effective
  • Provide clear written feedback on correctness, system design, scalability, and optimisation quality
  • Develop detailed rubrics for evaluating training pipelines, distributed systems reasoning, and kernel-level implementations
  • Collaborate with other technical specialists to maintain consistency across evaluation standards and training data

Benefits

  • Competitive hourly compensation
  • Full-time W-2 contingent employment arrangement
  • Fully remote role
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service