Research Engineer, ML Platform

MistralPalo Alto, CA

About The Position

This role focuses on building and operating the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. You will develop the infrastructure that enables researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions. You will work across the full ML lifecycle, from workload scheduling and capacity management to platform APIs, observability, and production operations. You will take ownership of critical systems and help turn complex infrastructure into reliable, self-service capabilities.

Requirements

  • 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.
  • Proficient in Python or Go and comfortable working with production-grade distributed systems.
  • Strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.
  • Understand technologies such as Kueue, Karpenter, Volcano, and Kyverno, and the problems they address in workload scheduling, provisioning, and policy enforcement.
  • Understand distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference.
  • Familiar with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking.
  • Understand concepts such as quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission.
  • Can diagnose performance and reliability problems across software, orchestration, networking, storage, and hardware.
  • Care about developer experience and enjoy turning complex infrastructure into simple, reliable interfaces.
  • Thrive in an ambiguous, fast-moving environment shaped by frontier AI research.

Responsibilities

  • Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.
  • Build systems for queueing, admission control, quotas, priorities, preemption, and topology-aware placement.
  • Improve how heterogeneous GPU resources are provisioned, allocated, and utilized across clusters.
  • Place workloads based on capacity, data locality, hardware requirements, and organizational priorities.
  • Create self-service workflows that make distributed workloads easy to launch, observe, debug, and reproduce.
  • Improve GPU utilization, scheduling latency, workload startup time, throughput, and infrastructure efficiency.
  • Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads.
  • Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure.

Benefits

  • healthcare coverage
  • parental leave
  • retirement plans
  • relocation support
  • wellness programs
  • meal and transportation allowances
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service