Senior Azure ML Infrastructure Engineer

Simpson Thacher & Bartlett LLPNew York, NY
$160,000 - $180,000Hybrid

About The Position

Simpson Thacher & Bartlett LLP is looking for a Senior Azure ML Infrastructure Engineer to lead the design, development, and optimization of scalable ML infrastructure on Microsoft Azure. In this role, you will be the technical lead for deploying and maintaining robust (Machine Learning) ML Ops frameworks, ensuring efficient collaboration between data science, engineering, and DevOps teams. You’ll be instrumental in scaling our machine learning capabilities from experimentation to production across multiple use cases.

Requirements

  • 5+ years of experience in ML infrastructure, cloud engineering, or MLOps
  • 2+ years of experience working in Azure environments.
  • Deep hands-on experience with Azure cloud services relevant to ML, including Azure Machine Learning, AKS, Blob Storage, Databricks, Azure Data Factory, and Synapse.
  • Strong expertise in containerization (Docker) and orchestration (Kubernetes, preferably AKS).
  • Proficient in Python and scripting languages (e.g., Bash, PowerShell).
  • Advanced knowledge of CI/CD tools such as Azure DevOps, GitHub Actions, or Jenkins for ML workloads.
  • Solid understanding of IaC tools: Terraform, Bicep, or ARM templates.

Nice To Haves

  • Microsoft Azure certifications (e.g., Azure AI Engineer Associate, Azure Solutions Architect Expert, or DevOps Engineer Expert).
  • Experience designing ML infrastructure in regulated industries (finance, healthcare, etc.).
  • Familiarity with feature stores, distributed training, and model monitoring frameworks.
  • Leadership experience in building infrastructure for ML at scale.
  • Legal IT experience a plus but not required

Responsibilities

  • Lead the architecture and implementation of production-grade ML infrastructure using Azure Machine Learning, AKS, Azure Data Lake, Azure Databricks, and related services.
  • Design scalable training and inference environments for deep learning and traditional ML workloads, optimizing performance and cost.
  • Define and implement MLOps best practices: versioning, CI/CD for ML pipelines, monitoring, and model governance.
  • Automate end-to-end ML workflows using tools such as MLFlow, Azure ML Pipelines, or Kubeflow.
  • Build reusable templates and frameworks to standardize ML deployment across teams.
  • Collaborate with data scientists to productionize models, offering guidance on infrastructure, deployment strategies, and performance optimization.
  • Partner with DevOps and platform engineering teams to align infrastructure with broader cloud strategies and compliance standards.
  • Mentor junior ML and platform engineers, sharing best practices and driving engineering excellence.
  • Implement enterprise-grade security and compliance controls using Azure Active Directory, RBAC, and data encryption strategies.
  • Integrate observability tooling (e.g., Azure Monitor, Prometheus, Grafana) for end-to-end monitoring of ML systems.
  • Ensure systems are highly available, reliable, and scalable to meet the demands of production ML workloads.

Benefits

  • Competitive salary, equity opportunities, and comprehensive benefits.
  • Continuous learning budget and Azure certification support.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service