Head of Infra, Enterprise Governance, Reliability

AnyscaleSan Francisco, CA
$314,000 - $372,000

About The Position

Anyscale is looking for an experienced Engineering leader to lead our Infrastructure, SRE and Enterprise Governance Engineering teams. Anyscale aims to provide the next generation of tools and infrastructure to make developing and running distributed AI applications in the cloud using Ray - the popular open source platform used by companies like Netflix, Uber, Instacart and others - seamless. In this position, you will guide the vision, technical direction of the team, and recruit, enable a high-performing engineering team that delivers critical values to developers and Anyscale customers by solving complex distributed systems challenges. You will oversee and drive the strategy and execution of components which includes cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability, billing stack, production database and related components. You will closely work with our customers and our field engineering team to solve their problems, understand their challenges and make sure they are successful.

Requirements

  • Solid engineering management experience leading productive, building high-performing teams.
  • Record of helping teams scale quickly while maintaining a good culture.
  • Ability to ensure a high hiring bar, motivate team, coach/mentor, and handle performance management issues.
  • Deep technical knowledge and experience in distributed systems.
  • Prior experience working on Kubernetes, VMs.
  • A great track record of execution.
  • A sense of urgency, a mindset towards achieving results, and excellent prioritization skills.
  • Effective communication.

Responsibilities

  • Guide the vision and technical direction of the Infrastructure, SRE and Enterprise Governance Engineering teams.
  • Recruit and enable a high-performing engineering team.
  • Deliver critical values to developers and Anyscale customers by solving complex distributed systems challenges.
  • Oversee and drive the strategy and execution of components including cluster launcher, cloud providers (AWS/GCP/Azure/etc.), Kubernetes support, cluster autoscaling, control plane, data plane, reliability, billing stack, production database and related components.
  • Closely work with customers and the field engineering team to solve their problems, understand their challenges and ensure their success.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service