Head of Cloud Infrastructure

YipitData
$200,000 - $250,000Remote

About The Position

YipitData is seeking a Head of Cloud Infrastructure (Director Level) to lead the modernization and scaling of its shared infrastructure. This role is responsible for cloud platform, Site Reliability Engineering (SRE), DevOps, and developer experience within an AWS environment utilizing Kubernetes and Terraform. It's a player-coach leadership position requiring strategic thinking, architectural decision-making, and technical credibility to lead a global team. The Head of Cloud Infrastructure will unify modernization efforts, establish reliability and disaster-recovery practices, and enhance the ease with which engineers can build, deploy, and operate software. The role also involves improving cloud efficiency and achieving cost savings without compromising reliability or developer velocity. While this role owns shared infrastructure, Data Platform, Data Engineering, and application teams will continue to own their respective applications and data products. Security and Enterprise IT are outside the scope of this role, with close collaboration expected with those functions, the Chief Architect, and engineering leaders. This is a remote-friendly opportunity within the US, with expected US working hours and sufficient overlap for global team collaboration. No travel is required.

Requirements

  • 6+ years of relevant infrastructure, SRE, platform engineering, or developer productivity experience, applied flexibly based on demonstrated scope and impact, with hands-on expertise in AWS, Kubernetes, and Terraform.
  • Led approximately 3–10 infrastructure, platform, SRE, or DevOps engineers and are excited to remain a player-coach within a collaborative, high-leverage team.
  • Both built or materially transformed a platform function and operated mature, business-critical production infrastructure.
  • Helped a data-intensive SaaS, analytics, or similarly complex technology organization scale through rapid growth.
  • Owned formal reliability practices, including SLOs, incident management, production operations, and disaster recovery, and can design a sustainable on-call model.
  • Improved developer velocity or platform experience through self-service capabilities, standardized workflows, internal platform adoption, or measurable reductions in engineering friction.
  • Can translate technical strategy into an actionable roadmap, make thoughtful tradeoffs, and influence the Chief Architect, engineering leaders, and partner teams through clear communication and credible technical judgment.
  • Delivered measurable cloud-cost or infrastructure-efficiency improvements and understand how to optimize technology costs beyond downsizing through smart architecture across services.
  • Hold a bachelor’s degree or have equivalent practical experience.

Responsibilities

  • Set and execute a prioritized cloud-infrastructure modernization roadmap, bringing existing initiatives together into a clear strategy for AWS, Kubernetes, Terraform, and the broader platform ecosystem.
  • Lead a small, high-leverage global team spanning cloud platform, SRE, and DevOps, combining clear direction and coaching with hands-on technical judgment.
  • Improve developer velocity and platform experience by delivering reliable paved roads, reusable infrastructure capabilities, self-service workflows, and faster feedback loops.
  • Establish and mature formal reliability practices, including service ownership, SLIs and SLOs, incident management, error-budget practices, observability, and accountable post-incident follow-through.
  • Own disaster-recovery strategy and readiness, including recovery objectives, testing, automation, and the closure of identified resilience gaps.
  • Design a sustainable 24/7 operational model, including on-call responsibilities, escalation paths, incident leadership, and healthy collaboration across a global team.
  • Partner with the Chief Architect and engineering leaders to define platform priorities, clarify ownership boundaries, and help application and data teams operate their workloads effectively on shared infrastructure.
  • Define baselines and measure progress across delivery performance, developer experience, platform adoption, reliability, operational health, and cloud unit economics.
  • Improve cloud efficiency through better cost attribution, capacity planning, Kubernetes resource utilization, automation, and measurable savings—without sacrificing reliability or delivery speed.

Benefits

  • flexible work hours
  • flexible vacation
  • generous 401K match
  • parental leave
  • team events
  • wellness budget
  • learning reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service