About The Position

We are sharing a full-time opportunity for an experienced Director of Infrastructure Engineering with deep expertise in AWS, GCP, infrastructure as code, CI/CD, platform engineering, reliability, security, and technical leadership to build and scale infrastructure supporting production AI systems. The role combines hands-on infrastructure engineering with strategic leadership across cloud architecture, developer platforms, observability, reliability, security, and engineering operations.

Requirements

  • 8+ years of experience in production infrastructure, platform engineering, DevOps, or SRE
  • 3+ years of engineering leadership experience
  • Deep expertise with AWS, GCP, or multi-cloud production environments
  • Strong Terraform or comparable infrastructure-as-code experience
  • Strong knowledge of Kubernetes and containerised infrastructure
  • Experience designing and operating modern CI/CD platforms
  • Demonstrated success building highly available, observable, and resilient systems
  • Strong understanding of infrastructure security, compliance, and operational risk
  • Experience scaling infrastructure and engineering teams in fast-moving environments
  • Excellent written and verbal communication and ability to influence technical strategy

Nice To Haves

  • AI/ML infrastructure or large-scale data-platform experience is highly valuable
  • Familiarity with model training, inference, evaluation, or data-pipeline infrastructure is advantageous
  • Experience with FedRAMP, GovCloud, CMMC Level 2, or comparable regulated environments is beneficial

Responsibilities

  • Own multi-cloud architecture across AWS and GCP
  • Define infrastructure strategy around scalability, reliability, security, and cost
  • Build and maintain infrastructure as code using Terraform or comparable tooling
  • Develop reusable platform abstractions, automation, and internal infrastructure tooling
  • Improve developer productivity while maintaining strong operational standards
  • Design and improve CI/CD systems for fast, reliable, and secure software delivery
  • Establish SLOs, error budgets, incident-response processes, and on-call practices
  • Lead disaster-recovery and resilience initiatives
  • Build observability across metrics, logs, traces, alerting, and operational signals
  • Use production data and postmortems to improve reliability and reduce deployment risk
  • Embed security into cloud architecture, platform tooling, and software-delivery workflows
  • Support compliance with frameworks such as ISO 27001, SOC 2, and CMMC
  • Lead and develop a high-performing Infrastructure or Platform Engineering team
  • Mentor engineers and influence infrastructure strategy across technical and executive stakeholders
  • Balance long-term platform strategy with hands-on production and incident-management responsibilities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service