Senior Cloud DevOps Engineer

AAACosta Mesa, CA
$132,600 - $176,900Hybrid

About The Position

This role specifically focuses on AI/ML enablement. The Senior Cloud DevOps Engineer will be responsible for designing, provisioning, and maintaining resilient, scalable multi-cloud infrastructure across AWS and GCP tailored for AI/ML workloads. The engineer will automate infrastructure provisioning, build end-to-end MLOps pipelines, enable and support modern data science platforms, and manage Kubernetes clusters. This is a hybrid role requiring the candidate to reside within a 50-mile radius of the Costa Mesa, CA office and be available to work on-site 3 days a week.

Requirements

  • Deep, production-level expertise in either AWS or GCP, with working knowledge of or high adaptability to the other.
  • AWS core services: EKS, S3, IAM, EC2, ECR, SageMaker AI, CloudWatch, Redshift, Lambda.
  • GCP core services: GKE, GCS, Cloud IAM, Artifact Registry, Vertex AI, Cloud Logging/Monitoring, BigQuery, Cloud Functions.
  • Advanced Docker containerization skills, including multi-stage builds and image optimization for heavy ML runtimes (PyTorch/TensorFlow).
  • Experience with building deployment pipelines using GitHub Actions for CI/CD.
  • Experience with IaC tools (CloudFormation/Terraform).
  • Solid understanding of cloud and container networking, routing, and protocols such as TCP/IP, TLS, HTTP/2, and data security best practices such as least privilege access management.
  • Strong scripting and programming proficiency in Python, Bash, SQL.
  • Familiarity with configuring Kubernetes (EKS or GKE) for ML workloads.
  • Experience supporting enterprise data/ML platforms and a willingness to learn and evaluate new vendor capabilities, such as Snowflake Cortex AI and Snowflake ML.
  • Interest and willingness to research, evaluate, and operationalize emerging AI/ML and cloud technologies.

Nice To Haves

  • Cloud Certification in DevOps in AWS or Google Cloud is a plus.

Responsibilities

  • Design, provision, and maintain resilient, scalable multi-cloud infrastructure across AWS and GCP tailored for AI/ML workloads.
  • Automate end-to-end cloud infrastructure and environment provisioning using Terraform and AWS CloudFormation.
  • Design and implement robust CI/CD/CT pipelines for automated model training, testing, validation, artifact packaging, deployment, monitoring and observability.
  • Enable and support modern data science platforms (e.g. Amazon SageMaker Unified Studio) and emerging platforms (Snowflake Cortex AI, and Snowflake ML).
  • Deploy and manage production-grade EKS or GKE clusters, optimizing node pools, autoscaling, and compute/GPU resources for ML training and real-time inference.
  • Partner with Information Security, Application Security, Data Center Services, and Cloud Infrastructure Services teams to enforce security, governance, and architectural best practices.
  • Partner closely with data scientists to architect and deploy flagship AI/ML products, establishing reusable blueprints, golden paths, and best practice patterns for enterprise-wide adoption.

Benefits

  • Health coverage for medical, dental, vision
  • 401(K) saving plans with company match AND Pension
  • Tuition assistance
  • Floating holidays and PTO for community volunteer programs
  • Paid parental leave
  • Wellness programs
  • Employee discounts (membership, insurance, travel, entertainment, services and more!)
  • Incentive program based upon the achievement of organization, team and personal performance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service