Azure and Google Cloud Engineer

TekCommandsPlano, TX

About The Position

This role focuses on maintaining the operational health of cloud environments across Azure and Google Cloud, ensuring uptime, performance, and stability. The engineer will be responsible for building, deploying, and operating Kubernetes platforms (AKS & GKE) with zero downtime, and designing and managing multi-cloud infrastructure using Terraform. A key aspect of the role is creating reusable, secure, and scalable Terraform modules and implementing zero-downtime deployment and upgrade strategies. The position also involves designing and managing cloud networking, troubleshooting complex issues, and automating infrastructure and operations using Python, Bash, or PowerShell. Building and maintaining CI/CD pipelines, applying an automation-first approach, and implementing monitoring, logging, alerting, and observability are crucial. The engineer will proactively identify issues, perform root cause analysis, and ensure systems are scalable, resilient, secure, and performant. A deep understanding of platform architecture and dependencies is required, along with driving operational best practices and continuous improvements. Collaboration with engineering and product teams is essential. The role also involves applying LLMs/AI to cloud operations, automation, diagnostics, and observability, and staying current with AI-driven DevOps/SRE tooling. The ideal candidate will be a self-starter, owning initiatives end-to-end, and demonstrating strong leadership, communication, and problem-solving skills.

Requirements

  • Azure and Google Cloud experience
  • Kubernetes platforms (AKS & GKE)
  • Terraform for multi-cloud infrastructure management
  • Python, Bash, or PowerShell for automation
  • CI/CD pipeline development and maintenance
  • Monitoring, logging, alerting, and observability implementation
  • Root cause analysis and troubleshooting of complex issues
  • Understanding of cloud networking (VNETs/VPCs, routing, peering, VPNs, DNS)
  • Knowledge of LLMs/AI in cloud operations
  • Familiarity with AI-driven DevOps/SRE tooling
  • Strong leadership, communication, and problem-solving skills

Responsibilities

  • Maintain cloud operational health across Azure and Google Cloud (uptime, performance, stability)
  • Build, deploy, and operate Kubernetes platforms (AKS & GKE) with zero downtime
  • Design and manage multi-cloud infrastructure using Terraform
  • Create reusable, secure, and scalable Terraform modules
  • Implement zero-downtime deployment and upgrade strategies
  • Design and manage cloud networking (VNETs/VPCs, routing, peering, VPNs, DNS)
  • Troubleshoot complex network, connectivity, and performance issues
  • Automate infrastructure, deployments, and operations using Python, Bash, or PowerShell
  • Build and maintain CI/CD pipelines and operational workflows
  • Apply an automation-first approach to minimize manual effort and risk
  • Implement and manage monitoring, logging, alerting, and observability
  • Proactively identify issues and perform root cause analysis
  • Ensure systems are scalable, resilient, secure, and performant
  • Develop deep understanding of platform architecture and dependencies
  • Drive operational best practices, standards, and continuous improvements
  • Collaborate closely with engineering and product teams
  • Apply LLMs/AI to cloud operations, automation, diagnostics, and observability
  • Stay current with AI-driven DevOps/SRE tooling
  • Act as a self-starter, owning initiatives end-to-end
  • Demonstrate strong leadership, communication, and problem-solving skills
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service