This role focuses on maintaining the operational health of cloud environments across Azure and Google Cloud, ensuring uptime, performance, and stability. The engineer will be responsible for building, deploying, and operating Kubernetes platforms (AKS & GKE) with zero downtime, and designing and managing multi-cloud infrastructure using Terraform. A key aspect of the role is creating reusable, secure, and scalable Terraform modules and implementing zero-downtime deployment and upgrade strategies. The position also involves designing and managing cloud networking, troubleshooting complex issues, and automating infrastructure and operations using Python, Bash, or PowerShell. Building and maintaining CI/CD pipelines, applying an automation-first approach, and implementing monitoring, logging, alerting, and observability are crucial. The engineer will proactively identify issues, perform root cause analysis, and ensure systems are scalable, resilient, secure, and performant. A deep understanding of platform architecture and dependencies is required, along with driving operational best practices and continuous improvements. Collaboration with engineering and product teams is essential. The role also involves applying LLMs/AI to cloud operations, automation, diagnostics, and observability, and staying current with AI-driven DevOps/SRE tooling. The ideal candidate will be a self-starter, owning initiatives end-to-end, and demonstrating strong leadership, communication, and problem-solving skills.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed