Cloud Engineer

VusionGroup SADallas, TX
Hybrid

About The Position

As a Cloud Engineer on the Cloud Operations team, you will make a positive impact by communicating with customers during incidents, identifying impacts and scope of issues, providing timely updates via status pages and communications channels (email, teams), leading and contributing to Root Cause Analysis (RCA) documentation, and participating in on-call rotations and ensuring high availability of cloud services. You will also be involved in CI/CD Pipeline Engineering & Automation, Infrastructure & Platform Engineering, development phases, Monitoring, Observing & Reliability, Ensuring Security, and continuous improvement. This role involves provisioning and managing cloud infrastructure using Infrastructure as Code (IaC) tools, managing containerized environments using Kubernetes, and ensuring scalability, reliability, and performance of cloud-native applications. You will also provide insights based on production environments, metrics, scalability, and constraints, challenge system design, and support production readiness reviews and deployment validations. Additionally, you will design and implement monitoring, alerting, and logging solutions, create dashboards for system health, performance, and business metrics, and proactively identify reliability risks. You will also take part in security reviews, monitor security tools, and implement improvements. Continuous improvement efforts include identifying improvements, automating repetitive tasks through scripting, optimizing cloud costs, and contributing to documentation and knowledge sharing.

Requirements

  • At least 3 years of experience in a similar role
  • Strong experience with cloud computing platforms, ideally Microsoft Azure
  • Experience with DevOps tools, such as Azure DevOps
  • Experience with containerization technologies such as Kubernetes
  • Fluently speaking and writing in English
  • Excellent verbal and written communication; ability to convey complex information in a clear and understandable manner
  • Ability to liaise with individuals across a wide variety of operational, functional, and technical disciplines and work within a virtual global team environment
  • Bachelor's degree in Computer Science or a related field, or equivalent experience

Responsibilities

  • Communicating with customers during incidents
  • Identifying impacts and scope of issues
  • Providing timely updates via status pages and communications channels (email, teams)
  • Leading and contributing to Root Cause Analysis (RCA) documentation
  • Participating in on-call rotations and ensuring high availability of cloud services
  • Designing, building, and maintaining CI/CD pipelines using tools such as Azure DevOps
  • Automating build, test, and deployment workflows to improve release efficiency and reliability
  • Implementing pipeline governance, versioning strategies, and release approvals
  • Troubleshooting pipeline failures and optimizing performance and reliability
  • Integrating security scans, code quality checks, and automated testing into pipelines
  • Provisioning and managing cloud infrastructure using Infrastructure as Code (IaC) tools (e.g., ARM templates, Terraform, Bicep)
  • Managing and optimizing containerized environments using Kubernetes
  • Ensuring scalability, reliability, and performance of cloud-native applications
  • Managing environment configurations across development, staging, and production
  • Providing insights based on production environments, metrics, scalability, and constraints
  • Challenging system design and proposing improvements as early as possible
  • Supporting production readiness reviews and deployment validations
  • Testing and validating disaster recovery and business continuity procedures
  • Designing and implementing monitoring, alerting, and logging solutions (e.g., Azure Monitor, Log Analytics)
  • Creating dashboards for system health, performance, and business metrics
  • Proactively identifying reliability risks and implementing preventive measures
  • Taking part in security reviews
  • Monitoring security tools
  • Identification of improvements
  • Implementing improvements
  • Identifying improvements, including monitoring or dashboarding
  • Taking part in lessons learnt from releases and incidents
  • Automating repetitive tasks through scripting (PowerShell, Bash, Python)
  • Taking part in processes optimization, documentation, audits (e.g. ISO 27001)
  • Optimizing cloud costs through rightsizing, reserved instances, and usage analysis
  • Contributing documentation, knowledge sharing, and process improvements

Benefits

  • Generous paid time off (PTO): 35 days PTO to enable work/life integration and promotes a culture of trust
  • Eligibility for healthcare benefits begin day one
  • Retirement savings plans
  • E-learning opportunities and workshops
  • Global mobility potential
  • Commute benefits: up to $100/month per employee for commuting expenses
  • Company matches employee donations up to $500 per year for causes close to your heart
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service