About The Position

The Cloud Engineering team's mission is to empower teams to reliably run production systems using modern cloud-native tools. They focus on building scalable, automated, and resilient solutions that benefit the entire engineering organization, managing the infrastructure that connects and serves over 26 million vehicles globally. The team is currently executing a multi-phase platform transformation, which includes decoupling services to reduce blast radius, building a shard-based architecture to scale to 2028 demand, and laying the groundwork for GCP migration. This role involves foundational platform engineering that shapes how the company's connected vehicle platform operates for the next five years. The primary customers are internal engineering teams, but the solutions built have a direct impact on vehicle reliability, operational cost, and the company's ability to launch new vehicle programs on time. The Software Engineer on the Cloud Engineering team will be responsible for designing, developing, and maintaining the cloud platform and automation tooling that powers this platform transformation. Contributions will directly impact the reliability, scalability, and cost-efficiency of a system serving millions of vehicles.

Requirements

  • Experience provisioning and managing Kubernetes clusters (EKS or GKE) in production
  • 3+ years managing cloud infrastructure on AWS and/or GCP
  • Strong Terraform skills - writing modular, production-grade IaC with GitOps workflows
  • Experience with Helm for packaging and deploying multi-environment Kubernetes workloads
  • Strong debugging and problem-solving skills for cloud infrastructure and distributed systems
  • Experience with monitoring and observability tools (Prometheus, Grafana, Datadog, or equivalent)
  • Bachelor's or Master's degree in Computer Science, Engineering, or related field
  • 5+ years of professional software or cloud engineering experience
  • Demonstrated experience shipping and operating production infrastructure at scale
  • 5+ years of relevant industry experience in Cloud Infrastructure, Platform Engineering, DevOps, or Site Reliability Engineering (SRE) roles.
  • Kubernetes Administration: Hands-on experience deploying, managing, troubleshooting, and optimizing production Kubernetes clusters.
  • Infrastructure Management: Strong background in cloud infrastructure engineering, including provisioning, configuration management, platform operations, reliability, scalability, and operational excellence.
  • Terraform Experience: Designing and managing Infrastructure as Code (IaC) solutions, including reusable modules and automated infrastructure deployment.
  • Scripting: Strong automation skills using Python, Bash, PowerShell, or similar scripting languages.
  • Helm: Experience creating, maintaining, and deploying Helm charts for Kubernetes-based applications.
  • Cloud Platforms (AWS/Azure): Hands-on experience with cloud services, networking, security, observability, and infrastructure management.
  • Application Development Technologies: Knowledge of Java, Spring Boot, and Python is desirable.

Nice To Haves

  • GCP experience: GKE, Pub/Sub, Cloud Storage, workload identity, VPC-SC, Cloud Operations
  • API gateway experience: Tyk, Apigee, or NGINX - routing, policy management, mTLS, SLI replication
  • Service mesh experience: Istio traffic management, mTLS, and observability
  • gRPC service design and protobuf schema management
  • Experience with CI/CD platforms: ArgoCD, Tekton, Concourse, or similar GitOps tooling
  • Programming proficiency in Python, Go, Java (Spring Boot), or Node.js
  • Security: IAM role isolation, mTLS, OAuth2/OIDC, securing multi-cloud environments at scale
  • Experience with Atlantis, Terragrunt, or similar collaborative Terraform workflows

Responsibilities

  • Design and build cloud-native solutions that enable automated, repeatable shard provisioning, including cluster bootstrapping, network configuration, service mesh setup, and health validation.
  • Build the control plane tooling that allows a new shard to be provisioned and traffic-ready in under 24 hours with zero manual steps.
  • Own and extend production Terraform modules for multi-cloud environments (AWS and GCP).
  • Maintain GitOps-driven infrastructure workflows using Atlantis or similar tooling.
  • Ensure infrastructure changes are reviewed, tested, and auditable before reaching production.
  • Provision, manage, and automate Kubernetes cluster lifecycle (EKS and GKE).
  • Write Helm charts for multi-shard deployments parameterized by shard ID, region, and environment.
  • Configure and operate Istio service mesh including traffic management, mTLS enforcement, and observability.
  • Build and operate GCP-native platform components: GKE workload deployment, Pub/Sub integration, Cloud Storage, Secret Manager, workload identity federation, and VPC-SC perimeter controls.
  • Support the phased migration from AWS to GCP as the architecture evolves.
  • Configure and operate API gateway infrastructure (Tyk, Apigee, or NGINX) including routing policies, mTLS enforcement, rate limiting and SLI/SLO metric replication.
  • Drive zero-downtime ingress migrations with automated runbook validation.
  • Develop and maintain CI/CD pipelines using ArgoCD, Tekton, or Concourse.
  • Automate deployment gates - cluster health checks, fleet distribution validation, rollback triggers - so that shard migrations execute reliably without human intervention at each step.
  • Implement IAM best practices across AWS and GCP: environment-specific role isolation, workload identity, least-privilege access, and automated access reviews.
  • Implement and maintain monitoring for quota thresholds, certificate expiry, and security policy drift.
  • Build and maintain monitoring, alerting, and dashboarding solutions (Prometheus, Grafana, Datadog, Cloud Operations) that give engineering teams real-time visibility into shard health, fleet distribution balance, and platform SLOs.
  • Write automation scripts and internal tools (Python, Go, Bash, Node.js) that improve platform workflows - migration orchestration CLIs, fleet distribution validators, runbook automation, and capacity planning tools.
  • Participate in on-call rotations, troubleshoot incidents, and apply SRE principles to improve system resilience.
  • Drive root cause elimination, not just mitigation.
  • Hands-on experience deploying, managing, troubleshooting, and optimizing production Kubernetes clusters.
  • Candidates should be comfortable with cluster operations, upgrades, networking, security, monitoring, and workload management.
  • Strong background in cloud infrastructure engineering, including provisioning, configuration management, platform operations, reliability, scalability, and operational excellence.
  • Experience designing and managing Infrastructure as Code (IaC) solutions, including reusable modules and automated infrastructure deployment.
  • Strong automation skills using Python, Bash, PowerShell, or similar scripting languages.
  • Experience creating, maintaining, and deploying Helm charts for Kubernetes-based applications.
  • Hands-on experience with cloud services, networking, security, observability, and infrastructure management.
  • Knowledge of Java, Spring Boot, and Python is desirable.

Benefits

  • Medical
  • dental
  • vision
  • 401k
  • Term life
  • Voluntary life and disability insurance
  • Optional Pre-paid legal plan
  • Optional Identity theft plan
  • Optional Medical and dependent FSA
  • Work-visa sponsorship
  • Opportunity for advancement
  • Long-term assignment with opportunity for hire by client
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service