About The Position

Rivian's Autonomy org needs a Staff Software Engineer, Cloud Infrastructure to join the Cloud Infra team in the AI Platform organization. This role brings deep cloud and data engineering expertise to Rivian's ADAS team, working directly with technical and business stakeholders to build, test, and release complex, mission-critical infrastructure services on AWS, GCP or Multi Cloud. This candidate needs to have a very good understanding of the Cloud Data Platform and Data/ML Ops processes that helps to build, test, and release complex mission critical infrastructure services for Rivian's Autonomy team on multi cloud. In this role You'll work with the AI Platform, Perception, Planning, Simulation, and Vehicle Integration, Product Management, and other technology partners to leverage best practices and reference architectures highlighting AWS/Multi Cloud Platforms and Data/Dev/ML Ops practices.

Requirements

  • 5+ Yrs. of software engineering or in ML/Dev/Data Ops role.
  • 5+ Yrs. Experience authoring, scaling, and managing production infrastructure.
  • 5+ Yrs. Experience with Kubernetes, cloud CI/CD tools, and core cloud services across compute, storage, networking, identity, and secrets management (e.g., AWS EKS/ECS/S3/Lambda/RDS/Systems Manager/Secrets Manager/CloudTrail, Azure AKS/Blob Storage/Functions, or GCP GKE/Cloud Storage/Cloud Functions).
  • 5+ Yrs. Infrastructure as Code and configuration management (Terraform, Pulumi, AWS CloudFormation, or AWS CDK).
  • 5+ Yrs. Experience monitoring applications across cloud environments using Datadog, Prometheus, or native tools such as AWS CloudWatch, or Google Cloud Operations.
  • 5+ Yrs. Experience debugging production systems and performing RCA on incidents.
  • 3+ Yrs. Hands-on with Python, Go, or Java, plus Git-based tooling (GitLab, GitHub, or similar) for automation.
  • 2+ Yrs. of CI/CD and/or GitOps patterns (using GitLab, Jenkins, Argo CD, or similar).
  • 2+ Yrs. Microservice-oriented architectures (using Kubernetes, container orchestration on any major cloud, or Docker Swarm).
  • 2+ Yrs. Knowledge of Agile development of accessible software tools.

Nice To Haves

  • Linux internals, networking, and distributed computing are a plus.
  • AWS or GCP certification, or a Cloud Native (CNCF) certification, is a plus.

Responsibilities

  • Lead, build, test and release complex mission-critical infrastructure services for Rivian's Autonomy team on cloud and/or on-prem.
  • Setup fault tolerant multi-region environments for data operations and data applications.
  • Own CI/CD pipeline for apps and data projects.
  • Define on-call strategy and participate in on-call rotations.
  • Make developers' lives smooth via automated workflows.
  • Build and optimize highly reliable, scalable, and distributed infra using microservice architecture.
  • Collaborate with the security & privacy team to perform audits and mitigate any findings.
  • Collaborate with cross-functional Autonomy teams for development and integrations.
  • Cost optimization in AWS or GCP across multiple accounts and services.

Benefits

  • paid vacation
  • paid sick leave
  • life insurance
  • medical insurance
  • dental insurance
  • vision insurance
  • short-term disability insurance
  • long-term disability insurance
  • 401(k) Plan
  • Employee Stock Purchase Program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service