About The Position

Rivian's Autonomy org needs a Staff Software Engineer, Compute & Storage to own the distributed compute and storage platform that every autonomy workload runs on. This sits in the Platform Services team in the AI Platform organization in the Autonomy team. Autonomy training, simulation base evaluation / validation, and autonomy visualization all depend on the same two things: available compute and fast access to data. The role requires deep expertise in Kubernetes-based distributed compute, large-scale object storage, and the performance and cost tradeoffs of running both at petabyte scale. You'll work with the AI Platform, Perception, Planning, Simulation, and Vehicle Integration, Product Management, and other technology partners to operate a platform serving a fleet of over 100,000 vehicles, hundreds of petabytes of drive and simulation data, and training clusters of thousands of GPUs across multiple clouds. This is a platform ownership role, and it's measured by what it lets other engineers do: how fast someone goes from idea to trained model or data pipeline, how many scenarios simulation runs per day, and what each of those costs.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • 8+ years of software engineering experience building and operating distributed systems in production.
  • 5+ years owning a distributed compute, batch execution, or scheduling platform used by other engineering teams: building the scheduler, orchestration layer, or execution engine itself, not just jobs that ran on one.
  • 5+ years with Kubernetes as an execution substrate: scheduling, resource management, custom controllers or operators, autoscaling, and failure modes at scale.
  • 5+ years with large-scale object storage: data layout, partitioning, lifecycle and tiering, caching, and the performance and cost tradeoffs among them.
  • 3+ years with a distributed processing or training framework (Ray, Spark, Flink, or Dask).
  • 3+ years hands-on with Go, C++ or Rust, plus strong Python.
  • 3+ years with infrastructure as code and configuration management (Terraform, AWS CDK, or CloudFormation), and container build and supply chain.
  • 3+ years with queueing and event-driven systems (SQS, Kafka, Kinesis, or equivalent), including autoscaling workloads on queue depth.
  • 3+ years with production monitoring and alerting (Prometheus, Grafana, Datadog, or CloudWatch): defining the metrics and building the dashboards, not just consuming them.
  • 3+ years debugging production distributed systems, including resource contention, stragglers, data skew, network saturation, and cascading failure, and running root-cause analysis on incidents.
  • 2+ years carrying an infrastructure cost target and meeting it.
  • Ability to turn ambiguous, high-level requirements into a detailed system design and drive it to completion unprompted.
  • Technical influence beyond your own commits: designs others build on, standards others adopt, engineers who improved from working with you.

Nice To Haves

  • GPU cluster management and utilization, including topology-aware placement, gang scheduling, multi-tenant sharing, and fragmentation.
  • Autonomous vehicle, robotics, or another domain with continuous high-volume sensor data.
  • Multi-cloud or hybrid cloud and on-premise operation, including data movement economics between providers.
  • Linux internals, including the CPU scheduler, memory management, file systems, and networking.

Responsibilities

  • Own the architecture and roadmap for Autonomy's distributed compute platform: job scheduling, quota and fair-share across teams, autoscaling, and spot and preemption strategy across multiple clouds.
  • Build scalable tools and APIs that turn high-level job requests into executed work, making large-scale computation accessible to engineers who aren't infrastructure specialists.
  • Read the jobs other teams run, profile them, and find inefficiencies and bottlenecks. Fix them directly, or give the team the tooling to see them.
  • Treat cluster efficiency as a primary metric: eliminate GPU fragmentation, right-size quota, and reclaim capacity stranded on partially filled nodes.
  • Own the storage architecture for autonomy data at petabyte scale, including layout, partitioning, tiering across hot, warm and cold, lifecycle policy, and the caching and prefetch layers that keep training and simulation jobs from starving on I/O.
  • Own the multi-cloud compute and data path as training extends beyond a single provider, including replication strategy, consistency, and cross-provider egress economics.
  • Drive throughput and turnaround time for the heaviest workloads: training data loading, large-scale log replay, and batch resimulation running tens of thousands of concurrent jobs.
  • Own the queue-based and event-driven infrastructure behind job submission and autoscaling, and keep it stable as queue depth moves by orders of magnitude.
  • Own platform observability: define the metrics and build the dashboards and alerting that make job throughput, queue health, cluster utilization, and per-team cost visible to you and the teams you serve.
  • Define SLAs for job admission, completion, and data availability. Set the on-call strategy and take part in the rotation.
  • Own compute and storage cost as a first-class engineering metric, instrumented per team and per workload, across multiple AWS accounts and services.
  • Work with the security & privacy team on data governance, access control, retention, and audit for vehicle-collected data.
  • Set technical standards for how distributed workloads are built and run, and raise the bar through design review, code review, and mentorship of senior engineers.

Benefits

  • paid vacation
  • paid sick leave
  • life insurance
  • medical insurance
  • dental insurance
  • vision insurance
  • short-term disability insurance
  • long-term disability insurance
  • 401(k) Plan
  • Employee Stock Purchase Program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service