Senior Software Engineer, GPU Cluster Infrastructure

FAR.AI,
$150,000 - $275,000Hybrid

About The Position

FAR.AI is a non-profit AI research institute focused on ensuring advanced AI is safe and beneficial. The Foundations team is responsible for the institute's infrastructure and engineering, including the compute platform, tools for researchers, workflow automation, and scaling experiments. This role is within the infrastructure sub-team that owns the GPU cluster fleet, focusing on adding capacity, managing networking and storage, infrastructure as code, and security. The position involves working across the entire infrastructure stack, with a particular emphasis on large-scale pre-training and post-training infrastructure, network fabric, cluster security, distributed storage systems, and batch scheduling for large GPU clusters. The engineer will collaborate with researchers and other engineers to ensure the performance and fault tolerance of large-scale experiments, and infrastructure challenges are often integral to the research itself.

Requirements

  • 3+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms.
  • Owned at least one system from design through operation.
  • Run production Kubernetes for GPU workloads with a batch layer on top (Slurm, Kueue, Volcano, or similar), including quotas, priority and preemption, and node health.
  • Owned infrastructure as code and observability for a production fleet, provisioning with Terraform or Ansible, deploying with Helm and ArgoCD, and monitoring with Prometheus, or their equivalents.
  • Strong programmer in at least one language commonly used for infrastructure (Python, Go, Rust, or C++), with automation and services maintained as shared code.
  • Ability to write clearly for engineers, researchers, and providers in design docs, incident summaries, or escalations.

Nice To Haves

  • Distributed training infrastructure: multi-node PyTorch and NCCL debugging, the NVIDIA node stack (drivers, GPU Operator, DCGM), InfiniBand or RoCE fabrics, topology-aware placement.
  • Distributed storage: VAST, Weka, Lustre, Ceph, or object storage at scale; checkpoint I/O.
  • Cluster security: admission control, RBAC, node and container hardening, sandboxed runtimes (gVisor, Kata, Firecracker), and isolating autonomous agents on shared infrastructure.
  • Scheduler internals: Kubernetes scheduler plugins or custom controllers, gang scheduling, fair-share and quota, and the utilization, fairness, and latency tradeoffs between them.
  • Multi-provider platforms: scheduling and storage across clusters at different providers so users see one system, including clusters with no shared network and uneven data locality.

Responsibilities

  • Operate the Kubernetes GPU fleet day to day, including node lifecycle, upgrades, driver and image rollouts, staged changes with safe rollback, and capacity planning.
  • Own batch scheduling and multi-tenancy, including queues, quotas, priorities, preemption, gang scheduling, and fair share across research teams.
  • Design and run the storage under the fleet, from high-performance shared filesystems for datasets and checkpoints to object storage tiers, quotas, and backups.
  • Keep multi-node training runs fault-tolerant by owning node health and automated draining, debugging NCCL and fabric problems, tracking down stragglers and flaky GPUs, and building checkpoint and restart patterns.
  • Harden the platform, covering identity and access, network policy, secrets, workload isolation, and sandboxing for AI agents.
  • Bring new capacity online by acceptance-testing providers on fabric, NCCL, and storage throughput, holding them to their SLAs, and integrating new clusters into the platform with infrastructure as code.
  • Work directly with research teams on their infrastructure problems and turn recurring issues into platform fixes.
  • Share on-call rotation, runbooks, and postmortems.

Benefits

  • Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date
  • 401(k) plan with up to 2% match
  • 25 days Paid Time Off per year, accrued weekly
  • Up to 10 days of paid sick leave per year
  • Paid Bereavement, Family, Medical and Pregnancy Disability Leave
  • Work computer and stipend provided for eligible employees
  • Catered lunches and dinners on workdays at our office (Berkeley Office Only)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service