About The Position

FAR.AI is a non-profit AI research institute focused on ensuring advanced AI is safe and beneficial for everyone. The Foundations team is responsible for the institute's infrastructure and engineering, including the compute platform, tools, frameworks, and automation of research workflows. This role specifically involves leading and engineering a new sub-team focused on the GPU cluster fleet, which is rapidly growing and includes bare-metal and managed Kubernetes clusters from multiple providers. The team will manage capacity, networking, storage, infrastructure as code, security, and the platform layer above. The work involves close collaboration with researchers to ensure large-scale experiments are performant and fault-tolerant, and includes setting technical direction, roadmap, and hiring for the team.

Requirements

  • Experience leading engineers as a manager, tech lead, or project lead, including setting technical direction, scoping work, and giving feedback.
  • 5+ years in systems or infrastructure engineering on production Linux, running GPU, HPC, or large-scale batch platforms, with ownership from design through operation.
  • Depth in at least one of scheduling, storage, networking, security, or GPU systems, with breadth to review designs in the others.
  • Experience running production Kubernetes for GPU workloads with a batch layer (Slurm, Kueue, Volcano, or similar), including quotas, priority, preemption, and node health.
  • Experience owning infrastructure as code and observability for a production fleet, using tools like Terraform or Ansible for provisioning, Helm and ArgoCD for deployment, and Prometheus for monitoring.
  • Strong programming skills in at least one infrastructure language (Python, Go, Rust, or C++), with automation and services maintained as shared code.
  • Ability to write clearly for engineers, researchers, and providers in various formats (roadmaps, design docs, incident summaries, escalations).

Nice To Haves

  • Real depth in distributed training infrastructure (multi-node PyTorch and NCCL debugging, NVIDIA node stack, InfiniBand or RoCE fabrics, topology-aware placement).
  • Real depth in distributed storage (VAST, Weka, Lustre, Ceph, object storage at scale, checkpoint I/O).
  • Real depth in cluster security (admission control, RBAC, node/container hardening, sandboxed runtimes, isolating autonomous agents).
  • Real depth in scheduler internals (Kubernetes scheduler plugins, custom controllers, gang scheduling, fair-share, quota, utilization/fairness/latency tradeoffs).
  • Real depth in multi-provider platforms (scheduling/storage across providers, handling no shared network and uneven data locality).
  • Experience in greenfield team-building, including hiring senior engineers and establishing on-call, incident, and review practices for a new team.

Responsibilities

  • Set the platform's technical direction and own its roadmap, deciding which systems to run, how to schedule and store across providers, and what to measure.
  • Own architecture and the scheduling and storage design, and debug failures across layers such as node health, GPU and fabric faults, and multi-node job hangs.
  • Hire and grow a small team of senior engineers, setting priorities, ownership, scoping projects, and providing feedback and coaching.
  • Set security direction for a shared cluster where AI agents run experiments, covering identity and access, workload isolation, and sandboxing.
  • Define operational practices including on-call rotation, incident response, postmortems, fault tolerance, and observability, and participate in them.
  • Serve as the escalation point for research teams and providers, and translate recurring problems into platform fixes.

Benefits

  • Health Insurance - 94% of premium paid by Organization
  • 401(k) plan with up to 2% match
  • 25 days Paid Time Off per year
  • Up to 10 days of paid sick leave per year
  • Paid Bereavement, Family, Medical and Pregnancy Disability Leave
  • Work computer and stipend provided for eligible employees
  • Catered lunches and dinners on workdays at our Berkeley office (for full-time employees in the US)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service