FAR.AI is a non-profit AI research institute focused on ensuring advanced AI is safe and beneficial. The Foundations team is responsible for the institute's infrastructure and engineering, including the compute platform, tools for researchers, workflow automation, and scaling experiments. This role is within the infrastructure sub-team that owns the GPU cluster fleet, focusing on adding capacity, managing networking and storage, infrastructure as code, and security. The position involves working across the entire infrastructure stack, with a particular emphasis on large-scale pre-training and post-training infrastructure, network fabric, cluster security, distributed storage systems, and batch scheduling for large GPU clusters. The engineer will collaborate with researchers and other engineers to ensure the performance and fault tolerance of large-scale experiments, and infrastructure challenges are often integral to the research itself.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed