We are seeking an experienced Site Reliability Engineering leader to build and lead a new SRE organization responsible for the production operation of a rapidly scaling AI supercomputing platform. The environment combines highly customized compute, high-performance networking, storage and supporting infrastructure, and will grow through multiple phases of deployment. This is a rare opportunity to establish the reliability function for a new platform from the ground up. The platform and its operational model are being developed in parallel and will ultimately support a 24x7x365 production service with stringent availability requirements. You will take the SRE organization from initial formation through production launch, stabilization and scale. This includes hiring and developing the team, defining the operating model, establishing production readiness and incident-management practices, and ensuring reliability and operability are engineered into the platform from the outset. SRE is responsible for the operational capability required to run the platform reliably in production, while partnering with engineering teams that remain accountable for the reliability and operability of the systems they build. This is not a purely managerial position. During the development and early production phases, the SRE Manager will be expected to work directly with engineering teams, develop a deep understanding of the platform, and participate in troubleshooting and incident response. Over time, success will increasingly mean building the people, processes, automation, tooling, and operational discipline that allow the organization to operate effectively without depending on you for day-to-day escalation.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed