Senior Manager, Production & Fleet Operations

Cerebras Systems•Sunnyvale, CA

About The Position

Cerebras is building and operating some of the world's most advanced AI infrastructure. As our infrastructure footprint expands across cloud, colocation facilities, and customer environments, we need a scalable operating model that provides continuous visibility into production infrastructure and ensures issues are rapidly identified, prioritized, and resolved. We are seeking a Senior Manager, Production & Fleet Operations to lead the day-to-day operational management of Cerebras-managed production infrastructure. Reporting to the Director of Central Operations, this leader will establish and operate the mechanisms required to understand fleet health, coordinate production response, manage maintenance and repair activities, and ensure infrastructure is safely and efficiently returned to service when failures occur. This role will work closely with SiteOps/DC Ops, Reliability & Incident Management, Service Ops & Enablement, Global Service Logistics & Inventory, Build & Deploy, and Engineering. The mission is to operate the Cerebras production fleet safely, reliably, and consistently, providing continuous visibility into infrastructure health, rapid restoration when failures occur, and an increasingly standardized and automated operating model as the fleet scales.

Requirements

  • 8+ years of experience in infrastructure operations, cloud operations, fleet operations, SRE, HPC operations, data center operations, or similar technical environments.
  • 3+ years leading technical operations teams.
  • Experience operating complex, highly available production infrastructure.
  • Strong understanding of monitoring, telemetry, incident response, maintenance, hardware troubleshooting, and service restoration.
  • Experience coordinating work between centralized operations and onsite technical teams.
  • Strong infrastructure troubleshooting and operational judgment.
  • Demonstrated ability to establish repeatable processes in rapidly evolving environments.
  • Strong cross-functional leadership and communication skills.
  • Demonstrated focus on automation and reduction of operational toil.

Nice To Haves

  • AI/HPC, large-scale compute, accelerator, cloud, or distributed infrastructure experience.
  • Experience with hardware-intensive production environments.
  • Experience building or operating NOC/Fleet Operations capabilities.
  • Experience managing hardware repair/RMA processes.
  • Experience operating geographically distributed infrastructure.

Responsibilities

  • Own Production Fleet Operations: Own day-to-day operational health of Cerebras-managed production infrastructure. Establish consistent operational processes across clusters, sites, and operating environments. Maintain clear visibility into production state, degradation, outages, maintenance, and infrastructure availability. Establish fleet monitoring, alert response, triage, escalation, and restoration processes. Develop appropriate 24x7 operational coverage as fleet requirements evolve. Establish clear ownership of production issues from detection through restoration.
  • Build the Fleet Operations Capability: Develop and mature the Cerebras Fleet Operations/NOC capability. Establish standards for monitoring, telemetry, alerting, dashboards, and fleet-health reporting. Develop shift, handoff, escalation, and on-call operating procedures. Ensure operators have the tooling, runbooks, procedures, and access required to safely operate production infrastructure. Partner with Service Ops & Enablement to develop operator competency and training requirements. Drive standardization across geographically distributed infrastructure.
  • Coordinate Physical Intervention: Determine when physical intervention is required. Dispatch and prioritize SiteOps/DC Ops activities based on production impact. Provide clear diagnostic information and execution procedures. Coordinate troubleshooting between centralized Operations and onsite technicians. Validate successful recovery following physical intervention. Authorize return-to-service from an operational perspective. Analyze recurring physical interventions and identify opportunities for improved diagnostics, procedures, serviceability, or automation.
  • Own Systems Repair & RMA Operations: Establish and manage the operational workflow for failed Cerebras systems and components. Coordinate diagnosis, replacement, repair, and RMA activities. Establish clear disposition paths for failed assets. Partner with Global Service Logistics & Inventory to ensure replacement assets and spares are available when required. Maintain visibility into repair status and repair-loop performance. Identify recurring repair patterns and feed them into Reliability & Incident Management and Engineering. Reduce time from failure detection through restored production capacity.
  • Manage Production Maintenance & Change: Establish operational governance for maintenance activities affecting production infrastructure. Coordinate planned maintenance across Central Ops, SiteOps, customers, Engineering, and infrastructure partners. Ensure production changes have appropriate validation, execution, and rollback procedures. Maintain clear maintenance windows and communication mechanisms. Measure and reduce production impact associated with planned maintenance.
  • Drive Operational Automation: Identify high-frequency and high-toil operational activities suitable for automation. Partner with Operations Strategy, Automation & Performance to prioritize automation opportunities. Provide operational requirements and acceptance criteria to Software/Engineering teams implementing automation. Increase automated detection, diagnosis, remediation, and validation where appropriate. Reduce unnecessary Engineering involvement in repeatable operational activities.

Benefits

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source their cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service