TensorWave is building the next generation of GPU cloud infrastructure, and our Global Operations Center is the backbone that keeps it running 24/7 across multiple data centers. As Lead Operations Engineer, you’ll be the technical backbone of the GOC and bridge the gap between our frontline operations engineers and the engineering teams that build and maintain our platform. You’re the person who makes the shift teams more effective: developing and validating the runbooks they execute, reviewing major incidents to drive systemic improvements, and working directly with engineering leads to build better alerting and identify tasks that can be safely pushed to the operations floor. You’ll work with the Head of Global Operations to own the operational maturity of the GOC and be the driving force behind turning reactive firefighting into proactive, repeatable operations.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed