We are hiring a hands-on NOC Manager to lead the 24/7 team responsible for monitoring, configuring, and troubleshooting our AI compute infrastructure. This is a working manager role: you will own shift coverage, escalation, and team development, and you will also be at the keyboard yourself — logged into devices, reading interface counters, tracing traffic paths, pushing configuration, and driving major incidents to resolution. This is a remote-hands operating model. The team works from our Richmond office, not the data center floor. All configuration and troubleshooting is done via CLI, out-of-band management, and console access, with physical work executed by on-site smart-hands technicians and colocation staff whom you will direct and hold accountable. The ideal candidate is a network engineer first and a manager second. AI training and inference workloads put unusual pressure on the network — dense east-west traffic, lossless fabric requirements, and jobs that fail expensively when a single link degrades. We need someone who understands what those failures look like at the packet and interface level, and who can build a team that catches them before customers do.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Manager
Education Level
No Education Listed