The OCI AI Infrastructure Network Operations team operates and improves the high-performance RDMA/RoCE network fabrics powering OCI’s largest AI, GPU, and HPC workloads. As a Senior Manager, you will lead a team responsible for building, operating, and scaling these critical network fabrics and supporting systems. You will combine deep networking expertise in RDMA/RoCE, Clos fabrics, congestion control, telemetry, and performance troubleshooting with strong software engineering and people leadership. You will drive automation, monitoring, resiliency, and operational readiness while partnering across Network Availability, Automation, Monitoring, GNOC, hardware engineering, and service teams. Your team will improve network performance and availability, resolve complex customer issues, and build fault-tolerant systems that support AI infrastructure at global cloud scale.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed