The OCI AI Infrastructure – Network Operations team operates the high-performance RDMA/RoCE network fabrics powering OCI’s largest AI, GPU, and HPC workloads. As a Senior Network Engineer, you will operate, support, and scale RDMA/RoCE network fabrics across OCI’s global cloud infrastructure. You will design and validate advanced network solutions for large-scale AI, GPU, HPC, and data center environments. Apply deep networking and automation expertise to improve network reliability, performance, scalability, and operational efficiency. Develop automation, testing frameworks, and tooling to streamline network operations and proactively monitor network health. Build test strategies, perform pre-production validation, and drive root cause analysis (RCA) for complex network issues. Analyze network telemetry and performance metrics to identify anomalies, capacity constraints, and opportunities for improvement. Support incident response, customer escalations, and complex troubleshooting across production environments. Partner with internal engineering teams, vendors, and project teams to validate solutions and deliver network initiatives. Improve monitoring, anomaly detection, and operational tooling for frontline support teams. Mentor junior engineers and contribute to networking standards, operational readiness, and engineering best practices.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed