Senior Network Operations Engineer

Oracle•Nashville, TN
•$81,100 - $187,000•Onsite

About The Position

In this role, you will monitor and troubleshoot network events, collect and analyze technical data, triage and mitigate incidents, coordinate escalations, and help drive continuous operational improvement. You will work alongside experienced engineers, partner teams, and vendors to maintain and optimize the infrastructure that supports Oracle customers and services worldwide. Our mission is to keep OCI’s global network highly available, performant, and resilient—while delivering exceptional service to customers and dependable operational support to our engineering and technical teams. For GNOC engineers, that mission translates into a broad, fast-moving role with real operational impact. You will help centrally manage OCI’s network infrastructure, respond to and resolve complex events, and develop automated solutions that reduce recurring operational work and improve reliability at scale.

Requirements

  • Strong knowledge of networking protocols and technologies, including BGP, OSPF, IS-IS, TCP/IP, IPv4/IPv6, DNS, DHCP, MPLS, VPNs, and TLS.
  • Broad hands-on experience with at least three of the following: Juniper, Cisco, Arista, InfiniBand, NVIDIA, firewalls, routers, switches, circuit management, and optical/network transport services.
  • Strong analytical skills, including the ability to gather, correlate, and interpret data from multiple sources.
  • Ability to diagnose, prioritize, resolve, or appropriately escalate network alerts and faults.
  • Experience in a large ISP, cloud provider, or similarly complex enterprise network environment.
  • Exposure to commodity Ethernet hardware and networking ASICs, including Broadcom and NVIDIA/Mellanox.
  • Hands-on experience supporting carrier circuits and transport services in a 24x7 NOC, ISP, cloud-provider, data center, telecommunications, or large enterprise environment.
  • Demonstrated experience troubleshooting fiber, optical, Ethernet, and WAN circuit failures across multiple carriers and vendors.
  • Experience coordinating carrier escalations, field dispatches, remote hands, circuit turn-ups, maintenance windows, and service restoration.
  • Working knowledge of optical power levels, transceivers, fiber paths, cross-connects, demarcation points, interface counters, and circuit-testing methods.
  • Experience with DWDM, dark fiber, DIA, MPLS, VPLS, microwave, SONET, PON, or comparable transport technologies.
  • Ability to correlate physical-layer and circuit conditions with routing adjacencies, packet loss, latency, congestion, and customer impact.
  • Strong incident-management, technical documentation, vendor-management, and root-cause-analysis skills.
  • Experience supporting GPU and RDMA network environments is highly desirable.
  • Experience supporting high-performance computing (HPC) environments is highly desirable.
  • Experience with InfiniBand and NVIDIA networking technologies, including Spectrum, is highly desirable.
  • Participate in network lifecycle management, including network build, refresh, and upgrade projects.
  • Participate in network solution design and design-review activities.
  • Self-motivated, proactive, and able to work independently.
  • Bachelor’s degree preferred, with at least 3–5 years of relevant network operations or engineering experience.
  • Strong organizational, time-management, verbal, and written communication skills.
  • Comfortable managing a broad range of priorities in a fast-paced operational environment.
  • Experience with incident-response plans, processes, and strategies.
  • Experience supporting large-scale enterprise infrastructure and cloud computing environments in a 24/7 network operations setting, including willingness to work rotational shifts.

Nice To Haves

  • Cisco, Arista and Juniper certifications are desirable.
  • Preferred experience with Python, Puppet, SQL, Ansible, network automation, and databases.

Responsibilities

  • Use established procedures and operational tooling to plan, implement, and safely complete network changes.
  • Mentor, onboard, and train junior network engineers.
  • Participate in operational rotations and provide break/fix and incident-response support.
  • Identify and triage actionable incidents through monitoring systems; analyze and mitigate network events; conduct or support root-cause analysis (RCA); and coordinate follow-up actions with internal support teams and vendors.
  • Provide on-call support as required, exercising sound independent judgment in a varied and complex operational environment.
  • Participate in major incident calls and use technical and analytical skills to resolve network issues affecting Oracle customers and services.
  • Manage fault detection, response, and escalation for OCI systems and networks, collaborating with third-party suppliers through resolution.
  • Collaborate with GNOC Shift Leads and management to ensure the efficient and timely completion of daily GNOC responsibilities.
  • Lead, contribute to, and participate in the identification, development, and evaluation of projects and tools that improve GNOC effectiveness.
  • Drive runbook audits and updates to maintain compliance and align operational processes with partner service teams.
  • Conduct interviews and participate in hiring junior-level engineers.
  • Lead and/or represent the GNOC in vendor meetings, service reviews, and governance boards.
  • Collaborate with network automation teams to integrate and improve operational support tooling.
  • Develop scripts and automation to reduce manual effort and improve the reliability of routine operational tasks.
  • Lead technical initiatives, including the development and improvement of runbooks, methods of procedure (MOPs), operational processes, and team onboarding materials.
  • Support the implementation of short-, medium-, and long-term plans to achieve project objectives.
  • Regularly engage senior management and network leadership to ensure team priorities and project objectives are met.
  • Takes proactive steps to design and architect infrastructure and/or service according to terms for reliability and functionality.
  • Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads.
  • Collaborates with the software development team to develop infrastructures and features that are reliable and scalable according to deployment requirements.
  • Independently identifies opportunities for and drives prototyping (e.g., testing new applications or infrastructures, assisting in onboarding).
  • Performs data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.
  • Independently monitors services, maintains up-to-date knowledge of their performance, and documents their condition.
  • Leverages comprehensive knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
  • Provides health and performance reporting and takes appropriate actions based on trends in data.
  • May independently perform provisioning to support infrastructure, applications, and services.
  • May perform standard and non-standard decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
  • Identifies opportunities for automation and assesses potential benefits.
  • Develops automation tools or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.
  • Independently conducts testing to ensure automation performs the task correctly and produces expected results.
  • Communicates the scale, capacity, security, performance attributes, and requirements of services and technology within and sometimes beyond immediate team.
  • Identifies and explains the potential impact of infrastructure, feature, and tool changes, considering their impact on team operations.
  • Provides operational support for technology, escalating incidents and other standard and non-standard issues arising within Oracle services.
  • Participates in on-call shifts to address issues.
  • Resolves technical issues spanning various services, investigating and debugging products in order to reach SLOs (service level objectives).
  • Documents incidents and performs root cause analyses according to standard reporting methods.
  • Independently performs post-mortem procedures to prevent incident reoccurrence.
  • Experiments with new tools and technologies to assess their potential impact on and improve infrastructure performance and reliability, ensuring adherence to security standards.
  • Independently identifies and executes improvements for performance bottlenecks and deployments to ensure efficient resource usage, speed, and scalability.
  • Develops knowledge of site reliability trends and shares new information with team members, management, and beyond to help others build, test, deploy and run services.
  • Performs standard and non-standard analyses and provides clear data on production to contribute to business development decisions (e.g., design changes).
  • Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
  • Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
  • Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices.
  • Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.
  • Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
  • May be eligible for bonus and equity.
  • Competitive benefits that support our people with flexible medical, life insurance, and retirement options.
  • Volunteer programs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service