Lead Principal Software Engineer

Oracle•Nashville, TN
•$146,300 - $306,400•Remote

About The Position

Oracle Cloud Infrastructure (OCI) delivers mission-critical cloud services to enterprises worldwide. The Physical Networking Automation and Tooling team builds and operates the software platforms that enable Network Engineers to manage OCI’s global physical network through automation, observability, and actionable insights. We are building intelligent network automation platforms that combine network telemetry, topology, device state, operational workflows, and AI/ML to automate work across the network lifecycle. You will set the technical direction for AI agents that analyze network context, interact with approved tools, and execute controlled workflows with appropriate human oversight. You will define how these capabilities integrate into reliable, scalable software platforms that support production network operations. As a Lead Principal Software Engineer, you will own architecture and technical direction for these platforms, aligning engineering teams and driving complex initiatives from concept through production. You will remain hands-on in critical design and implementation, establish engineering standards for reliability and safe automation, and mentor engineers and technical leaders. Working closely with engineering and network operations leaders, you will shape how OCI operates its physical network as it grows.

Requirements

  • At least 15 years of experience in software engineering, network automation, cloud infrastructure, or a related field, including technical leadership of initiatives spanning multiple teams.
  • Bachelor’s degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience.
  • Strong programming skills in Java and Python, with experience building and operating production software.
  • Deep experience designing distributed systems, APIs, event-driven data pipelines, workflow systems, or cloud-native services.
  • Experience defining technical architecture and leading complex projects from concept through production operation.
  • Experience with network operations, data center infrastructure, or other large-scale infrastructure systems.
  • Experience building AI/ML or LLM-based capabilities to operational problems, including evaluating their output and defining when human review is required.
  • Strong knowledge of Linux, software development practices, CI/CD, observability, and production incident response.
  • Demonstrated ability to influence senior stakeholders, resolve architectural tradeoffs, and mentor experienced engineers.

Nice To Haves

  • Experience developing AI/ML automation workflows with network telemetry, device connectivity, configuration validation, or automated remediation in large-scale networks.
  • Track record of architecting agentic systems, including planner and executor patterns, tool orchestration, agent handoffs, state management, checkpointing, retries, and durable workflows.
  • Hands-on work integrating agents with enterprise tools through MCP, A2A, or similar protocols, including authentication, tool schemas, context handling, approval gates, and failure handling.
  • Knowledge of AI evaluation, quality, observability, and governance practices for production systems.
  • Proficiency with Kubernetes, containers, and operating cloud-native platforms in production.
  • Practical use of AI-assisted development tools for coding, testing, review, and debugging.
  • Contributions to industry standards groups or open source projects related to network automation or data center infrastructure, such as the IETF, IEEE 802, or the Open Compute Project.

Responsibilities

  • Define the long-term architecture and technical roadmap for physical network automation, validation, observability, and AI-driven operations across OCI.
  • Lead the design of distributed, event-driven systems that ingest and correlate network failures, performance, power, capacity, topology, and change data at global scale.
  • Architect AI decision systems for anomaly detection, fault diagnosis, capacity forecasting, and remediation. Establish how models and agents use network context, express uncertainty, and escalate decisions to engineers when needed.
  • Direct the development of agent workflows that discover and invoke approved capabilities, coordinate tools and services, maintain execution state, and recover safely from partial failures.
  • Establish controls for actions affecting production networks, including identity and authorization, approval gates, validation, limits on the impact of failures, cancellation, rollback, and audit trails.
  • Define evaluation and observability standards for AI-driven capabilities. Measure diagnostic accuracy, false positives, action quality, latency, cost, and operational outcomes before and after deployment.
  • Set architecture and engineering standards across APIs, backend services, data pipelines, workflow execution, testing, deployment, and production operations.
  • Advise senior leaders on technical investments and risk. Mentor principal and senior engineers and strengthen engineering practices across teams.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service