Director, Platform Software Engineering

Oracle•Nashville, TN
•$122,500 - $355,400•Onsite

About The Position

Oracle Cloud Infrastructure (OCI) builds and operates cloud services that help customers solve some of their largest technical and business challenges. Oracle Kubernetes Engine (OKE), OCI’s managed Kubernetes service, enables customers to deploy, run, scale, and secure Kubernetes workloads with integrated compute, networking, storage, identity, and observability. We are seeking a Director of Platform Software Engineering to lead teams responsible for core OKE platform capabilities. In this role, you will manage and develop engineering managers and senior technical leaders, set technical direction, and own delivery and operational outcomes for critical components of a highly available, globally distributed, 24x7 cloud service. Your organization will advance Kubernetes cluster lifecycle management, orchestration, control plane reliability, scalability, performance, security, automation, and integration with OCI infrastructure. You will help evolve OKE to support larger clusters, more demanding enterprise workloads, and emerging AI and accelerated computing use cases. This role requires deep Kubernetes knowledge, cloud infrastructure experience, strong distributed systems fundamentals, and a demonstrated ability to deliver through multiple engineering teams. You should be comfortable examining architecture and production behavior in detail, challenging technical assumptions, and guiding difficult decisions while empowering managers and engineers to own execution. As a leader within OKE, you will partner with product management, senior architects, operations, and other OCI service teams to translate customer needs into a clear strategy and an achievable roadmap. Success requires sound judgment under ambiguity, disciplined execution, strong communication, and a commitment to developing people and improving the customer experience. You will also guide the adoption of responsible AI-assisted and agentic engineering practices across design, implementation, testing, debugging, documentation, and operations. We expect you to help teams improve productivity while maintaining clear accountability for correctness, security, and production quality.

Requirements

  • Extensive experience designing, building, and operating production software, including cloud infrastructure or distributed platform services.
  • Demonstrated success leading multiple engineering teams, managing engineering managers, developing technical leaders, and delivering complex initiatives across organizational boundaries.
  • Deep Kubernetes expertise and practical understanding of control plane architecture, cluster lifecycle, networking, storage, scalability, and production failure modes.
  • Strong distributed systems fundamentals, including availability, consistency, fault tolerance, performance, and operational tradeoffs.
  • Hands-on cloud infrastructure experience with OCI, AWS, Azure, GCP, or a comparable large-scale environment.
  • Strong software development background and the ability to guide design and implementation in Go and Java, supported by practical Linux, networking, and debugging knowledge.
  • Experience owning production service operations, incident response, safe change management, and sustained reliability improvements.
  • Strong judgment, communication, and execution skills, with the ability to lead through ambiguity and organizational change.

Nice To Haves

  • Experience building managed Kubernetes services or operating Kubernetes infrastructure at substantial scale.
  • Experience with Kubernetes networking and storage integrations, including CNI, CSI, Cilium, Calico, or equivalent technologies.
  • Experience with AI/ML infrastructure, GPU clusters, distributed training or inference, GPU scheduling, device plugins, or high-performance networking.
  • Experience improving engineering productivity through automation and responsible AI-assisted or agentic workflows.
  • Contributions to Kubernetes or related cloud native open-source projects.

Responsibilities

  • Lead, hire, coach, and develop engineering managers and software engineers. Build leadership capacity, establish clear expectations, and create a culture of ownership, collaboration, and technical excellence.
  • Define and execute the technical strategy and roadmap for core OKE platform capabilities, balancing customer needs, feature delivery, reliability, security, performance, and long-term maintainability.
  • Own delivery across multiple teams, including prioritization, staffing, dependencies, milestones, and risk management. Turn ambiguous requirements into clear plans and measurable outcomes.
  • Guide architecture and design for distributed systems that create, update, scale, repair, and operate Kubernetes clusters across OCI regions.
  • Provide technical leadership across Kubernetes control planes, controllers and operators, APIs, etcd, scheduling, autoscaling, container runtimes, and cluster and node lifecycle management.
  • Partner with OCI compute, networking, storage, identity, and security teams to deliver reliable integrations and resolve issues across service and organizational boundaries.
  • Own service health and operational outcomes for your organization’s components of OKE, including availability, capacity, performance, operational readiness, and customer escalations.
  • Lead effective incident response and ensure corrective actions address underlying causes and prevent recurring failures.
  • Establish rigorous standards for design reviews, testing, observability, production readiness, safe deployments, canary validation, upgrades, and rollback.
  • Drive automation that improves fleet health, detects failures earlier, accelerates diagnosis and recovery, and reduces manual operations.
  • Partner with product management and customers to understand workload requirements and use customer feedback and production data to guide investment decisions.
  • Prepare the platform for demanding AI/ML and GPU workloads, working across teams on scalability, orchestration, resource management, and infrastructure integration.
  • Introduce and scale responsible AI-assisted and agentic engineering workflows, measuring improvements in productivity and quality while maintaining security and human accountability.
  • Communicate strategy, delivery progress, service health, risks, and tradeoffs clearly to engineering teams, partners, and senior leadership.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service