Global Vice President, Compute Control Plane

Oracle•United States,
•$264,500 - $519,900

About The Position

The Global Vice President, Compute Control Plane will lead the engineering organization responsible for the core software systems that power the Compute infrastructure portfolio. This leader will own the evolution, scale, reliability, and operational excellence of the common compute platform and the control plane responsible for provisioning, placement, lifecycle management, health management, capacity orchestration, and fleet operations. The role requires deep expertise in distributed systems, cloud infrastructure software, control-plane architecture, large-scale service design, and operating mission-critical systems at hyperscale. The GVP will drive architectural consistency, simplify complex systems, improve engineering velocity, and raise the bar on reliability and operational performance. A major component of the role will be using AI transformation to take the platform and engineering organization to the next level. The GVP will drive practical adoption of AI across software engineering, testing, operations, reliability, incident management, fleet orchestration, and engineering productivity. This is both an engineering and business leadership role. The GVP will connect platform and AI-transformation investments to customer experience, capacity availability, cost, efficiency, reliability, and measurable business outcomes.

Requirements

  • 15+ years of experience in distributed systems, cloud infrastructure, systems software, large-scale platform engineering, or related technical domains.
  • Significant engineering leadership experience managing large, complex software organizations.
  • Deep expertise in designing and operating distributed systems at scale.
  • Strong experience with control-plane architecture, service-oriented systems, APIs, state management, orchestration, and automation.
  • Proven experience operating highly available, mission-critical software services.
  • Demonstrated success improving the scalability, reliability, and operational efficiency of large infrastructure platforms.
  • Strong operational mindset and experience using service health and business metrics to drive engineering priorities.
  • Demonstrated ability to adopt emerging technologies, including AI, to improve software engineering and operational outcomes.

Nice To Haves

  • Senior engineering leadership experience at a hyperscale cloud or large-scale infrastructure provider.
  • Experience leading a cloud compute control plane, orchestration platform, resource-management system, or comparable distributed infrastructure platform.
  • Deep experience with large-scale scheduling, placement, provisioning, lifecycle management, and fleet orchestration.
  • Strong background in highly available distributed systems and fault-tolerant software design.
  • Experience modernizing complex infrastructure software without disrupting large production environments.
  • Experience applying AI-assisted development, operations, or automation techniques across a large engineering organization.
  • Strong understanding of infrastructure economics and the relationship between software architecture, utilization, and operating cost.

Responsibilities

  • Own the engineering strategy, roadmap, execution, and operational health of the common compute software platform.
  • Drive continued evolution toward greater scale, consistency, modularity, reliability, and operational simplicity.
  • Own the architecture, engineering, and operation of the compute control plane across provisioning, scheduling, placement, capacity allocation, lifecycle management, configuration and state management, health monitoring, remediation, fleet orchestration, failure recovery, and service APIs.
  • Drive adoption of AI across the software development and operational lifecycle.
  • Use AI-assisted development tools to improve developer productivity, code quality, testing, debugging, modernization, and engineering velocity.
  • Apply AI to observability, anomaly detection, root-cause analysis, incident response, troubleshooting, automated remediation, capacity forecasting, scheduling, placement, and fleet optimization.
  • Set a high technical bar for distributed systems design across consistency models, state management, concurrency, fault tolerance, partitioning, replication, recovery, idempotency, dependency management, and service isolation.
  • Own the operational health and reliability of the compute platform and control plane. Drive continuous reduction in customer-impacting incidents, provisioning failures, service degradation, and operational toil.
  • Advance automated remediation, self-healing systems, and intelligent operational tooling.
  • Ensure the compute platform can support continued rapid growth in fleet size, request volume, service complexity, and customer demand. Identify and remove scalability bottlenecks and improve throughput, latency, service efficiency, and software infrastructure cost.
  • Drive a coherent API and service model across the compute platform. Improve consistency, usability, observability, lifecycle management, tooling, automation, and AI-assisted development capabilities.
  • Own the software systems that translate compute capacity into reliable, consumable cloud capacity. Improve discovery, allocation, placement, balancing, onboarding, draining, maintenance, recovery, and lifecycle transitions.
  • Apply AI and predictive analytics to improve forecasting, placement, and fleet health.
  • Ensure the compute platform and control plane can support rapidly growing AI and accelerated compute workloads.
  • Extend common provisioning, lifecycle management, scheduling, health management, and orchestration capabilities to GPU and accelerator environments.
  • Operate as both an engineering leader and a business leader. Maintain detailed visibility into compute availability, provisioning success and latency, control-plane availability, capacity availability, fleet utilization, incidents, mean time to recovery, operational toil, automation coverage, infrastructure cost, and engineering productivity.
  • Lead a large engineering organization spanning distributed systems, control plane, cloud infrastructure software, platform engineering, fleet orchestration, and infrastructure operations.
  • Build a culture centered on technical excellence, simplicity, ownership, accountability, operational discipline, and AI-enabled engineering.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service