Compute Software Architect

Oracle•United States,

About The Position

The Software Architect, Compute Platform and Control Plane will be a senior individual contributor responsible for shaping the architecture and technical direction of the core software systems that power the Compute infrastructure platform. This role will work across the compute platform and control plane, spanning provisioning, scheduling, placement, lifecycle management, capacity orchestration, fleet management, service APIs, and the foundational distributed systems that support large-scale cloud infrastructure. The ideal candidate brings exceptional depth in software architecture, distributed systems, and cloud infrastructure, with the ability to reason across complex systems, identify structural weaknesses, simplify architectures, and define designs that can operate reliably at hyperscale. This architect will work closely with senior engineers, engineering leaders, product teams, and adjacent infrastructure organizations to evolve the compute platform to the next level of scalability, reliability, efficiency, and engineering velocity.

Requirements

  • 12+ years of experience in software engineering, distributed systems, cloud infrastructure, or large-scale platform development.
  • Deep expertise in software architecture and distributed systems.
  • Strong experience designing and operating large-scale, highly available services.
  • Significant experience with cloud infrastructure or large-scale infrastructure platforms.
  • Strong understanding of control-plane architecture, service APIs, state management, orchestration, and automation.
  • Experience designing systems for scalability, reliability, fault tolerance, and operational efficiency.
  • Ability to reason across large and complex software systems and identify architectural simplifications.
  • Proven ability to influence technical direction across multiple engineering teams and organizations.
  • Strong programming and systems fundamentals.
  • BS or MS degree or equivalent experience relevant to functional area.
  • 10 more years of software engineering or related experience.

Nice To Haves

  • Experience designing cloud compute platforms or large-scale infrastructure control planes.
  • Deep experience with scheduling, placement, provisioning, resource management, or fleet orchestration systems.
  • Experience operating systems across multiple regions and failure domains.
  • Experience modernizing large-scale infrastructure software while maintaining production stability.
  • Strong understanding of cloud infrastructure economics and the relationship between architecture, utilization, and cost.
  • Experience applying AI-assisted development or AI-driven operational techniques in production engineering environments.
  • Experience working with GPU or accelerated compute infrastructure from a software platform perspective.

Responsibilities

  • Define and evolve the architecture of the core compute software platform.
  • Establish clear architectural principles, service boundaries, APIs, interfaces, and ownership models across compute platform components.
  • Provide deep technical leadership across compute provisioning, scheduling and placement, capacity discovery and allocation, instance and host lifecycle management, configuration and state management, fleet orchestration, health monitoring and remediation, failure recovery, service APIs, and regional isolation.
  • Serve as a technical authority for distributed systems design. Drive architectural rigor in consistency, concurrency, state management, partitioning, replication, idempotency, failure recovery, dependency management, and service isolation.
  • Apply deep cloud infrastructure expertise to the design and evolution of the compute platform. Understand the end-to-end lifecycle of cloud compute capacity, from resource discovery and orchestration through provisioning, customer consumption, maintenance, failure recovery, and retirement.
  • Embed reliability and resilience into the architecture of the platform. Design for fault containment, graceful degradation, recovery, retry safety, deployment safety, and failure isolation.
  • Identify scaling limitations across the compute platform and control plane. Develop architectures that improve throughput, latency, efficiency, and resource consumption while reducing unnecessary infrastructure overhead.
  • Identify legacy complexity, duplicated services, and architectural patterns that limit engineering velocity or operational efficiency. Define pragmatic modernization strategies that allow critical systems to evolve without unnecessary disruption.
  • Help drive the use of AI to advance both the compute platform and software engineering practices. Identify opportunities to apply AI to software development, testing, code analysis, debugging, operational diagnostics, incident analysis, anomaly detection, failure prediction, root-cause analysis, capacity optimization, and automated remediation.
  • Operate as a senior technical leader across organizational boundaries. Lead complex architecture reviews, drive alignment on critical technical decisions, mentor senior engineers, and partner with Distinguished Engineers, Architects, Directors, and Vice Presidents on long-term platform strategy.

Benefits

  • flexible medical
  • life insurance
  • retirement options
  • volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service