About The Position

We are looking for experienced engineers to join an incubation team responsible for initiating, prototyping, and transitioning new projects with broad impact across OCI. One of our key long-term initiatives is the development of a new platform designed to power OCI services at cloud scale. The platform encompasses low-level execution runtimes, application and lifecycle management, and high-level change-management workflows, providing OCI developers with the infrastructure and abstractions they need to build and operate services efficiently. Our goal is simple: enable OCI developers to focus on building innovative services while the platform provides the scalability, reliability, security, and operational capabilities required to run them at cloud scale. As part of this team, you will challenge existing engineering assumptions, explore new architectural approaches, and apply your expertise in high-performance and reliable systems to help evolve OCI's infrastructure.

Requirements

  • Experienced engineers to join an incubation team responsible for initiating, prototyping, and transitioning new projects with broad impact across OCI.
  • Development of a new platform designed to power OCI services at cloud scale.
  • Platform encompasses low-level execution runtimes, application and lifecycle management, and high-level change-management workflows.
  • Provide OCI developers with the infrastructure and abstractions they need to build and operate services efficiently.
  • Enable OCI developers to focus on building innovative services while the platform provides the scalability, reliability, security, and operational capabilities required to run them at cloud scale.
  • Challenge existing engineering assumptions, explore new architectural approaches, and apply expertise in high-performance and reliable systems to help evolve OCI's infrastructure.
  • Design and build distributed systems at cloud scale.
  • Optimize critical code and data paths for high-throughput, hyperscale workloads.
  • Leverage data-plane platforms for large-scale retrieval, storage, and processing.
  • Architect resilient and fault-tolerant systems.
  • Design systems that support in-service upgrades, safe patching, updates, and rollbacks.
  • Apply load shedding, throttling, rate limiting, backpressure, and other resilience mechanisms.
  • Define and maintain Service Level Objectives (SLOs), Key Performance Indicators (KPIs), telemetry, dashboards, and proactive alerting.
  • Develop sophisticated validation strategies—including fault injection, brownout testing, failure simulation, and resilience testing.
  • Design and improve replication and synchronization mechanisms.
  • Proactively investigate and resolve complex production issues.
  • Perform deep technical analysis across system boundaries to identify root causes and long-term improvements.
  • Drive operational readiness.
  • Implement robust security controls and remediation strategies.
  • Build Infrastructure as Code (IaC), automation, and deployment tooling.
  • Contribute to architectural direction, technical standards, and engineering best practices.
  • Mentor peers and raise the technical bar across the team.
  • Prototype and evaluate new technologies and architectural approaches.
  • Help transition successful incubations into production systems and broader OCI adoption.
  • Design, build, and operate highly scalable, reliable, and secure distributed systems that support OCI services at hyperscale.
  • Take ownership of critical components from architecture and implementation through production operations.
  • Contribute to technical direction, engineering excellence, and the growth of the broader team.
  • Lead the design, development, and implementation of components for scalable, elastic distributed systems.
  • Support both horizontal and vertical scaling as workload demands evolve.
  • Define scalability and performance requirements for owned components.
  • Ensure requirements are reflected throughout design, implementation, testing, and production operation.
  • Optimize critical code paths and system architectures for high-throughput, large-scale data processing and hyperscale workloads.
  • Design systems with elasticity in mind.
  • Enable resources to scale efficiently both up and down based on demand.
  • Leverage distributed state-management technologies and data-plane platforms.
  • Support large-scale data retrieval, storage, processing, and coordination.
  • Develop comprehensive performance, scalability, capacity, and load-testing strategies.
  • Validate system behavior under expected and extreme workloads.
  • Design and build fault-tolerant, highly available systems.
  • Remain operational during failures, maintenance, and in-service updates.
  • Implement redundancy, replication, automatic failover, and recovery mechanisms.
  • Minimize disruption and protect critical workloads.
  • Design systems to operate predictably during infrastructure and network failures, including partitions and partial dependency failures.
  • Make appropriate tradeoffs among consistency, availability, and partition tolerance.
  • Implement and optimize resilience mechanisms including load shedding, throttling, rate limiting, backpressure, retries, and graceful degradation.
  • Establish appropriate availability and durability expectations through clearly defined Service Level Objectives (SLOs).
  • Use SLOs to guide architectural and operational decisions.
  • Design systems and components to support upgrades and maintenance with minimal or no customer-visible downtime.
  • Define meaningful Key Performance Indicators (KPIs), telemetry, and health signals.
  • Measure system performance, reliability, capacity, and operational health.
  • Build and customize dashboards, telemetry pipelines, monitoring systems, and alerting mechanisms.
  • Proactively identify degradation and emerging issues.
  • Use production telemetry and performance data to identify bottlenecks, capacity constraints, and opportunities for architectural improvement.
  • Define and implement functional and correctness requirements for complex features, components, and distributed systems.
  • Develop data replication and synchronization mechanisms that maintain correctness, integrity, consistency, durability, and availability across distributed components.
  • Identify complex failure modes and incorporate appropriate safeguards into system architecture and implementation.
  • Take a proactive role in diagnosing, debugging, and resolving complex issues across production components and distributed systems.
  • Maintain deep technical expertise in owned systems to support effective troubleshooting, performance optimization, and production operations.
  • Design and implement strategies that enable zero- or minimal-downtime maintenance.
  • Reduce or eliminate the need for customer-facing maintenance windows.
  • Ensure systems meet operational-readiness requirements before entering production, including observability, capacity planning, recovery procedures, documentation, and failure handling.
  • Participate in operational support rotations.
  • Provide technical leadership during incident response, mitigation, and recovery.
  • Lead or contribute to root cause investigations.
  • Identify systemic improvements.
  • Ensure lessons from incidents are incorporated into future designs.
  • Mentor engineers in debugging, incident response, operational practices, and distributed-systems troubleshooting.
  • Design and implement robust security controls for applications and infrastructure operating in multi-tenant cloud environments.
  • Apply appropriate encryption, authentication, authorization, access-control, and data-protection mechanisms.
  • Identify security gaps and execute remediation plans to reduce risk and strengthen system security.
  • Ensure infrastructure and services meet applicable security, compliance, and regulatory requirements.
  • Maintain accurate security and compliance documentation.
  • Incorporate security considerations throughout the development lifecycle.
  • Develop and maintain automation, tooling, and Infrastructure as Code (IaC) to provision, configure, operate, and manage cloud infrastructure reliably at scale.
  • Establish and follow change-management practices for application and infrastructure patching, upgrades, deployments, and rollbacks.
  • Design systems and components that enable these processes to become increasingly automated, repeatable, observable, and safe.
  • Improve deployment and operational tooling to reduce manual intervention, minimize risk, and accelerate recovery when changes do not behave as expected.
  • Lead and coordinate moderately complex engineering initiatives.
  • Manage priorities, dependencies, timelines, and deliverables to ensure successful execution.
  • Provide technical oversight across multiple workstreams.
  • Balance short-term delivery with long-term architectural objectives.
  • Prioritize and delegate work effectively.
  • Monitor progress, identify risks early, and adjust execution plans as resources, requirements, or timelines evolve.
  • Drive projects toward completion while maintaining high standards for engineering quality, reliability, security, and operational readiness.
  • Collaborate across engineering teams and organizational boundaries.
  • Align technical direction, expectations, dependencies, and shared objectives.
  • Develop a strong understanding of the needs of business leaders, stakeholders, customers, and partner teams.
  • Ensure proposed solutions address meaningful requirements.
  • Communicate technical decisions, tradeoffs, risks, and recommendations clearly to both technical and non-technical stakeholders.
  • Foster an inclusive engineering environment by actively seeking diverse perspectives, encouraging constructive discussion, and ensuring team members feel heard and respected.
  • Analyze complex technical problems using data, system behavior, telemetry, experimentation, and engineering judgment.
  • Identify effective solutions.
  • Investigate issues across component and organizational boundaries rather than limiting analysis to individual services.
  • Proactively escalate critical or unresolved issues with a clear assessment of impact, risks, alternatives, and recommended solutions.
  • Document problem-solving approaches, architectural decisions, tradeoffs, and lessons learned.
  • Improve organizational knowledge and future decision-making.
  • Continuously expand expertise in distributed systems, cloud infrastructure, reliability engineering, security, automation, and emerging technologies.
  • Stay current with relevant industry trends, technologies, architectural patterns, and engineering best practices.
  • Actively seek and incorporate feedback to strengthen technical and leadership capabilities.
  • Coach and mentor engineers.
  • Share technical knowledge and help others develop stronger design, implementation, debugging, and operational skills.
  • Promote knowledge sharing within and across teams.
  • Identify opportunities to simplify and improve engineering processes, architectures, tools, protocols, and operational workflows.
  • Develop and recommend improvements that increase engineering velocity, reliability, scalability, security, and operational efficiency.
  • Collaborate with partner teams to implement improvements that span organizational or system boundaries.
  • Evaluate the impact of proposed changes on customers, developers, operators, and other stakeholders.
  • Solicit feedback and continuously explore alternative approaches to improve technical and organizational effectiveness.
  • Contribute to building and strengthening the engineering organization through technical mentorship and knowledge sharing.
  • Participate in candidate interviews.
  • Assess technical and problem-solving capabilities.
  • Provide thoughtful hiring recommendations.
  • Help maintain a high engineering bar while supporting the development and success of existing and incoming team members.

Responsibilities

  • Lead the development and begin architecting key components of scalable, elastic distributed systems, taking ownership of their performance, reliability, operability, and security.
  • Design and build distributed systems at cloud scale, defining and enforcing scalability requirements for the components you own.
  • Optimize critical code and data paths for high-throughput, hyperscale workloads, leveraging data-plane platforms for large-scale retrieval, storage, and processing.
  • Architect resilient and fault-tolerant systems using redundancy, replication, failover, and well-defined policies for handling network and infrastructure partitions.
  • Design systems that support in-service upgrades, safe patching, updates, and rollbacks while minimizing customer impact.
  • Apply load shedding, throttling, rate limiting, backpressure, and other resilience mechanisms to maintain service objectives under overload, partial failures, and unreliable network conditions.
  • Define and maintain Service Level Objectives (SLOs), Key Performance Indicators (KPIs), telemetry, dashboards, and proactive alerting to provide clear visibility into system health and performance.
  • Develop sophisticated validation strategies—including fault injection, brownout testing, failure simulation, and resilience testing —to verify system behavior under adverse conditions.
  • Design and improve replication and synchronization mechanisms that preserve correctness, consistency, and durability across distributed components.
  • Proactively investigate and resolve complex production issues, performing deep technical analysis across system boundaries to identify root causes and long-term improvements.
  • Drive operational readiness, ensuring services are observable, supportable, recoverable, and prepared for production at scale.
  • Implement robust security controls and remediation strategies, while maintaining the documentation and processes necessary to satisfy security and compliance requirements.
  • Build Infrastructure as Code (IaC), automation, and deployment tooling that enables repeatable, reliable, and safe infrastructure and service changes.
  • Contribute to architectural direction, technical standards, and engineering best practices while mentoring peers and raising the technical bar across the team.
  • Prototype and evaluate new technologies and architectural approaches, helping transition successful incubations into production systems and broader OCI adoption.
  • Help design, build, and operate highly scalable, reliable, and secure distributed systems that support OCI services at hyperscale.
  • Take ownership of critical components from architecture and implementation through production operations, while contributing to technical direction, engineering excellence, and the growth of the broader team.
  • Lead the design, development, and implementation of components for scalable, elastic distributed systems, supporting both horizontal and vertical scaling as workload demands evolve.
  • Define scalability and performance requirements for owned components and ensure those requirements are reflected throughout design, implementation, testing, and production operation.
  • Optimize critical code paths and system architectures for high-throughput, large-scale data processing and hyperscale workloads.
  • Design systems with elasticity in mind, enabling resources to scale efficiently both up and down based on demand.
  • Leverage distributed state-management technologies and data-plane platforms to support large-scale data retrieval, storage, processing, and coordination.
  • Develop comprehensive performance, scalability, capacity, and load-testing strategies to validate system behavior under expected and extreme workloads.
  • Design and build fault-tolerant, highly available systems capable of remaining operational during failures, maintenance, and in-service updates.
  • Implement redundancy, replication, automatic failover, and recovery mechanisms that minimize disruption and protect critical workloads.
  • Design systems to operate predictably during infrastructure and network failures, including partitions and partial dependency failures, while making appropriate tradeoffs among consistency, availability, and partition tolerance.
  • Implement and optimize resilience mechanisms including load shedding, throttling, rate limiting, backpressure, retries, and graceful degradation.
  • Establish appropriate availability and durability expectations through clearly defined Service Level Objectives (SLOs) and use them to guide architectural and operational decisions.
  • Design systems and components to support upgrades and maintenance with minimal or no customer-visible downtime.
  • Define meaningful Key Performance Indicators (KPIs), telemetry, and health signals to measure system performance, reliability, capacity, and operational health.
  • Build and customize dashboards, telemetry pipelines, monitoring systems, and alerting mechanisms that proactively identify degradation and emerging issues.
  • Use production telemetry and performance data to identify bottlenecks, capacity constraints, and opportunities for architectural improvement.
  • Define and implement functional and correctness requirements for complex features, components, and distributed systems.
  • Develop data replication and synchronization mechanisms that maintain correctness, integrity, consistency, durability, and availability across distributed components.
  • Identify complex failure modes and incorporate appropriate safeguards into system architecture and implementation.
  • Take a proactive role in diagnosing, debugging, and resolving complex issues across production components and distributed systems.
  • Maintain deep technical expertise in owned systems to support effective troubleshooting, performance optimization, and production operations.
  • Design and implement strategies that enable zero- or minimal-downtime maintenance, reducing or eliminating the need for customer-facing maintenance windows.
  • Ensure systems meet operational-readiness requirements before entering production, including observability, capacity planning, recovery procedures, documentation, and failure handling.
  • Participate in operational support rotations and provide technical leadership during incident response, mitigation, and recovery.
  • Lead or contribute to root cause investigations, identify systemic improvements, and ensure lessons from incidents are incorporated into future designs.
  • Mentor engineers in debugging, incident response, operational practices, and distributed-systems troubleshooting.
  • Design and implement robust security controls for applications and infrastructure operating in multi-tenant cloud environments.
  • Apply appropriate encryption, authentication, authorization, access-control, and data-protection mechanisms.
  • Identify security gaps and execute remediation plans to reduce risk and strengthen system security.
  • Ensure infrastructure and services meet applicable security, compliance, and regulatory requirements.
  • Maintain accurate security and compliance documentation and incorporate security considerations throughout the development lifecycle.
  • Develop and maintain automation, tooling, and Infrastructure as Code (IaC) to provision, configure, operate, and manage cloud infrastructure reliably at scale.
  • Establish and follow change-management practices for application and infrastructure patching, upgrades, deployments, and rollbacks.
  • Design systems and components that enable these processes to become increasingly automated, repeatable, observable, and safe.
  • Improve deployment and operational tooling to reduce manual intervention, minimize risk, and accelerate recovery when changes do not behave as expected.
  • Lead and coordinate moderately complex engineering initiatives, managing priorities, dependencies, timelines, and deliverables to ensure successful execution.
  • Provide technical oversight across multiple workstreams while balancing short-term delivery with long-term architectural objectives.
  • Prioritize and delegate work effectively, monitor progress, identify risks early, and adjust execution plans as resources, requirements, or timelines evolve.
  • Drive projects toward completion while maintaining high standards for engineering quality, reliability, security, and operational readiness.
  • Collaborate across engineering teams and organizational boundaries to align technical direction, expectations, dependencies, and shared objectives.
  • Develop a strong understanding of the needs of business leaders, stakeholders, customers, and partner teams to ensure proposed solutions address meaningful requirements.
  • Communicate technical decisions, tradeoffs, risks, and recommendations clearly to both technical and non-technical stakeholders.
  • Foster an inclusive engineering environment by actively seeking diverse perspectives, encouraging constructive discussion, and ensuring team members feel heard and respected.
  • Analyze complex technical problems using data, system behavior, telemetry, experimentation, and engineering judgment to identify effective solutions.
  • Investigate issues across component and organizational boundaries rather than limiting analysis to individual services.
  • Proactively escalate critical or unresolved issues with a clear assessment of impact, risks, alternatives, and recommended solutions.
  • Document problem-solving approaches, architectural decisions, tradeoffs, and lessons learned to improve organizational knowledge and future decision-making.
  • Continuously expand expertise in distributed systems, cloud infrastructure, reliability engineering, security, automation, and emerging technologies.
  • Stay current with relevant industry trends, technologies, architectural patterns, and engineering best practices.
  • Actively seek and incorporate feedback to strengthen technical and leadership capabilities.
  • Coach and mentor engineers, sharing technical knowledge and helping others develop stronger design, implementation, debugging, and operational skills.
  • Promote knowledge sharing within and across teams.
  • Identify opportunities to simplify and improve engineering processes, architectures, tools, protocols, and operational workflows.
  • Develop and recommend improvements that increase engineering velocity, reliability, scalability, security, and operational efficiency.
  • Collaborate with partner teams to implement improvements that span organizational or system boundaries.
  • Evaluate the impact of proposed changes on customers, developers, operators, and other stakeholders.
  • Solicit feedback and continuously explore alternative approaches to improve technical and organizational effectiveness.
  • Contribute to building and strengthening the engineering organization through technical mentorship and knowledge sharing.
  • Participate in candidate interviews, assess technical and problem-solving capabilities, and provide thoughtful hiring recommendations.
  • Help maintain a high engineering bar while supporting the development and success of existing and incoming team members.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position.
  • Accrued Vacation is provided to all other employees eligible for vacation benefits.
  • For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment.
  • Vacation accrual is prorated for employees working between 20 and 34 hours per week.
  • 11 paid holidays
  • 72 hours of paid sick leave upon date of hire.
  • Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service