Principal Core Infrastructure Engineer

OracleNashville, TN
$114,600 - $234,600

About The Position

We are looking for experienced engineers to join an incubation team responsible for initiating, prototyping, and transitioning new projects with broad impact across OCI. One of our key long-term initiatives is the development of a new platform designed to power OCI services at cloud scale. The platform encompasses low-level execution runtimes, application and lifecycle management, and high-level change-management workflows, providing OCI developers with the infrastructure and abstractions they need to build and operate services efficiently. Our goal is simple: enable OCI developers to focus on building innovative services while the platform provides the scalability, reliability, security, and operational capabilities required to run them at cloud scale. As part of this team, you will challenge existing engineering assumptions, explore new architectural approaches, and apply your expertise in high-performance and reliable systems to help evolve OCI's infrastructure.

Requirements

  • Experienced engineers
  • Expertise in high-performance and reliable systems

Responsibilities

  • Lead the development and begin architecting key components of scalable, elastic distributed systems, taking ownership of their performance, reliability, operability, and security.
  • Design and build distributed systems at cloud scale, defining and enforcing scalability requirements for the components you own.
  • Optimize critical code and data paths for high-throughput, hyperscale workloads, leveraging data-plane platforms for large-scale retrieval, storage, and processing.
  • Architect resilient and fault-tolerant systems using redundancy, replication, failover, and well-defined policies for handling network and infrastructure partitions.
  • Design systems that support in-service upgrades, safe patching, updates, and rollbacks while minimizing customer impact.
  • Apply load shedding, throttling, rate limiting, backpressure, and other resilience mechanisms to maintain service objectives under overload, partial failures, and unreliable network conditions.
  • Define and maintain Service Level Objectives (SLOs), Key Performance Indicators (KPIs), telemetry, dashboards, and proactive alerting to provide clear visibility into system health and performance.
  • Develop sophisticated validation strategies—including fault injection, brownout testing, failure simulation, and resilience testing—to verify system behavior under adverse conditions.
  • Design and improve replication and synchronization mechanisms that preserve correctness, consistency, and durability across distributed components.
  • Proactively investigate and resolve complex production issues, performing deep technical analysis across system boundaries to identify root causes and long-term improvements.
  • Drive operational readiness, ensuring services are observable, supportable, recoverable, and prepared for production at scale.
  • Implement robust security controls and remediation strategies, while maintaining the documentation and processes necessary to satisfy security and compliance requirements.
  • Build Infrastructure as Code (IaC), automation, and deployment tooling that enables repeatable, reliable, and safe infrastructure and service changes.
  • Contribute to architectural direction, technical standards, and engineering best practices while mentoring peers and raising the technical bar across the team.
  • Prototype and evaluate new technologies and architectural approaches, helping transition successful incubations into production systems and broader OCI adoption.

Benefits

  • Flexible medical
  • Life insurance
  • Retirement options
  • Volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service