Principal Site Reliability Engineer

OracleNashville, TN
$84,900 - $209,500

About The Position

As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve service reliability, operational efficiency, and developer productivity across large-scale distributed systems. You will lead complex reliability initiatives, drive automation-first operational practices, and develop software solutions that eliminate manual toil. You will partner closely with software engineering, cloud infrastructure, security, and product teams to architect resilient platforms that meet aggressive availability, scalability, and performance objectives. Success in this role requires deep expertise in distributed systems, cloud infrastructure, coding, automation, observability, incident management, and operational excellence. You will leverage modern AI technologies, machine learning, and intelligent automation to streamline operations, accelerate incident response, improve troubleshooting, and enable autonomous system management. You are expected to be a technical leader who influences architecture, establishes engineering best practices, mentors other engineers, and drives continuous improvements across multiple services and organizations.

Requirements

  • Deep expertise in distributed systems.
  • Deep expertise in cloud infrastructure.
  • Deep expertise in coding.
  • Deep expertise in automation.
  • Deep expertise in observability.
  • Deep expertise in incident management.
  • Deep expertise in operational excellence.

Responsibilities

  • Designing, building, and operating highly available, scalable, secure, and resilient cloud services.
  • Combining software engineering with infrastructure expertise to improve service reliability, operational efficiency, and developer productivity across large-scale distributed systems.
  • Leading complex reliability initiatives.
  • Driving automation-first operational practices.
  • Developing software solutions that eliminate manual toil.
  • Partnering closely with software engineering, cloud infrastructure, security, and product teams to architect resilient platforms that meet aggressive availability, scalability, and performance objectives.
  • Leveraging modern AI technologies, machine learning, and intelligent automation to streamline operations, accelerate incident response, improve troubleshooting, and enable autonomous system management.
  • Influencing architecture, establishing engineering best practices, mentoring other engineers, and driving continuous improvements across multiple services and organizations.

Benefits

  • Flexible medical
  • Life insurance
  • Retirement options
  • Volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service