Principal Site Reliability Engineer

OracleNashville, TN
$84,900 - $209,500

About The Position

The successful candidate will serve as a senior technical authority, establish reliability standards, guide complex technical decisions, and lead improvements that reduce operational risk and manual effort. This individual must be comfortable moving between architecture and hands-on execution, including accessing deployed hosts, troubleshooting failed services, reviewing logs, correcting configurations, and validating production changes. Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.

Requirements

  • Serve as a senior technical authority.
  • Establish reliability standards.
  • Guide complex technical decisions.
  • Lead improvements that reduce operational risk and manual effort.
  • Comfortable moving between architecture and hands-on execution.
  • Accessing deployed hosts.
  • Troubleshooting failed services.
  • Reviewing logs.
  • Correcting configurations.
  • Validating production changes.

Responsibilities

  • Designs and architects infrastructure and service to ensure reliability and functionality.
  • Forecasts demands and responds to capacity needs.
  • Collaborates with software development teams to develop reliable and scalable infrastructures.
  • Exercises judgment when performing data collection to maintain and optimize operations and reliability.
  • Leverages advanced knowledge to perform incident response and/or maintenance tasks.
  • Provides comprehensive health and performance reporting.
  • Identifies and recommends opportunities for automation.
  • Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes.
  • Provides comprehensive support for technology and documents incidents.
  • Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.

Benefits

  • Flexible medical
  • Life insurance
  • Retirement options
  • Volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service