Senior Manager, Site Reliability Engineering

OracleReston, VA
$121,500 - $264,100

About The Position

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing [email protected] [[email protected]] or by calling 1-888-404-2494 in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Requirements

  • Supervises team members and provides direction to ensure accurate forecasting of demands for infrastructure and response to capacity needs, ensuring systems have sufficient resources to handle current and future workloads and identifying resource gaps.
  • Maintains a collaborative relationship with the software development team to develop infrastructures, ensuring features are reliable and scalable according to deployment requirements.
  • Monitors data collection, triage, technical analysis, and redirection, ensuring team members maintain and optimize operations and infrastructure reliability.
  • Provides support to team members monitoring services, ensuring they maintain up-to-date knowledge of performance and document their condition.
  • Leverages advanced knowledge to aid team members in performing incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
  • Monitors comprehensive health and performance reporting and ensures team members take appropriate actions based on trends in data.
  • Ensures team members adhere to procedures when performing provisioning to support infrastructure, applications, and services.
  • Encourages team members to experiment with new approaches for and perform decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
  • Implements standards for identifying and recommending opportunities for automation and assesses potential benefits to enhance operational efficiency.
  • Takes a proactive role in reviewing and offering feedback on design, automation tools, or scripts, acting as a leader during implementation.
  • Shares strategies for conducting testing on automations to ensure they perform tasks correctly and produce expected results.
  • Reviews and provides feedback on release notes and ensures team members communicate comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams.
  • Proactively anticipates and articulates the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations.
  • Serves as a resource to team members on what information to communicate and how to communicate.
  • Serves as a senior management escalation point for incidents and complex issues arising within Oracle services.
  • Monitors the resolution of technical issues spanning multiple services, ensuring effective investigation and debugging techniques are leveraged to achieve SLOs (service level objectives).
  • Shares expectations for documenting incidents performing root cause analyses, guiding team members to capture essential information for analysis and future reference.
  • Implements guidelines for post-mortem procedures to prevent incident reoccurrence.
  • Ensures team members adhere to service level agreements (SLAs) made with customers.
  • Sets expectations for conducting experiments and evaluating cutting-edge tools and technologies to optimize infrastructure performance and reliability, taking proactive steps to adhere to security standards.
  • Manages and contributes to the prioritization of initiatives to improve performance bottlenecks and deployments, ensuring efficient resource usage, speed, and scalability.
  • Implements standards for developing and maintaining knowledge of site reliability trends and sharing valuable insights and information with team members, management, and beyond to promote innovative building, testing, deploying, and running services.
  • Leverages analyses and data from teams to contribute to business development decisions (e.g., design changes).

Responsibilities

  • Supports team members designing and architecting infrastructure and/or service, sharing guidance on practices and terms for reliability and functionality.
  • Supervises team members and provides direction to ensure accurate forecasting of demands for infrastructure and response to capacity needs, ensuring systems have sufficient resources to handle current and future workloads and identifying resource gaps.
  • Maintains a collaborative relationship with the software development team to develop infrastructures, ensuring features are reliable and scalable according to deployment requirements.
  • Implements expectations for identifying opportunities for prototyping and manages prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding) to explore novel approaches.
  • Monitors data collection, triage, technical analysis, and redirection, ensuring team members maintain and optimize operations and infrastructure reliability.
  • Provides support to team members monitoring services, ensuring they maintain up-to-date knowledge of performance and document their condition.
  • Leverages advanced knowledge to aid team members in performing incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).
  • Monitors comprehensive health and performance reporting and ensures team members take appropriate actions based on trends in data.
  • Ensures team members adhere to procedures when performing provisioning to support infrastructure, applications, and services.
  • Encourages team members to experiment with new approaches for and perform decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
  • Implements standards for identifying and recommending opportunities for automation and assesses potential benefits to enhance operational efficiency.
  • Takes a proactive role in reviewing and offering feedback on design, automation tools, or scripts, acting as a leader during implementation.
  • Shares strategies for conducting testing on automations to ensure they perform tasks correctly and produce expected results.
  • Reviews and provides feedback on release notes and ensures team members communicate comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams.
  • Proactively anticipates and articulates the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations.
  • Serves as a resource to team members on what information to communicate and how to communicate.
  • Serves as a senior management escalation point for incidents and complex issues arising within Oracle services.
  • Monitors the resolution of technical issues spanning multiple services, ensuring effective investigation and debugging techniques are leveraged to achieve SLOs (service level objectives).
  • Shares expectations for documenting incidents performing root cause analyses, guiding team members to capture essential information for analysis and future reference.
  • Implements guidelines for post-mortem procedures to prevent incident reoccurrence.
  • Ensures team members adhere to service level agreements (SLAs) made with customers.
  • Sets expectations for conducting experiments and evaluating cutting-edge tools and technologies to optimize infrastructure performance and reliability, taking proactive steps to adhere to security standards.
  • Manages and contributes to the prioritization of initiatives to improve performance bottlenecks and deployments, ensuring efficient resource usage, speed, and scalability.
  • Implements standards for developing and maintaining knowledge of site reliability trends and sharing valuable insights and information with team members, management, and beyond to promote innovative building, testing, deploying, and running services.
  • Leverages analyses and data from teams to contribute to business development decisions (e.g., design changes).

Benefits

  • flexible medical
  • life insurance
  • retirement options
  • volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service