Senior Site Reliability Engineer

Oracle•Nashville, TN
•$81,100 - $187,000

About The Position

As a Senior Site Reliability Engineer, you will help design, deploy, maintain, and improve reliable, secure, and scalable infrastructure and services. You’ll proactively identify operational risks and potential failure points, troubleshoot system and application issues, support production environments, and develop automation that improves consistency and reduces manual effort. You will work across Windows Server and Linux environments and participate in deployments, patching, incident response, root-cause analysis, security remediation, and operational documentation.

Requirements

  • Hands-on experience administering Windows Server and/or Linux systems.
  • Ability to access deployed hosts and perform post-deployment configuration and validation.
  • Experience installing, configuring, and validating applications in Windows Server and Linux environments.
  • Ability to troubleshoot operating-system-level, service-level, and application-level issues.
  • Working knowledge of system services, permissions, configuration files, logs, and resource utilization.
  • Experience performing structured build and deployment tasks using runbooks, deployment guides, and technical procedures.
  • Ability to execute scripts, validate outputs, and correct common build or configuration issues.
  • Familiarity with build handoffs, environment-readiness checks, deployment validation, and post-build verification.
  • Ability to follow detailed implementation steps while identifying and documenting exceptions or deviations.
  • Ability to investigate failed services, installation errors, patching failures, application startup problems, permissions issues, and connectivity failures.
  • Experience reviewing logs, event viewers, service status, configuration files, ports, certificates, and permissions.
  • Strong analytical and problem-solving skills, with the ability to isolate root causes and recommend practical solutions.
  • Ability to escalate issues clearly by documenting symptoms, troubleshooting steps, findings, impact, and recommended actions.
  • Experience supporting production or other business-critical environments.
  • Hands-on experience with PowerShell, Bash, Python, Ansible, Chef.
  • Candidates should be able to run, modify, validate, and troubleshoot existing scripts and understand basic automation concepts.
  • Experience applying operating system, middleware, and application patches.
  • Ability to follow patching procedures, validate successful completion, and troubleshoot failures.
  • Understanding of maintenance windows, change control, rollback planning, and post-change validation.
  • Familiarity with Oracle Cloud Infrastructure or another major cloud platform.
  • Understanding of cloud compute, storage, networking, identity, and environment-provisioning concepts.
  • Experience supporting applications in cloud-hosted or hybrid environments.
  • Working knowledge of DNS, firewalls, routing, load balancers, ports, and certificates.
  • Ability to identify basic connectivity issues between hosts, applications, and services.
  • Familiarity with standard network diagnostic tools.
  • Familiarity with STIGs, vulnerability remediation, system hardening, and compliance-driven configuration.
  • Ability to support cybersecurity remediation activities without disrupting application functionality.
  • Ability to follow detailed technical instructions, runbooks, and change procedures.
  • Strong attention to detail when documenting completed work, issues, deviations, and validation results.
  • Experience working in ticketing, incident-management, or change-management systems.
  • Clear written and verbal communication skills for status updates, handoffs, and escalations.
  • Ability to collaborate effectively with engineering, operations, cybersecurity, networking, and client-facing teams.

Nice To Haves

  • Experience supporting federal clients or government-hosted environments.
  • Must be able to obtain and maintain US Federal Security Clearance.
  • Knowledge of cybersecurity workflows, STIG implementation, or federal compliance requirements.
  • Experience with Citrix technologies.
  • Legacy infrastructure or application-support experience.
  • Millennium or Cerner application knowledge.
  • Experience with infrastructure-as-code or configuration-management tools.
  • Production support, incident response, or SRE operational experience.
  • Knowledge of monitoring, alerting, centralized logging, observability, and reliability practices.

Responsibilities

  • Design and architect reliable, secure, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements.
  • Identify operational risks, dependencies, performance issues, and potential failure points before they affect service.
  • Translate business and application requirements into practical technical solutions.
  • Install, configure, deploy, and validate applications across Windows Server and Linux environments.
  • Perform structured builds and deployments using runbooks, scripts, readiness checks, and post-build validation.
  • Troubleshoot operating system, service, application, installation, patching, permissions, certificate, and connectivity issues.
  • Monitor system performance and implement improvements to availability, reliability, and operational efficiency.
  • Develop and maintain scripts and automation that reduce manual effort and improve consistency.
  • Plan and execute operating system, middleware, and application patching using established change-control and rollback procedures.
  • Lead or support incident investigations, root-cause analyses, and corrective actions.
  • Support vulnerability remediation, system hardening, STIG compliance, and other security-driven changes.
  • Maintain accurate runbooks, deployment procedures, troubleshooting guides, and operational records.
  • Provide technical guidance and mentorship to junior engineers.
  • Communicate status, risks, blockers, and escalation details clearly to stakeholders.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service