Site Reliability Operations Engineer

SalesforceSeattle, WA
$94,000 - $142,300Hybrid

About The Position

As a Site Reliability Operations Engineer you'll be part of our internal DET Site Reliability Operations team supporting our employees globally. This role combines incident command, reliability engineering, and hands-on technical support. You'll help keep critical systems running while working with teams across different time zones.

Requirements

  • 5-8 years in IT operations, incident management, or site reliability work. Experience in a 24x7 high availability environment with enterprise systems preferred.
  • Demonstrated ability to manage high severity incidents under pressure. Establish impact, evaluate solutions with subject matter experts, and make decisions that balance technical and business needs.
  • Strong verbal and written communication skills to explain complex technical issues to both technical and executive audiences. Create clear incident updates and status reports.
  • Demonstrated technical troubleshooting ability across Windows and Linux servers, networking, cloud platforms, and virtualization technologies. Diagnose problems quickly using logs, monitoring tools, and common diagnostic approaches.
  • Experience with cloud platforms (e.g. AWS) and monitoring of IT infrastructure. You should know core cloud concepts and be comfortable with monitoring tools.
  • Understanding of ITIL framework, particularly incident, problem, and change management processes.
  • A related technical degree required.

Nice To Haves

  • Salesforce platform experience and certifications
  • Industry certifications like ITIL, AWS, CCNA, MCSA, or RHCE
  • Scripting ability in Python, Bash, PowerShell, or similar languages to help automate, reduce manual work, and improve efficiency.
  • Experience with monitoring and visualization tools like Splunk, Grafana, or Tableau. Ability to analyze data and identify trends for improving reliability.
  • Background with automation tools like Puppet or Chef

Responsibilities

  • Respond to and manage major incidents affecting internal business operations. Serve as Incident Commander to coordinate technical teams, establish impact, and drive rapid service restoration.
  • Monitor and troubleshoot enterprise systems including infrastructure, applications, and network components. Use your technical skills to diagnose complex problems across multiple platforms and vendors before they impact users.
  • Work with teams globally to improve incident response by creating and improving runbooks, developing SOPs, and driving automation.
  • Coordinate emergency changes and infrastructure updates to resolve incidents. Work with cross-functional teams to maintain business continuity during critical situations.
  • Analyze incident data and KPI metrics to identify trends. Develop actionable recommendations to reduce impact duration and improve performance, then present findings to stakeholders.
  • Lead problem management activities, investigating recurring incidents, documenting root cause analyses, and tracking known errors.
  • Participate in on-call rotation as part of regional coverage. Handle escalations during your shift and serve as Duty Manager for high severity incidents when needed.
  • Track on-call burden and surface toil reduction opportunities with measurable impact.

Benefits

  • time off programs
  • medical
  • dental
  • vision
  • mental health support
  • paid parental leave
  • life and disability insurance
  • 401(k)
  • employee stock purchasing program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service