Site Reliability Engineer

Charles Schwab Inc.Austin, TX
$120,000 - $155,000Onsite

About The Position

At Schwab, you’re empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together. We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site 4-days per week during night shifts (2pm-10pm CST), and weekends as needed, in the specified location(s). As a Site Reliability Engineer, you will play a critical role in protecting the stability, performance, and resiliency of Schwab’s Order Management System, supporting the technology that enables our clients to navigate their financial futures with confidence. In this highly visible role, you will assess and resolve complex production incidents, drive rapid restoration of critical systems, and collaborate across engineering, infrastructure, databases, and vendor teams to minimize business impact and improve client experiences. Success in this role requires strong problem-solving capabilities, sound decision-making during high-pressure situations, and the ability to navigate complex distributed environments. You will partner closely with technology teams to strengthen production readiness, improve operational excellence, and identify opportunities to reduce recurring incidents through automation, observability, and continuous improvement initiatives. As a senior member of the team, you will also help shape operational best practices, mentor fellow engineers, and contribute to a culture of collaboration, accountability, and knowledge sharing. At Schwab, we succeed together as One Schwab. You will have the opportunity to make a meaningful impact while developing your technical expertise, leadership capabilities, and operational excellence skills in a collaborative environment that values innovation, adaptability, and continuous learning.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
  • 5+ years of experience in Production Support, Site Reliability Engineering (SRE), Software Operations, or a related technology support role.
  • Advanced experience troubleshooting Java-based distributed applications and SQL-backed enterprise systems.
  • Strong experience with application monitoring and observability tools, including AppDynamics, Splunk, Grafana, InfluxDB, and Control-M.
  • Strong Oracle Database knowledge with experience in SQL analysis, database troubleshooting, and performance optimization.
  • Strong Linux administration experience, preferably supporting RHEL 7/8/9 environments.
  • Experience using scripting or automation technologies such as Python or Shell scripting to improve operational efficiency.
  • Experience leading incident response activities and managing high-severity production incidents.
  • Strong written and verbal communication skills with the ability to communicate effectively with both technical and non-technical stakeholders.
  • Availability to support night shifts, weekends, and participation in a rotating on-call support model.

Nice To Haves

  • Experience driving operational excellence initiatives that improve availability, reliability, and mean time to resolution (MTTR).
  • Experience developing or enhancing monitoring strategies, operational runbooks, and escalation procedures.
  • Demonstrated ability to identify root causes, implement permanent corrective actions, and reduce incident recurrence.
  • Experience partnering with software engineering teams to improve application supportability and production readiness.
  • Experience leading change implementation planning and supporting medium-to-high risk production releases.
  • Demonstrated mentoring, coaching, or technical leadership experience within production support or SRE teams.
  • Knowledge of change management, risk management, security, and compliance practices in highly regulated environments.

Responsibilities

  • Assess and resolve complex production incidents.
  • Drive rapid restoration of critical systems.
  • Collaborate across engineering, infrastructure, databases, and vendor teams to minimize business impact and improve client experiences.
  • Partner closely with technology teams to strengthen production readiness, improve operational excellence, and identify opportunities to reduce recurring incidents through automation, observability, and continuous improvement initiatives.
  • Shape operational best practices.
  • Mentor fellow engineers.
  • Contribute to a culture of collaboration, accountability, and knowledge sharing.
  • Identify root causes, implement permanent corrective actions, and reduce incident recurrence.
  • Partner with software engineering teams to improve application supportability and production readiness.
  • Lead change implementation planning and support medium-to-high risk production releases.

Benefits

  • In addition to the salary range, this role is eligible for bonus or incentive opportunities.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service