Site Reliability Engineer

FidelityWestlake, TX
Onsite

About The Position

We are looking for a Site Reliability Engineer who solves operational problems by building software. In this role, you will improve reliability, reduce toil, and enhance production systems by writing code, building automations, and leveraging modern AI-assisted development tools. The position requires a deep understanding of application and infrastructure support and expert with supporting cloud computing environments. This is a hands-on engineering role not a traditional support position.

Requirements

  • 2 plus years of experience in systems and platform operations and technology management
  • Experience in Cloud computing(Azure and AWS), VMs, Windows and Linux.
  • Experience in managing Kubernetes cluster administration and expected to have good experience troubleshooting Kubernetes.
  • Experience in Python scripting and PowerShell skills are highly preferred.
  • Experience with analytics and monitoring tools such as Grafana, Splunk and Datadog
  • Experience supporting 24/7, continuous availability production and managed environments.
  • Good understanding of software architecture helps empower software developers and engineers to build platforms with greater resiliency and fault tolerance in mind.
  • Great communication, collaboration, and interpersonal skills

Nice To Haves

  • SRE (Site Reliability Engineering) principles a plus – resiliency, observability, and governance gating
  • Experience with continuous integration tools, such as Jenkins and AWX
  • Good to have knowledge of networking, firewalls and load balancers.
  • Knowledge of best practices for IT operations in an always-on, always-available service model
  • Ideal candidates will have a background in Site Reliability with a strong desire to expand into the other domain, or prior experience as a Site Reliability Engineer.
  • We are looking for a systems-thinking Site Reliability Engineer who has helped teams scale through production insights, operational automation, developer enablement, real-time metrics, and continuous improvement.
  • Optional certifications: AWS, Azure related credentials.

Responsibilities

  • Build automation, scripts, and lightweight tools to eliminate repetitive manual work and improve operational efficiency.
  • Participate in on-call rotations, respond to incidents, execute runbooks, and ensure clear communication and handoffs.
  • Continuously analyze system performance in production, troubleshoot consumer reported issues, and proactively identify areas in need of optimization.
  • Lead change in a creative and collaborative manner with business and technical partners.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service