DevOps/SRE Engineer

Saxon GlobalPlano, TX

About The Position

The MAPS Quartz team owns the reliability and operational health of the Quartz platform. We actively monitor and support core infrastructure and components, execute automation initiatives, build and enhance internal tools, resolve infrastructure issues, and lead incident response to keep the platform running at scale. We are looking for a Site Reliability Engineer (SRE) who will operate hands on across the stack to improve platform and application observability, drive reliability improvements, and deliver measurable gains in operational efficiency across Global Markets. This role will work closely with Quartz core teams to execute on platform modernization, harden production systems, and evolve support tooling. This position is critical to maintaining execution velocity, reducing operational risk, and ensuring Quartz continues to meet its reliability and performance objectives. Position Summary Collaborates with a diverse set of engineers, architects, and teams to design, develop, test, and implement secure, robust, highly available and scalable solutions for Global Market applications and platforms. Collaborates other software engineers and teams to design and implement deployment approaches using highly scalable, automated, continuous integration and continuous delivery pipelines. Responsible for all aspects of reliability, collaborates with technical experts, key stakeholders, and team members to resolve complex problems, owning the issue until you are sure it will not reoccur. Deep understanding of SRE practices, service level indicators, and service level objectives; proactively utilize them to resolve issues before they impact customers. Gather, analyze, synthesize, and develop visualizations and reporting from large, diverse data sets in service of continuous improvement of the platform. Identify opportunities to eliminate toil and automate the triage of issues to improve overall operational stability. Collaborate with a global team to identify, analyze, and resolve platform vulnerabilities. Proactively promotes the adoption of site reliability engineering best practices within the team and organization.

Requirements

  • At least 5 years of combined experience in either SRE, software development, or infrastructure engineering or related technical field.
  • Strong experience in implementing, monitoring, and maintaining a highly scalable and resilient Application Services and platforms.
  • Strong experience with monitoring tools such as or OpenTelemetry (OTel). ELK (Elasticsearch, Logstash, Kibana), Splunk, and Dynatrace.
  • Knowledge in Python/Shell/Perl scripting.
  • Proficiency in implementing CI/CD pipelines with tools such as git and Jenkins.
  • Advanced knowledge of networking (firewalls, DNS, Load Balancing, Proxies, etc.)
  • Advanced understanding of Linux operating system including shell scripting - Able to use core Linux commands and tools to build automation scripts.
  • Ansible – Writing playbooks and using Ansible core modules
  • Excellent interpersonal, organizational and communication (written, verbal, and presentation) skills are a must.
  • Self-motivated and results oriented with excellent analytical, problem solving, interpersonal, presentation and communication skills.

Nice To Haves

  • UI/UX experience to provide oversight on best practices for tooling to be used by the production support team across Global Market.
  • Infrastructure as Code (IaC): Hands-on experience with Terraform for automating infrastructure deployment.
  • Background in large Enterprise experience.
  • Candidate should be able to connect the dots and conclude complex Infrastructure issues quickly.
  • Able to work in a fast-paced environment whilst meeting deadlines.

Responsibilities

  • Operate hands on across the stack to improve platform and application observability, drive reliability improvements, and deliver measurable gains in operational efficiency across Global Markets.
  • Work closely with Quartz core teams to execute on platform modernization, harden production systems, and evolve support tooling.
  • Design, develop, test, and implement secure, robust, highly available and scalable solutions for Global Market applications and platforms.
  • Design and implement deployment approaches using highly scalable, automated, continuous integration and continuous delivery pipelines.
  • Resolve complex problems, owning the issue until you are sure it will not reoccur.
  • Proactively utilize SRE practices, service level indicators, and service level objectives to resolve issues before they impact customers.
  • Gather, analyze, synthesize, and develop visualizations and reporting from large, diverse data sets in service of continuous improvement of the platform.
  • Identify opportunities to eliminate toil and automate the triage of issues to improve overall operational stability.
  • Collaborate with a global team to identify, analyze, and resolve platform vulnerabilities.
  • Proactively promote the adoption of site reliability engineering best practices within the team and organization.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service