The MAPS Quartz team owns the reliability and operational health of the Quartz platform. We actively monitor and support core infrastructure and components, execute automation initiatives, build and enhance internal tools, resolve infrastructure issues, and lead incident response to keep the platform running at scale. We are looking for a Site Reliability Engineer (SRE) who will operate hands on across the stack to improve platform and application observability, drive reliability improvements, and deliver measurable gains in operational efficiency across Global Markets. This role will work closely with Quartz core teams to execute on platform modernization, harden production systems, and evolve support tooling. This position is critical to maintaining execution velocity, reducing operational risk, and ensuring Quartz continues to meet its reliability and performance objectives. Position Summary Collaborates with a diverse set of engineers, architects, and teams to design, develop, test, and implement secure, robust, highly available and scalable solutions for Global Market applications and platforms. Collaborates other software engineers and teams to design and implement deployment approaches using highly scalable, automated, continuous integration and continuous delivery pipelines. Responsible for all aspects of reliability, collaborates with technical experts, key stakeholders, and team members to resolve complex problems, owning the issue until you are sure it will not reoccur. Deep understanding of SRE practices, service level indicators, and service level objectives; proactively utilize them to resolve issues before they impact customers. Gather, analyze, synthesize, and develop visualizations and reporting from large, diverse data sets in service of continuous improvement of the platform. Identify opportunities to eliminate toil and automate the triage of issues to improve overall operational stability. Collaborate with a global team to identify, analyze, and resolve platform vulnerabilities. Proactively promotes the adoption of site reliability engineering best practices within the team and organization.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed