MX is a fintech company focused on empowering individuals financially by building technology for banks, credit unions, and fintechs to offer improved financial experiences. The company is experiencing renewed momentum and growth, valuing thoughtful execution, innovation, and individual ownership. MX fosters a culture of curiosity, accountability, and impact, encouraging employees to question assumptions, design better solutions, and contribute to company growth. The role of Site Reliability Engineer is crucial as MX's infrastructure supports financial applications used by millions and processes billions of transactions, making reliability a core product. The company is establishing a new observability function that mirrors its incident response approach, where the system handles the bulk of the work and humans manage judgment, customers, and exceptions. This Senior Observability Engineer will build and operate an observability control plane, establishing baselines, scoring coverage, and leveraging incidents to improve platform detection. This is a multiplier role aimed at elevating the standards for all teams through automation rather than manual dashboard creation. The 'shepherd model' involves guiding Datadog usage and partnering with product engineering teams to ensure proper signal observation, providing service owners with clear insights and leadership with program metrics on coverage and health. This role includes shared on-call responsibilities, participating in the incident response roster, and acting as an Incident Commander when needed, which is a fundamental aspect of the position.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
Associate degree