Sr. Site Reliability Engineer

MX Technologies, Inc.Lehi, UT
Hybrid

About The Position

MX is a fintech company focused on empowering individuals financially by building technology for banks, credit unions, and fintechs to offer improved financial experiences. The company is experiencing renewed momentum and growth, valuing thoughtful execution, innovation, and individual ownership. MX fosters a culture of curiosity, accountability, and impact, encouraging employees to question assumptions, design better solutions, and contribute to company growth. The role of Site Reliability Engineer is crucial as MX's infrastructure supports financial applications used by millions and processes billions of transactions, making reliability a core product. The company is establishing a new observability function that mirrors its incident response approach, where the system handles the bulk of the work and humans manage judgment, customers, and exceptions. This Senior Observability Engineer will build and operate an observability control plane, establishing baselines, scoring coverage, and leveraging incidents to improve platform detection. This is a multiplier role aimed at elevating the standards for all teams through automation rather than manual dashboard creation. The 'shepherd model' involves guiding Datadog usage and partnering with product engineering teams to ensure proper signal observation, providing service owners with clear insights and leadership with program metrics on coverage and health. This role includes shared on-call responsibilities, participating in the incident response roster, and acting as an Incident Commander when needed, which is a fundamental aspect of the position.

Requirements

  • BS in Computer Science or equivalent experience
  • 5+ years running production observability, SRE, or DevOps
  • 5+ years automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency
  • AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs
  • Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis
  • Shared on-call, Incident Commander-capable

Nice To Haves

  • Fintech experience with MX-like architectures
  • Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast
  • Google SRE practices: toil elimination, incident management, automation for self-healing
  • Cross-functional influence without authority. You've improved teams that don't report to you
  • Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends)
  • OpenTelemetry instrumentation
  • Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
  • Golang and Ruby on Rails (the MX stack)

Responsibilities

  • Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
  • After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
  • Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
  • Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
  • Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
  • Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
  • Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
  • Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
  • Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
  • After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
  • Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.

Benefits

  • company-paid meals
  • sports simulator
  • gym
  • mother’s lounge
  • meditation room
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service