Senior Infrastructure Engineer, SRE

Rocket MoneyWashington, DC
$150,000 - $185,000Hybrid

About The Position

Rocket Money's mission is to empower people to live their best financial lives. Rocket Money offers members a unique understanding of their finances and a suite of valuable services that save them money and time. The Cloud Infrastructure team is looking to expand with a Senior Infrastructure Engineer, SRE to lead the reliability and operational evolution of their platform. This role will focus on building and improving the reliability and resiliency of systems and services, establishing SLIs, SLOs, and error budgets, owning and evolving the disaster recovery strategy, partnering with product engineering teams, evolving the observability platform and standards, strengthening the incident practice, and contributing to day-to-day Cloud Infrastructure work. The role supports millions of people and ensures the platform can operate reliably and at scale.

Requirements

  • 5+ years of hands-on cloud or infrastructure engineering experience, with substantial time spent on reliability and production operations at scale
  • Defined SLIs and SLOs for real production services, and can talk about what changed as a result.
  • Hands-on experience with an observability platform in production; Datadog strongly preferred
  • Comfortable writing code (Python, Go, TypeScript, or similar) for internal tooling, production debugging, and automation
  • Write production Terraform and are comfortable in AWS
  • Built or operated a disaster recovery plan: set the recovery goals, wrote the failover and restore steps, and ran the drills that proved it works
  • Been on-call for services you helped build

Nice To Haves

  • Led a reliability or observability modernization project where you defined the vision, approach, and delivered the implementation
  • Built internal tooling, libraries, or instrumentation standards that made it easier for other teams to operate their services well
  • Run game days, chaos experiments, or DR exercises, and fixed the problems they uncovered
  • Cut observability spend while keeping the coverage you needed

Responsibilities

  • Building and improving the reliability and resiliency of our systems and services
  • Establishing SLIs, SLOs, and error budgets for our most critical services and user journeys, and reviewing them regularly with the teams that own them
  • Owning and evolving our disaster recovery strategy: recovery objectives, failover and restore paths, and regular exercises that prove they work
  • Partnering with product engineering teams so they can own and operate their own services, with metrics that reflect real user experience
  • Evolving our observability platform and standards across metrics, tracing, and logs: including instrumentation paved roads, alert quality, and observability cost
  • Strengthening our incident practice: tuning paging thresholds, keeping runbooks current, and following through on postmortem action items
  • Contributing to day-to-day Cloud Infrastructure work alongside your reliability specialty — infrastructure build-outs, platform backlog, and a shared on-call rotation (1 week out of every 6 weeks)

Benefits

  • Health, Dental & Vision Plans
  • Competitive Pay
  • 401k Matching
  • Unlimited PTO
  • Lunch daily (in-office only)
  • Snacks & Coffee (in-office only)
  • Commuter benefits (in-office only)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service