Senior Engineering Manager, Site Reliability

Upstart
$195,300 - $270,400Remote

About The Position

Upstart is a leading AI lending marketplace that partners with banks and credit unions to expand access to affordable credit through technology. The company is digital-first, offering flexibility in work location while fostering connection through team events. Upstart is seeking individuals energized by tackling meaningful problems and motivated by work that matters. The Site Reliability Engineering (SRE) team at Upstart is responsible for ensuring the reliability, observability, and resilience of the company's systems at scale. They manage company-wide incident response, set reliability standards, oversee operational readiness, and develop capabilities to help engineering teams manage production issues. The SRE team aims to integrate reliability into the software development lifecycle, providing engineering teams with the necessary signals, automated safeguards, and operational practices to innovate rapidly while protecting customers and the business. Their work involves advancing observability, incident detection and response, service level objectives, operational readiness, and implementing systemic improvements based on incident learnings. SRE collaborates with product engineering, infrastructure, security, and platform teams to enhance reliability across the organization.

Requirements

  • 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering
  • Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems
  • Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes
  • Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations
  • Experience leading high severity incident response and improving incident management practices at scale
  • Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes
  • Track record of hiring, developing, and retaining high performing engineers and engineering leaders
  • Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams

Nice To Haves

  • Experience operating large scale, highly available distributed systems
  • Experience implementing or evolving service-level objectives and error-budget practices
  • Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies
  • Experience developing incident management, operational readiness, or resilience programs across a large engineering organization
  • Familiarity with Kubernetes, AWS, and modern cloud native architectures
  • Experience supporting major platform or architectural transitions
  • Strong product mindset when building internal reliability capabilities
  • Experience establishing executive level reliability reporting and operating reviews

Responsibilities

  • Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering
  • Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function
  • Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures
  • Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts
  • Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership
  • Set a high bar for technical quality, operating rigor, and executive communication
  • Develop engineers and leaders who can independently own complex reliability initiatives
  • Evolve Upstart’s incident management program to improve detection, response, coordination, communication, and recovery
  • Establish clear standards for managing high severity incidents and provide visible leadership during critical events
  • Improve postmortem quality and ensure incident learnings result in durable engineering improvements
  • Identify recurring failure patterns and drive systemic solutions across teams
  • Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction
  • Improve the quality, accessibility, and trustworthiness of signals used to understand production health
  • Drive consistent practices across metrics, logs, traces, alerting, and service health
  • Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions
  • Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil
  • Define measurable reliability outcomes and use data to prioritize investments and communicate impact
  • Partner with platform and product engineering teams to embed reliability into standard engineering workflows
  • Establish scalable operational readiness standards for new services, major launches, and architectural changes
  • Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response
  • Identify systemic reliability risks and partner with engineering teams to prioritize and address them
  • Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards
  • Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through
  • Align stakeholders and dependencies before critical launches and engineering decisions

Benefits

  • Competitive compensation, including base pay, bonus opportunities, and annual equity grants that vest quarterly
  • Retirement benefits to help you plan for the future, including a 401(k) or Group Retirement Savings Plan with a company match of $2 for every $1 contributed, up to $15,000 annually (USD in the US, CAD in Canada)
  • Employee Stock Purchase Plan (ESPP) with discounted stock purchase options for eligible employees (US only)
  • Comprehensive health coverage designed to support you and your family, including medical, dental, vision, and wellness resources for US and supplemental health coverage for Canada.
  • Health Savings Account contributions from Upstart for eligible plans (US only)
  • Income protection benefits, including life insurance and disability coverage for added financial security
  • Paid time off, sick leave, and company holidays, in line with local requirements
  • Paid family and parental leave to support caregiving and major life moments (duration varies by country)
  • Family-centered benefits to support fertility, parenthood, and caregiving needs
  • Employee Assistance Program (EAP) offering mental health support and life-centered resources
  • Financial wellness resources, including access to financial planning tools and a financial concierge service (US Only)
  • Annual wellness allowance to support your physical and emotional well-being and personal development, based on what matters most to you
  • Annual productivity allowance to invest in relevant tools and resources you need to do your best work, no matter where you work from
  • Connection and community through team events, all-company updates, and employee resource groups (ERGs)
  • Onsite perks, including catered lunches and fully stocked micro-kitchens when working from one of our offices in the Bay Area, Austin, Columbus, and New York City (opening Summer 2026!)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service