Engineering Manager, Site Reliability

Instacart
•$178,000 - $226,000•Remote

About The Position

Instacart is transforming the grocery industry by building technology that connects customers, shoppers, retailers, and brands through a reliable, convenient online marketplace. The Site Reliability Engineering team helps ensure that this experience remains resilient, scalable, and dependable as Instacart grows. We are seeking an Engineering Manager to lead a team of Site Reliability Engineers responsible for the systems, tools, and practices that support the reliability of Instacart’s technology platform. In this role, you will manage and develop a team of engineers while partnering closely with engineering teams across the company to improve availability, scalability, performance, observability, and operational excellence. This is an opportunity to shape the future of reliability at significant scale. You will help establish sound engineering practices, guide complex technical initiatives, and create an environment where teams can build and operate dependable systems. The role is well suited for a collaborative, hands-on people leader who is energized by complex problems, thrives in a fast-changing environment, and is motivated by helping others do their best work.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • Seven or more years of experience in software engineering, infrastructure engineering, Site Reliability Engineering, or a related field.
  • Two or more years of experience managing, mentoring, or leading engineering teams.
  • Professional experience with cloud infrastructure, distributed systems, networking, containers, orchestration platforms, or related production technologies.
  • Experience leading or participating in production incident response, post-incident reviews, reliability improvement initiatives, and operational readiness practices.
  • Experience communicating technical risks, priorities, and tradeoffs to engineering leaders and cross-functional stakeholders.

Nice To Haves

  • Experience leading Site Reliability Engineering, platform engineering, infrastructure engineering, or developer productivity teams.
  • Experience operating highly available services at significant scale and improving service-level objectives, observability, capacity, or disaster recovery capabilities.
  • Experience with infrastructure as code, continuous delivery, monitoring, logging, tracing, and automated remediation.
  • Experience building or evolving reliability programs across multiple engineering teams, including shared standards, operational reviews, and service ownership practices.
  • Demonstrated ability to create alignment across teams, navigate ambiguity, and turn complex technical challenges into clear, achievable plans.
  • A leadership approach grounded in empathy, transparency, direct communication, collaboration, and a commitment to inclusive team development.

Responsibilities

  • Lead, mentor, and develop a team of Site Reliability Engineers, establishing clear goals, providing actionable feedback, and supporting career growth and professional development.
  • Set the team’s technical direction and priorities for improving the reliability, scalability, availability, performance, and operational readiness of Instacart’s systems.
  • Partner with engineering, product, security, infrastructure, and other cross-functional teams to define reliability standards, influence system design, and deliver initiatives that improve the customer and developer experience.
  • Drive incident management and operational excellence, including incident response, post-incident learning, service-level objectives, capacity planning, observability, and continuous risk reduction.
  • Promote automation and self-service tooling that reduce operational toil, improve deployment confidence, and enable engineering teams to own and operate their services effectively.
  • Balance near-term operational needs with long-term investments, making thoughtful tradeoffs in a high-growth environment where priorities and requirements can change quickly.
  • Communicate clearly with technical and non-technical stakeholders, bringing transparency to reliability risks, project status, tradeoffs, and decisions.
  • Make decisions with incomplete information, respond calmly during incidents, and help teams learn from failure without assigning blame.

Benefits

  • Highly market-competitive compensation and benefits in each location where our employees work.
  • New hire equity grant as well as annual refresh grants.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service