Senior Manager, Reliability & Platform Engineering

McMaster-CarrChicago, IL
$351,000 - $403,000Hybrid

About The Position

We're looking for an exceptional engineering leader to help define the future of engineering at McMaster-Carr, someone who can move from solving hard technical problems themselves to leading others through those same challenges. You'll join a department where site reliability is already championed by strong technical leaders across every team you touch, but this role carries a unique charge: setting the vision for how our reliability practices evolve company-wide. Success here means cultivating strong engineering talent, providing technical leadership across multiple simultaneous efforts, and influencing outcomes well beyond your own direct contributions. Our engineering teams operate within domains: distinct, high-impact areas of our platform that let engineers dive deep, build expertise, and release work that matters. In this role, you and your team will shape the reliability, performance, and security of the hybrid systems that power these domains, spanning on-premise data centers, cloud platforms, and the connective tissue between them. You'll partner closely with teams across Hybrid Infrastructure, mcmaster.com, Customer Navigation, AI & Data Systems, and Fulfillment & Automation. Across these domains, you’ll drive clarity, reduce operational friction, and build the guardrails that allow teams to move quickly without compromising reliability or security.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or Platform Engineering.
  • Deep ownership of production systems.
  • Direct experience managing and developing engineers.
  • A track record of growing engineers through complex, large-scale systems work.
  • Experience leading systemic improvements: dissecting complex failures across distributed systems, isolating root causes, and reducing toil.
  • Deep experience designing, building, and operating resilient, large-scale distributed systems.
  • Expertise operating hybrid environments, including on-premise data centers and cloud platforms (AWS, Azure, or GCP).
  • Strong command of networking fundamentals, including routing, load balancing, DNS, VPNs, service mesh, and zero-trust architectures.
  • Demonstrated experience implementing cloud security best practices, such as IAM design, secrets management, network segmentation, and workload hardening.
  • Ability to lead cross-functional initiatives with engineering and operations teams to translate architecture into business impact.

Nice To Haves

  • Knowing when to dive into the architecture yourself and when to maximize leverage by coaching someone else through it.
  • Understanding how to cut through noise to isolate root causes.

Responsibilities

  • Setting the vision for how reliability practices evolve company-wide.
  • Cultivating strong engineering talent.
  • Providing technical leadership across multiple simultaneous efforts.
  • Influencing outcomes well beyond your own direct contributions.
  • Shaping the reliability, performance, and security of hybrid systems spanning on-premise data centers, cloud platforms, and the connective tissue between them.
  • Partnering with Hybrid Infrastructure teams to extend our on-premise footprint and evolve toward cloud and colocation.
  • Revolutionizing foundational compute, storage, and networking layers.
  • Owning the reliability and performance of mcmaster.com.
  • Strengthening the reliability of search, browsing, and systems that help customers navigate millions of SKUs.
  • Scaling AI agents and AI infrastructure for low-latency inference, secure data flows, and resilient distributed systems.
  • Building and hardening systems that integrate with warehouse automation, delivery orchestration, and customer service operations.
  • Driving clarity, reducing operational friction, and building guardrails for teams.
  • Understanding current architecture, identifying opportunities to improve performance, and building fluency in operational tooling.
  • Improving targeted components.
  • Taking ownership of ambiguous, cross-cutting challenges.
  • Partnering with engineering leaders to define long-term architectural direction.
  • Designing resilient patterns for hybrid service connectivity.
  • Leading efforts to harden cloud and on-premise environments.
  • Driving incident response maturity.
  • Upskilling your team through mentorship, technical coaching, and creating an environment where engineers can do their best work.
  • Creating clarity where requirements are fuzzy.
  • Building momentum across teams.
  • Delivering durable solutions that raise the reliability bar for the entire organization.

Benefits

  • 100% tuition reimbursement
  • Informal and formal mentorship
  • Employee resource groups
  • Medical, dental, pharmacy, and vision plans with no monthly premiums
  • Inclusive, all-gender benefits
  • Paid parental leave for all new parents
  • Adoption and surrogacy assistance
  • First-time home buyer assistance
  • Industry-leading company-funded retirement accounts
  • Paid vacation and personal time
  • Relocation stipend (if applicable)
  • Signing bonus
  • Profit sharing based on company profitability
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service