Senior Manager, Reliability & Platform Engineering

McMaster-CarrChicago, IL
Hybrid

About The Position

McMaster-Carr is trusted by industrial customers to help keep manufacturing lines running, operations moving, and customers innovating. We earn that trust by offering the right products, making them easy to find, and delivering them quickly. Our website, mcmaster.com, is a huge part of that promise, and it's earned a reputation among engineers for being fast, reliable, and refreshingly easy to use. Customers count on it whether they're replacing a critical part, keeping operations moving, or finding exactly what they need under real time pressure. At McMaster-Carr, engineering is central to everything we do. Our systems power a business that customers rely on every day, and the reliability, scalability, and performance of those systems directly shape the experience we deliver. We intentionally cultivate a culture focused on clear execution and long-term growth. We build and maintain nearly all our critical systems ourselves, outsourcing very little, because we believe engineers who own their systems end-to-end build better ones, and many of the systems running today have been evolving here for decades. That responsibility means engineering work starts with a deep understanding of the problem and its impact, grounded in clear ownership, open communication, and direct feedback. Our teams are trusted to make thoughtful decisions about how work gets done, balancing a high bar for quality with practical execution. We've built much of our infrastructure on-premise by choice, not default, and we're bringing that same diligence with us as we extend into cloud and colocation. Our leaders understand that great engineering organizations are not built by simply solving today’s problems. They are built by developing strong technical teams, creating systems that scale with the business, and establishing practices that allow engineers to do their best work.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or Platform Engineering, with deep ownership of production systems and direct experience managing and developing engineers.
  • A track record of growing engineers through complex, large-scale systems work, knowing when to dive into the architecture yourself and when to maximize your leverage by coaching someone else through it.
  • Experience leading systemic improvements: you dissect complex failures across distributed systems, cut through noise to isolate root causes, and reduce toil on our reliability team.
  • Deep experience designing, building, and operating resilient, large-scale distributed systems, from architecture and capacity planning through launch and iterative refinement, while remaining close to the details of execution.
  • Expertise operating hybrid environments, including on‑premise data centers, cloud platforms (AWS, Azure, or GCP), and a strong command of networking fundamentals, including routing, load balancing, DNS, VPNs, service mesh, and zero‑trust architectures.
  • Demonstrated experience implementing cloud security best practices, such as IAM design, secrets management, network segmentation, and workload hardening.
  • Ability to lead cross‑functional initiatives with engineering and operations teams to translate architecture into business impact.

Responsibilities

  • Setting the vision for how our reliability practices evolve company-wide.
  • Cultivating strong engineering talent.
  • Providing technical leadership across multiple simultaneous efforts.
  • Influencing outcomes well beyond your own direct contributions.
  • Shaping the reliability, performance, and security of the hybrid systems that power our domains, spanning on-premise data centers, cloud platforms, and the connective tissue between them.
  • Partnering closely with teams across Hybrid Infrastructure, mcmaster.com, Customer Navigation, AI & Data Systems, and Fulfillment & Automation.
  • Driving clarity, reducing operational friction, and building the guardrails that allow teams to move quickly without compromising reliability or security.
  • Understanding our current architecture, identifying opportunities to improve the performance of data flow across our systems, and building fluency in our operational tooling.
  • Improving targeted components—small enough to ramp quickly, substantial enough to matter.
  • Taking ownership of ambiguous, cross‑cutting challenges such as defining long‑term architectural direction, designing resilient patterns for hybrid service connectivity, leading efforts to harden cloud and on‑premise environments, and driving incident response maturity.
  • Upskilling your team through mentorship, technical coaching, and creating an environment where engineers can do their best work.
  • Creating clarity where requirements are fuzzy, building momentum across teams, and delivering durable solutions that raise the reliability bar for the entire organization.

Benefits

  • Profit sharing based on company profitability.
  • Relocation stipend (if applicable).
  • Signing bonus.
  • 100% tuition reimbursement.
  • Informal and formal mentorship.
  • Employee resource groups.
  • Medical, dental, pharmacy, and vision plans with no monthly premiums.
  • Paid parental leave for all new parents.
  • Adoption and surrogacy assistance.
  • First-time home buyer assistance.
  • Industry-leading company-funded retirement accounts.
  • Paid vacation and personal time.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service