Staff, Reliability Engineer

TenstorrentToronto, ON
Hybrid

About The Position

Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI. This role is hybrid, based out of Toronto, Canada. We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.

Requirements

  • 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems.
  • Comfortable with statistical analysis, including HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA.
  • Ability to work through technical problems in a thermal lab and clearly explain risks and trade-offs to leadership.
  • Skilled at bringing people together (mechanical, electrical, thermal, software, and System Dev & Compliance Validation teams), especially under tight timelines.
  • Bachelor's or Master's in Mechanical Engineering, Electrical Engineering, Reliability Engineering, or a related field.

Responsibilities

  • Set the reliability strategy for our next-generation AI computing systems, with MTBF modeling as a core piece.
  • Lead root-cause investigations and drive fixes across engineering and the supply chain.
  • Summarize findings and feed them back to the Systems Engineering design team, reliability as an ongoing loop, not a one-time check.
  • Partner closely with System Dev & Compliance Validation, helping set hardware up for success ahead of formal testing and certification.
  • Travel to third-party test facilities for HALT/HASS, to preplan, oversee testing, resolve DUT issues, and assess design risk in person.

Benefits

  • Highly competitive compensation package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service