Performance Architect, CPU Cluster

TenstorrentSanta Clara, CA
$100,000 - $500,000Remote

About The Position

Tenstorrent is a leader in cutting-edge AI technology, aiming to revolutionize performance, ease of use, and cost efficiency. As AI redefines computing, solutions must integrate innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team has developed a high-performance RISC-V CPU from scratch and is passionate about building the best AI platform. We value collaboration, curiosity, and solving complex problems. We are expanding our team and seeking contributors at all seniority levels. Tenstorrent is looking for a CPU Cluster Performance Architect to influence the performance and scalability of our next-generation CPUs. This role focuses on the interaction of multiple CPU cores within a cluster, including cache hierarchies, coherent interconnects, memory bandwidth, and overall system behavior. It's a hands-on architecture position for individuals who enjoy using performance models and simulations to identify bottlenecks, evaluate tradeoffs, and guide design decisions early in the process. You will collaborate with CPU architects, RTL designers, software and compiler teams, and system architects to understand workload behavior and translate performance insights into architectural improvements. This is an opportunity to make a significant impact by analyzing cluster performance scaling across cores, memory systems, and chiplets. This role is remote and based in North America. We are open to candidates at various experience levels, and the final level and offer will be determined during the interview process.

Requirements

  • Strong foundation in CPU microarchitecture and performance.
  • Understanding of how cores, caches, interconnects, and memory systems interact.
  • Experience using modeling, simulation, and workload analysis to understand cluster-level performance bottlenecks and scaling challenges.
  • Comfort moving between detailed microarchitecture and broader system-level questions around latency, bandwidth, quality of service, utilization, and scalability.
  • Analytical and hands-on approach with the ability to translate performance data into clear architectural recommendations.
  • Strong communication skills and enjoyment of cross-functional collaboration with architecture, RTL, compiler, software, and system teams.

Responsibilities

  • Analyze and optimize CPU cluster performance, including cache hierarchies, interconnects, coherence, and memory access behavior.
  • Build and use performance models and simulation environments such as Gem5 or equivalent to evaluate architectural concepts and identify performance opportunities.
  • Study real workloads to understand bottlenecks related to cache misses, coherence traffic, memory bandwidth, latency, contention, and core-to-core communication.
  • Lead architectural tradeoff studies across performance, scalability, bandwidth, latency, power, and implementation complexity.
  • Collaborate with CPU, cache, interconnect, memory, software, and system architects to translate performance analysis into concrete design decisions.

Benefits

  • Highly competitive compensation package
  • Benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service