This is the team that makes "it runs on Tenstorrent hardware" actually true for real models, at real scale, and builds the infrastructure that keeps the rest of engineering moving fast. This role is on-site based out of Austin, TX or Santa Clara, CA. This posting spans multiple teams within Kernels, Models, Inference, Scaleout & our Runtime teams. One application, one recruiter screen, then we match you to the specific team and location that fits best. Teams include: Kernels: Develops high performance kernels on Tenstorrent hardware Models: Optimizes ML models (LLMs, vision models, video and image generation, and other architectures) for our hardware Inference Server: development and serving-side optimization Runtime: Builds the software engine that manages memory, task scheduling, and code execution on hardware while an application is actively running. Scale Out: communication and coordination between devices in a distributed AI system. Compiler & Infra: Create tools that optimize AI models into high-performance programs on Tenstorrent hardware, covering memory planning, profiling, debug, and emulation.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Intern