AI Software Intern

Tenstorrent University Jobs•Austin, TX
•$50 - $70•Onsite

About The Position

This is the team that makes "it runs on Tenstorrent hardware" actually true for real models, at real scale, and builds the infrastructure that keeps the rest of engineering moving fast. This role is on-site based out of Austin, TX or Santa Clara, CA. This posting spans multiple teams within Kernels, Models, Inference, Scaleout & our Runtime teams. One application, one recruiter screen, then we match you to the specific team and location that fits best. Teams include: Kernels: Develops high performance kernels on Tenstorrent hardware Models: Optimizes ML models (LLMs, vision models, video and image generation, and other architectures) for our hardware Inference Server: development and serving-side optimization Runtime: Builds the software engine that manages memory, task scheduling, and code execution on hardware while an application is actively running. Scale Out: communication and coordination between devices in a distributed AI system. Compiler & Infra: Create tools that optimize AI models into high-performance programs on Tenstorrent hardware, covering memory planning, profiling, debug, and emulation.

Requirements

  • Currently pursuing a BS, MS, or PhD in Comp Sci, Comp Eng, Physics/Math or a related field.
  • Coursework or projects in parallel processing, machine learning or distributed systems.
  • High comfort level as an end user of AI and agentic flows.
  • Experience with model quantization, kernel fusion, or other optimization techniques; or low level high performance programming.
  • Hands on experience working with Pytorch, experimenting with and deploying models, for inference or training.
  • Experience with C/C++ or kernel development in languages like CUDA, or Python and at least one ML framework (PyTorch, TensorFlow, JAX).

Nice To Haves

  • Knowledge of RTL, HDL is nice to have on certain teams.

Responsibilities

  • Develops high performance kernels on Tenstorrent hardware
  • Optimizes ML models (LLMs, vision models, video and image generation, and other architectures) for our hardware
  • Development and serving-side optimization for Inference Server
  • Builds the software engine that manages memory, task scheduling, and code execution on hardware while an application is actively running
  • Handles communication and coordination between devices in a distributed AI system
  • Create tools that optimize AI models into high-performance programs on Tenstorrent hardware, covering memory planning, profiling, debug, and emulation

Benefits

  • Highly competitive compensation package
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service