Compiler Engineer, Graph Compiler Performance Optimization

MetaMenlo Park, CA
$183,997 - $257,000

About The Position

In this role, you will be a member of the MTIA (Meta Training & Inference Accelerator) Software team and part of the bigger AI and Compute Foundations team. The Graph Compiler team drives the development of the top-of-stack compilation pipeline for MTIA — taking PyTorch models, tracing the model graph, lowering and optimizing through FX/Inductor, and mapping to high-performance kernels. Your specific focus will be on performance optimization within the graph compiler: designing and implementing compiler passes that maximize throughput and minimize latency for AI workloads on MTIA hardware. You will work closely with AI researchers to understand emerging model architectures and translate performance requirements into compiler optimization strategies. You will partner with hardware design teams to drive hardware-software co-design, ensuring the compiler exploits new silicon capabilities from day one. You will also collaborate with the Triton/DSL and LLVM compiler teams to deliver cross-stack performance improvements.

Requirements

  • Experience with C/C++ and Python programming
  • Experience in compiler development, performance optimization, or accelerating deep learning models on hardware architectures
  • Understanding of deep learning model execution (graphs, operators, data flow) and how compiler transformations affect end-to-end performance
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience

Nice To Haves

  • A Bachelor's degree in Computer Science, Computer Engineering, relevant technical field and 7+ years of experience in compiler development or accelerating deep learning models on hardware architectures OR a Master's degree and 4+ years OR a PhD and 3+ years of relevant experience
  • Experience with graph-level compiler optimizations such as operator fusion, graph scheduling, memory allocation optimization, dead code elimination, or constant folding in ML compiler stacks
  • Experience with PyTorch internals, PyTorch 2.0 compilation stack (TorchDynamo, FX IR, Inductor), or similar ML compilation frameworks (XLA, TVM, MLIR, Glow)
  • Experience with performance profiling and analysis: identifying compute/memory/I/O bottlenecks, understanding roofline models, and developing systematic approaches to performance tuning
  • Experience with traditional compiler optimizations (loop transformations, vectorization, parallelization, instruction scheduling) and how they apply to ML workloads
  • Experience with AI hardware accelerator architectures (GPUs, TPUs, or custom ASICs) and understanding of how hardware constraints inform graph-level optimization decisions
  • Experience working with deep learning frameworks (PyTorch, TensorFlow, JAX) and understanding their compilation and execution models
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)

Responsibilities

  • Design, implement, and validate graph-level compiler optimization passes targeting performance within the PyTorch Inductor / FX IR compilation pipeline for MTIA
  • Profile and analyze deep learning models to identify graph-level performance bottlenecks such as suboptimal fusion boundaries, excessive memory traffic, and scheduling inefficiencies — then develop compiler solutions to address them
  • Develop and extend automatic fusion strategies to unlock peak hardware utilization across MTIA chip generations
  • Implement memory optimizations including data placement strategies, memory footprint reduction, and data movement elimination to reduce latency and improve bandwidth utilization
  • Build and improve performance analysis tooling to accelerate optimization iteration cycles
  • Collaborate with hardware design teams on hardware-software co-design: informing hardware features from compiler needs and rapidly developing compiler support for new chip capabilities
  • Ensure compiler optimizations are portable and scalable across MTIA chip generations, enabling rapid software bring-up for new silicon
  • Partner with AI researchers to co-design model architectures and compiler optimizations, ensuring models run performantly on MTIA out of the box

Benefits

  • bonus
  • equity
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service