Senior AI Performance Engineer

Graphcore•Austin, TX
•Hybrid

About The Position

Turn complex AI workload signals into faster, more efficient systems at data center scale. As a Senior AI Performance Engineer, you will analyze and optimize AI training and inference workloads across large-scale distributed systems. Your work will connect compute, memory, communication and software behavior. You will own complex performance investigations from problem definition through validation. You will turn profiling, benchmarking and modeling results into improvements engineers can build from. You will design benchmarks, investigate system bottlenecks and optimize distributed communication software across technologies such as MPI, NCCL, UCX and RDMA. Your work will help Graphcore deliver efficient, reliable AI systems for large-scale deployment. This role gives you rare scope across hardware, software, networking and system architecture. It is based in Austin, Texas. The team and culture The System Engineering Performance team architects, evaluates and optimizes high-performance infrastructure for large-scale data center deployments. The team works across the computing stack to understand real system behavior. Work moves through evidence, ownership and clear technical judgment. You will define investigations, coordinate across teams and validate impact with reliable performance data. Decisions are shaped through benchmarks, models, simulations and practical engineering tradeoffs. The team values engineers who think big, act fast, take responsibility, speak up and lead beyond their own area.

Requirements

  • Strong experience profiling and optimizing AI, machine learning or high-performance computing workloads.
  • Experience with distributed systems and communication libraries such as MPI, NCCL, UCX or libfabric.
  • Strong C++ and Python skills, including reliable tools or performance-sensitive software.
  • Deep understanding of compute, memory and communication behavior in large-scale systems.
  • Ability to own complex technical work and coordinate improvements across teams.

Nice To Haves

  • Familiarity with MLPerf, accelerated architectures, ML frameworks or high-performance interconnects.
  • Transferable skills and diverse experiences.
  • Engineers returning to the profession after a career break, including through returnship routes.

Responsibilities

  • Analyze and optimize AI training and inference workloads across large-scale distributed systems.
  • Connect compute, memory, communication and software behavior.
  • Own complex performance investigations from problem definition through validation.
  • Turn profiling, benchmarking and modeling results into improvements engineers can build from.
  • Design benchmarks, investigate system bottlenecks and optimize distributed communication software across technologies such as MPI, NCCL, UCX and RDMA.
  • Define investigations, coordinate across teams and validate impact with reliable performance data.

Benefits

  • Medical, dental, and vision coverage, with options that may extend to eligible dependents.
  • Mental health, wellness, and employee assistance resources.
  • Retirement savings benefits and company contributions where applicable.
  • Paid vacation, sick time, company holidays, and parental or family leave in accordance with applicable plans and policies.
  • Life insurance and short-term or long-term disability coverage.
  • Flexible working hours and hybrid working arrangements where compatible with the role and team requirements.
  • Professional-development resources, learning programs, office amenities, and team-led activities.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service