Staff Systems Engineer, AI/ML

GlobalFoundriesAustin, TX
$106,000 - $205,000Onsite

About The Position

We are looking for a seasoned Staff AI/ML Systems Engineer to lead workload-driven architecture strategy across hardware and software boundaries. You will define how we study, model, and optimize AI/ML workloads for current and next-generation products, drive alignment across HW and SW engineering organizations, and serve as a technical authority on performance and architecture tradeoffs. This is a senior individual contributor role with significant cross-functional scope and organizational influence.

Requirements

  • A BS or MS (MS preferred) in Electrical Engineering, Computer Engineering, Computer Science, or equivalent, with 4+ years of industry experience in systems engineering, hardware architecture, ML systems, or performance engineering, with a track record of technical leadership.
  • Exceptional mathematical reasoning is a core requirement at this level. You should be able to derive and defend analytical performance models from first principles, reason rigorously about the numerical behavior of quantized and sparse models, construct bandwidth-latency tradeoff curves across memory hierarchy levels, and identify when an approximation in a model is safe versus misleading. You will also be expected to evaluate the mathematical soundness of others' models and call out gaps in cross-functional reviews.
  • Deep expertise in CPU and SoC architecture is expected — you should be fluent in how modern processors handle memory hierarchies, out-of-order execution, vector/SIMD pipelines, and power management, and understand how these interact with AI/ML workloads. You should have strong command of memory bandwidth constraints at the system level (DDR/LPDDR bandwidth, channel configuration, utilization efficiency) and know how to reason quantitatively about when workloads are memory-bound vs. compute-bound.
  • You have built and validated analytical performance models (roofline, bandwidth-latency, first-principles throughput models) and know their limits. You have experience with AI/ML acceleration on edge devices — NPUs, dedicated inference accelerators, DSP-based pipelines — and understand the HW/SW co-design challenges involved. Experience with model quantization, sparsity, or other efficiency techniques and their interaction with hardware capabilities is a strong plus.
  • Familiarity with AI compiler infrastructure is preferred and increasingly important in this role. Experience with MLIR-based toolchains, IREE, TVM, or equivalent compilation and lowering pipelines — understanding how high-level graph representations are transformed, tiled, scheduled, and lowered to hardware — will meaningfully improve your ability to engage with software teams and identify where compiler strategy and hardware architecture must be co-designed. Prior work contributing to or evaluating such toolchains is a significant differentiator.
  • You are an effective cross-functional collaborator who can drive technical consensus without direct authority. You write clearly, present persuasively, and can calibrate your level of technical depth for different audiences.
  • Language Fluency - English (Written & Verbal)
  • Travel - Up to 10% of the time
  • This is a 100% in-office role (Dallas/Austin/San Jose)

Nice To Haves

  • Prior experience defining or co-defining SoC architecture requirements from workload analysis, contributions to MLIR/IREE or similar compiler infrastructure in a performance or backend capacity, contributions to internal or external publications or technical standards, and experience mentoring and growing junior systems engineers.
  • Knowledge of RISC-V architecture and Vector/Matrix extensions is a strong plus.

Responsibilities

  • Own the end-to-end process of workload characterization and hardware performance analysis for AI/ML systems — from selecting the right representative workloads and defining measurement methodology, to building analytical models that project system-level KPIs against candidate architectures. Your findings will directly inform SoC architecture decisions, memory subsystem design, and HW/SW co-optimization strategy.
  • Lead architectural discussions with hardware teams (CPU, SoC, memory, interconnect) and software teams (compilers, runtimes, ML frameworks), serving as the connective tissue between workload reality and design decisions. You will identify where the critical bottlenecks lie — whether in compute throughput, DRAM bandwidth, on-chip memory capacity, data movement latency, or software overhead — and build the case for specific architectural changes or optimization investments.
  • Define the performance KPI framework for AI/ML workloads across the product portfolio: what metrics matter, how to measure them accurately, how to estimate them pre-silicon, and how to use them to make architectural bets. You will set the standard for how the team does this work and mentor more junior engineers in applying it.
  • Regularly present findings and recommendations to senior engineering leadership and product stakeholders. Your communication must bridge deep technical content and strategic implication — you should be as comfortable writing a one-page architectural recommendation as a detailed technical memo.
  • Perform all activities in a safe and responsible manner and support all Environmental, Health, Safety & Security requirements and programs.

Benefits

  • GlobalFoundries makes possible the technologies and systems that transform industries and give customers the power to shape their markets.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service