AI Systems Engineer (OCI/AI Infrastructure)

OracleNashville, TN
$96,800 - $306,400Onsite

About The Position

Oracle Hardware Platform Development Engineering is seeking a highly driven AI Systems Engineer to evaluate and characterize next-generation GPU and AI accelerator platforms for Oracle Cloud Infrastructure (OCI). This is a hands-on engineering role focused on bringing up new hardware platforms, enabling AI training and inference software stacks, running representative workloads, and analyzing system performance under real operating conditions. The engineer will identify whether workloads are HBM/memory-bandwidth, compute, scale-up, or scale-out bound, while characterizing power, thermals, memory behavior, utilization, scaling, and performance efficiency. Working directly in the lab, you will debug hardware/software integration issues, design and execute experiments, and develop data-driven insights that explain system behavior beyond benchmark results. A key part of the role is comparative architecture analysis across GPUs and emerging AI accelerators. You will evaluate architectural tradeoffs and translate performance findings into clear, actionable recommendations on which platforms are best suited for specific AI training and inference workloads. You will work closely with internal hardware and software teams as well as technology partners to help shape Oracle’s next generation of high-performance AI infrastructure. Position Overview: This position is ideal for someone who loves deep systems engineering, debugging complex hardware–software interactions, and optimizing performance at every layer of the ML stack. You will play a pivotal role in enabling the training and deployment of next-generation LLMs and generative AI models.

Requirements

  • Deep systems engineering experience.
  • Experience debugging complex hardware–software interactions.
  • Experience optimizing performance at every layer of the ML stack.

Responsibilities

  • Evaluate and characterize next-generation GPU and AI accelerator platforms for Oracle Cloud Infrastructure (OCI).
  • Bring up new hardware platforms.
  • Enable AI training and inference software stacks.
  • Run representative workloads and analyze system performance under real operating conditions.
  • Identify whether workloads are HBM/memory-bandwidth, compute, scale-up, or scale-out bound.
  • Characterize power, thermals, memory behavior, utilization, scaling, and performance efficiency.
  • Debug hardware/software integration issues.
  • Design and execute experiments.
  • Develop data-driven insights that explain system behavior beyond benchmark results.
  • Perform comparative architecture analysis across GPUs and emerging AI accelerators.
  • Evaluate architectural tradeoffs.
  • Translate performance findings into clear, actionable recommendations on which platforms are best suited for specific AI training and inference workloads.
  • Work closely with internal hardware and software teams as well as technology partners to help shape Oracle’s next generation of high-performance AI infrastructure.
  • Enable the training and deployment of next-generation LLMs and generative AI models.

Benefits

  • Flexible medical
  • Life insurance
  • Retirement options
  • Volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service