High Performance Computing (HPC) AI Engineer

NYU Langone HealthNew York, NY
$101,494 - $140,000Onsite

About The Position

The High Performance Computing (HPC) AI Engineer will leverage technical expertise to bridge the gap between traditional high-performance computing and modern artificial intelligence workloads. Reporting to the Director of the NYULH HPC Facility, this role is critical to supporting the computational needs of world-class researchers in both clinical and basic science at NYU Langone Health (NYULH). The AI Engineer will focus on managing HPC development tools, optimizing AI applications, and seamlessly integrating AI frameworks with our broader computing infrastructure, including the UltraViolet supercomputer and our specialized GPU clusters.

Requirements

  • Bachelors Degree Required.
  • 5 Years of relevant experience.
  • Demonstrated understanding of parallel storage engineering and systems administration
  • Advanced skills with large-scale parallel file systems and storage tiering
  • Experience with SMB and NFS network storage protocols
  • Experience with Linux/UNIX operating systems and bash scripting
  • Extensive knowledge of ethernet, Infiniband, and fiber channel networking
  • Experience with scripting languages such as Bash, Python, or Lua
  • Qualified candidates must be able to effectively communicate with all levels of the organization.

Nice To Haves

  • Experience supporting AI-driven research computing in biomedical, life sciences, or clinical environments.
  • Familiarity with Large Language Models (LLMs), fine-tuning techniques, and Retrieval-Augmented Generation (RAG) pipelines.
  • Knowledge of security best practices, including the handling of Protected Health Information (PHI) when training or hosting models.

Responsibilities

  • Deploy, manage, and optimize HPC development tools and software stacks tailored for machine learning, deep learning, and data science workloads.
  • Lead the integration of HPC systems with enterprise AI frameworks and manage scalable AI model hosting and serving environments.
  • Benchmark, troubleshoot, and optimize AI application performance across our hardware infrastructure, ensuring efficient utilization of computational resources.
  • Develop and deliver specialized training sessions, tutorials, and workshops on AI model development, deployment, and optimization within an HPC environment.
  • Create and maintain comprehensive technical documentation, user guides, and best practices for AI development workflows and tools.
  • Collaborate with scientists, principal investigators, and software engineers to deploy customized, scalable AI solutions for complex research challenges.
  • Monitor AI job queues and system performance within resource management and scheduling systems to ensure equitable resource distribution.
  • Consult with researchers to assess AI computational needs, assisting with project execution, pipeline optimization, and resource planning.

Benefits

  • financial security benefits
  • a generous time-off program
  • employee resources groups for peer support
  • holistic employee wellness program
  • physical, mental, nutritional, sleep, social, financial, and preventive care
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service