AI HPC Infrastructure Engineer

Analysis GroupBoston, MA
Hybrid

About The Position

The AI HPC Infrastructure Engineer owns the operation, performance, and growth of a hybrid high-performance computing (HPC) and AI/GPU infrastructure environment. The engineer maintains the Linux-based clustered computing platform that supports both traditional HPC/analytical workloads and large-scale AI/ML training and inference, ensuring systems run efficiently, GPUs and other accelerators are current and well-utilized, and operations are monitored, documented, and reported — including change management and performance statistics — across both domains.

Requirements

  • Bachelor's degree required; degree in computer science, electrical engineering, or a related field preferred.
  • A minimum of 5 years of experience as a hands-on Linux Systems Administrator in a research, HPC, or production setting.
  • Experience managing Posit Workbench (RStudio Server Pro), Python, and R environments; strong Posit Workbench administration experience is a significant plus.
  • Experience with SLURM, Platform LSF, or other job schedulers required; experience scheduling GPU resources strongly preferred.
  • Hands-on experience with GPFS (IBM Spectrum Scale) required.
  • Demonstrated experience tuning LLM training and/or inference performance (e.g., batching, quantization, KV-cache management, parallelism strategies) required.
  • Excellent hardware troubleshooting experience, including GPU-specific diagnostics.
  • Knowledge of applicable data privacy practices and laws.
  • Strong interpersonal, written, and oral communication skills.
  • Highly self-motivated and directed, with keen attention to detail.
  • Proven analytical and problem-solving abilities.
  • Strong customer service orientation.
  • Experience working in a collaborative environment.
  • An inclusive and growth-oriented mindset, strong interpersonal skills, and an ability to work across functions.
  • To the extent permitted by applicable law, eligible candidates must be authorized to work in the United States, without sponsorship or restriction, now and in the future.

Nice To Haves

  • An ideal candidate will have 5 to 10 years of substantive relevant experience.
  • Hands-on experience with NVIDIA GPU infrastructure and software stack (CUDA, cuDNN, NCCL, NVIDIA GPU Operator) strongly preferred.
  • Experience with Kubernetes and container orchestration for AI/ML workloads highly desired.
  • Familiarity with ML/AI frameworks (PyTorch, TensorFlow) and distributed training patterns highly desired.
  • Experience with MLOps tooling (MLflow, Kubeflow, Weights & Biases, or similar) is a plus.
  • Experience with Bright Cluster Manager is highly desired.
  • Experience with Ansible is highly desired.
  • Experience with containerization (Docker, Singularity/Apptainer) is highly desired.
  • Proficiency with remote access technologies and tools such as RDP, SSH, and emulation software
  • Experience with AI Gateways (e.g., LiteLLM, Kong AI Gateway, Portkey, or similar) is a very nice to have.

Responsibilities

  • Maintain, tune, and manage the analytical and AI computing environment for researchers and data scientists, including Posit Workbench (RStudio Server Pro) environments.
  • Optimize systems and infrastructure performance using parallelization technologies (MPI, OpenMP) and distributed/multi-GPU training strategies (e.g., PyTorch Distributed, Horovod, DeepSpeed).
  • Design, deploy, and maintain GPU-accelerated compute infrastructure for large-scale model training and inference.
  • Manage GPU scheduling, multi-tenancy, and utilization across SLURM and/or Kubernetes-based environments.
  • Administer the NVIDIA software stack — drivers, CUDA, cuDNN, NCCL — and coordinate firmware and health monitoring across GPU fleets.
  • Tune and optimize LLM training and inference performance — including batching, quantization, KV-cache utilization, parallelism strategies, and throughput/latency across GPU clusters.
  • Build and maintain MLOps pipelines for model training, versioning, deployment, and monitoring (e.g., MLflow, Kubeflow).
  • Manage container orchestration and runtimes (Docker, Kubernetes, Singularity/Apptainer) supporting both HPC jobs and ML workloads.
  • Manage access authentication including PAM, LDAP integration, and single sign-on.
  • Design and develop scripts for system administration, automating tasks, monitoring, and usage reporting across HPC and AI resources.
  • Manage high-performance storage and data pipelines for AI training datasets and HPC workloads, primarily on GPFS (IBM Spectrum Scale).
  • Troubleshoot, isolate, and resolve application, systems, and other technical problems (hardware, software, network, and GPU-specific issues).
  • Develop and implement backup and recovery programs.
  • Research, deploy, and manage general infrastructure, including development of policies and procedures for both HPC and AI/ML environments.
  • Migrate data from heterogeneous environments to Linux, on-prem clusters, or cloud.
  • Collaborate with data scientists and ML engineers to support the model development lifecycle and translate research needs into infrastructure requirements.
  • Evaluate emerging AI hardware, accelerators, and cloud AI services, and recommend adoption where beneficial.
  • Monitor performance, troubleshoot problem areas, and provide statistics and reports across compute, storage, and network.
  • Create and maintain documentation related to system configuration, processes, change management, inventory, and service records.
  • Ensure continuous network connectivity of all equipment.
  • Conduct research and report on products, services, protocols, and standards to remain abreast of developments in HPC and AI infrastructure.
  • Participate in a 24x7 on-call rotation; troubleshoot and resolve issues remotely or onsite as necessary.

Benefits

  • competitive compensation
  • comprehensive benefits package
  • discretionary annual bonus
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service