HPC Engineer - Privacy Research Program

Tufts UniversityUNAVAILABLE, Massachusetts

About The Position

Tufts Technology Services (TTS) is seeking an HPC Engineer to design, develop, and deploy high-performance computing systems and applications. This role will focus on building and deploying new services and infrastructure for a secure and compliant research computing environment as part of the Privacy Program and Secure Research Computing Development Project. This is a two-year limited-term appointment with the possibility of extension. The engineer will be responsible for integrating new user-facing services into existing HPC web gateways, installing and testing software, utilizing configuration management and infrastructure as code, and collaborating with other technical staff on network, storage, virtualization, and firewall issues. The role also involves providing system administration services, consulting on innovative architectures, and maintaining connections with the broader research computing community.

Requirements

  • Bachelor's Degree in science or equivalent experience.
  • High School Diploma plus 8 or more years related experience in a higher education, research, scientific or technical computing environment.
  • Understanding of and experience with high performance computing, scientific gateways from both an architecture, subsystems and networking perspective as well as daily usage, support, and application-level knowledge.
  • Demonstrated experience maintaining technologies used in research and high-performance computing such as job schedulers (Slurm), Containers (singularity, docker), RDMA over ethernet, Infiniband, GPUDirect, etc.
  • Extensive experience with scripting (e.g., Shell, Batch, Perl, Python, etc.).
  • Broad experience installing, maintaining open source or commercial research computing web gateway solutions such as OpenOnDemand.
  • Strong experience installing, configuring, maintaining, troubleshooting common frameworks and software used in research and high-performance computing such as scikit-learn, TensorFlow/TensorBoard, Keras, Theano, Caffe, Pytorch, MXNet, DGL, GPU libraries such as NVIDIA RAPIDS suite (cuDF, cuML, cuGraph, cuDNN). on both GPU and CPU architectures. Management/monitoring frameworks such as DCGM.
  • Broad experience and resourcefulness with all aspects of the system management and development cycle from analysis through evaluation and documentation when approaching system engineering challenges.
  • Proficiency with modern system administration devops and design patterns to automate Linux HPC clusters, operating system, software installation via scripting as well as configuration management systems such as Ansible and Warewulf.
  • Willing and able to learn technologies and required domain knowledge at a rapid pace.
  • Background supporting academic researchers (e.g., faculty, staff, students, etc.).
  • Strong communication, presentation, customer service, problem-solving skills in pursuit of system management and innovation.
  • Demonstrated ability to work effectively in a dynamic, collaborative environment with colleagues and build partnerships across technical disciplines, job functions and departments.
  • Dedication to taking ownership of projects that include identifying problems, developing testing protocols, and developing and implementing solutions.

Nice To Haves

  • Master's Degree in science or engineering field plus 2 or more years related experience in a higher education, research, scientific or technical computing environment.
  • Familiarity and experience with resources at private or public sector HPC research computing environments or national centers.
  • Familiarity and experience with running HPC workloads in cloud computing environments.
  • Experience deploying or working with NIST 800-171 or CMMC compliant research computing environments.
  • Knowledge of the continuum of research computing and scalability from desktop to HPC, cloud and national centers.

Responsibilities

  • Build and deploy new secure, compliant services within the research computing ecosystem as identified as part of the Privacy Program and Secure Research Computing Development Project.
  • Integrate new user-facing services into existing HPC web gateways such as Open OnDemand, Jupyter Notebooks and RStudio.
  • Install, maintain and test open source and commercial software as needed for the new services being built in this project.
  • Utilize configuration management, infrastructure as code, and security best practices for all project work.
  • Document all work and provide regular progress updates.
  • Work closely with other RT and TTS technical staff to ensure new services are part of a cohesive plan, including network, storage administration, virtualization layer and firewall issues.
  • Respond to outage, emergency or urgent issues as needed.
  • Provide system administration services and consult with other team members to evaluate Proof of Concept (POC) systems to foster innovative architectures and solutions in the selection of the new software and hardware components needed for this project.
  • Maintain ties with the larger system administration and research computing community to better understand the paradigms, methods, and opportunities others are using to solve the problems identified in this project.

Benefits

  • Two-year limited-term appointment with possibility of extension.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service