AI Solutions Engineer

Johns Hopkins Applied Physics Laboratory•Laurel, MD
•Onsite

About The Position

The AI Solutions Engineer combines AI platform engineering with hands-on researcher enablement. You will serve as a technical bridge between researchers and the organization's scalable AI computing platform. You will also work directly with researchers to move AI workloads from experimentation to large-scale GPU execution, providing expertise across Jupyter, containers, Kubernetes, GPU scheduling, distributed training, and AI frameworks.

Requirements

  • Hold a Bachelor of Science degree or equivalent years of related professional work experience.
  • Have at least three (3) year experience using container-based orchestration (e.g. Kubernetes) and run-time environments (e.g. Docker).
  • Have hands-on experience with Jupyter environments and GPU-accelerated AI frameworks such as PyTorch, TensorFlow, Hugging Face, or similar technologies.
  • Have working knowledge of Kubernetes and GPU computing concepts sufficient to configure, integrate, and troubleshoot AI applications running in shared computing environments.
  • Have experience supporting researchers or developers with model training, fine-tuning, inference, workload optimization, and troubleshooting.
  • Have proficiency with Python and shell scripting for automation, troubleshooting, and platform integration.
  • Have working knowledge of machine learning concepts and model development workflows.
  • Can demonstrate strong problem-solving, communication, collaboration, prioritization, and continuous-learning skills.
  • Are able to work effectively with leadership to prioritize competing tasks.
  • Are able to obtain Interim Secret level security clearance by your start date and can ultimately obtain Secret level clearance. If selected, you will be subject to a government security clearance investigation and must meet the requirements for access to classified information. Eligibility requirements include U.S. citizenship.

Nice To Haves

  • Experience with AI workload orchestration and GPU scheduling platforms such as Run:ai, Slurm, Kubernetes-based GPU scheduling, or equivalent technologies.
  • Experience with NVIDIA GPU platforms and technologies such as CUDA, NVIDIA GPU Operator, NCCL, or NVIDIA Container Toolkit.
  • Experience supporting distributed AI training across multiple GPUs and/or multiple compute nodes.
  • Experience building and supporting containerized Jupyter environments and creating reusable AI/ML development environments.
  • Experience with AI/ML experiment tracking and lifecycle platforms such as MLflow, Weights & Biases, Kubeflow, or similar technologies.
  • Understanding of high-performance storage and networking considerations for GPU-intensive AI workloads.
  • Familiarity with a cloud computing platform (AWS, Azure or GCP).

Responsibilities

  • Own the technical configuration and evolution of the AI platform by evaluating new capabilities, determining appropriate adoption, and establishing application-level configurations, standards, and best practices.
  • Partner with Linux and Kubernetes administrators on platform deployments, upgrades, infrastructure integration, and troubleshooting issues that span the AI platform and underlying computing environment.
  • Provide hands-on technical assistance to researchers and data scientists using GPU-accelerated AI/ML platforms for model development, training, fine-tuning, and inference.
  • Troubleshoot AI workloads across the stack, including Python environments, Jupyter, containers, Kubernetes, GPU scheduling and allocation, storage, networking, and AI frameworks.
  • Help researchers optimize GPU utilization and scale workloads from interactive or single-GPU experimentation to multi-GPU and multi-node distributed execution.
  • Develop and maintain reusable environments, workflows, automation, documentation, and operational guidance that improve platform reliability, usability, and researcher productivity.

Benefits

  • robust education assistance program
  • unparalleled retirement contributions
  • healthy work/life balance
  • retirement plans
  • paid time off
  • medical
  • dental
  • vision
  • life insurance
  • short-term disability
  • long-term disability
  • flexible spending accounts
  • education assistance
  • training and development
  • sign-on bonus
  • relocation benefits
  • locality allowance
  • discretionary payments for exceptional performance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service