Systems & Platform Engineer - TS/SCI

Sunayu•Bethesda, MD
•Onsite

About The Position

Sunayu, LLC is seeking a highly skilled platform engineer with deep expertise in operating systems, hardware, GPU, and high-speed networking. In this role, you will design, develop, and optimize Kubernetes clusters that power enterprise AI for mission customers. This is a 100% on-site position, requiring all work to be performed at the customer site in Bethesda at the Intelligence Community Campus.

Requirements

  • Bachelor's or higher degree in Computer Science, Electrical Engineering, or a related field. Additional years of experience may be considered in lieu of a degree.
  • 5+ years in Platform Engineering or System Engineering experience.
  • Strong expertise with Linux distributions (RHEL, Ubuntu, Oracle Linux, and Rocky).
  • Experience administering Kubernetes clusters, including deploying, scaling, and maintaining containerized workloads.
  • Hands-on experience creating, managing, and troubleshooting Docker containers and container images throughout the software development lifecycle.
  • Experience with Kubernetes cluster management and AI/ML workflow orchestration (Argo, Airflow, and Kubeflow).
  • Strong track record with consuming, and troubleshooting RESTful APIs for platform integration and automation.
  • Excellent problem-solving skills and the ability to collaborate within a team.
  • Candidate must, at a minimum, meet DoD 8570.11- IAT Level II certification requirements (currently Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP along with an appropriate computing environment (CE) certification). An IAT Level III certification would also be acceptable (CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, CCSP).
  • Active TS/SCI clearance with Polygraph required OR active TS/SCI and willingness to obtain and maintain a Poly.
  • US Citizenship is required.

Nice To Haves

  • Experience in managing NVIDIA GPU data center platforms (DGX, HGX, H200, H100, 200, B300, L40S).
  • Experience with NVIDIA enterprise tools such as Base Command Manager, Run:AI, Nvidia AI Enterprise.
  • Knowledge of enterprise server components (storage/network controllers, HBA, SSDs).
  • Familiarity with GPU virtualization and cloud computing.
  • Experience developing and deploying infrastructure in AWS.
  • Knowledge of distributed resource scheduling systems (Slurm, LSF, Open MPI, etc.).

Responsibilities

  • Design, configure, and maintain enterprise Kubernetes platforms.
  • Collaborate with a multidisciplinary team to define and optimize Kubernetes architecture, ensuring they meet performance, efficiency, and feature requirements.
  • Develop and manage Infrastructure as Code (IaC) using tools such as Terraform, Salt, Ansible, Bash, Python or similar frameworks.
  • Collaborate with development teams to design and implement secure, automated, and repeatable pipelines (e.g., GitLab CI/CD).
  • Troubleshoot complex systems issues across cloud, network, and platform layers.
  • Maintain technical documentation, architectural specifications, and Linux best practices.
  • Support ATO (Authority to Operate) and ensure compliance with federal security standards.

Benefits

  • 3 Medical Plan Options
  • Dental and Vision
  • FSA, DCFSA, HSA
  • Life/AD&D Insurance
  • Short-Term & Long-Term Disability
  • Employee Assistance Program (EAP)
  • Training and Educational Assistance
  • Paid Time Off (PTO)
  • 11 Federal holidays
  • 401k plan with up to a 6% match (100% immediate vesting)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service